Interaction processing method and device, electronic equipment, computer readable storage medium and computer program product
By performing speech encoding and feature denoising on the interactive audio and combining it with prompt words to generate reply text, the problem of insufficient paralinguistic information and contextual understanding in existing interactive systems is solved, the accuracy and anthropomorphism of the replies are improved, and the naturalness of human-computer interaction is enhanced.
Patent Information
- Application Number
- CN202510660073.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-23
AI Technical Summary
Existing interactive systems fail to fully understand the paralinguistic information and contextual content in the interactive audio when generating responses, resulting in misjudgment of user intent and generation of inappropriate responses.
By performing speech encoding on the interactive audio, extracting linguistic and paralinguistic features, and performing feature denoising, the response text is generated based on the prompt words, and a multimodal large language model is used to improve the accuracy of the dialogue system.
The accuracy and anthropomorphism of generated response texts are improved, enhancing the naturalness of human-computer interaction and user experience.
Smart Images

Figure CN120687559A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an interactive processing method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] In related technologies, interactive systems typically focus primarily on interactive content, neglecting the importance of emotional factors in communication. To enhance user experience, emotional understanding has become a key research direction for the next generation of interactive systems. Interactive systems with emotional understanding capabilities are crucial for achieving efficient and natural human-computer interaction. The key to developing an emotional interaction system lies in understanding historical interaction content and user expressions. By leveraging the text processing capabilities of large language models, multiple rounds of understanding of human-computer interaction can be achieved in interactive scenarios, thereby generating coherent responses. Related technologies focus solely on understanding semantic information, resulting in inaccurate responses generated in interactive systems. Summary of the Invention
[0003] The embodiments of the present application provide an interactive processing method, device, electronic device, computer-readable storage medium, and computer program product, which can effectively improve the accuracy of generated reply text.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] This embodiment of the present application provides an interaction processing method, including:
[0006] Performing a first speech encoding on the first audio to obtain a first linguistic feature and a first paralinguistic feature;
[0007] performing feature denoising of a corresponding paralinguistic dimension on the first linguistic feature to obtain a second linguistic feature, and performing feature denoising of a corresponding linguistic dimension on the first paralinguistic feature to obtain a second paralinguistic feature;
[0008] A reply text is generated based on the second paralinguistic feature, the second linguistic feature, and a first prompt word for the first audio, wherein the reply text is used to interact with the first audio according to an instruction of the first prompt word.
[0009] The present invention provides an interactive processing device, including:
[0010] an encoding module, configured to perform a first speech encoding on the first audio to obtain a first linguistic feature and a first paralinguistic feature;
[0011] a mapping module configured to perform feature denoising of a corresponding paralinguistic dimension on the first linguistic feature to obtain a second linguistic feature, and perform feature denoising of the corresponding linguistic dimension on the first paralinguistic feature to obtain a second paralinguistic feature;
[0012] A generation module is configured to generate a response text based on the second paralinguistic feature, the second linguistic feature, and a first prompt word for the first audio, wherein the response text is used to interact with the first audio according to an instruction of the first prompt word.
[0013] An embodiment of the present application provides an electronic device, including:
[0014] a memory for storing computer-executable instructions or computer programs;
[0015] The processor is used to implement the interactive processing method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0016] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which is used to implement the interactive processing method provided in the embodiment of the present application when executed by a processor.
[0017] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the interactive processing method provided in the embodiment of the present application is implemented.
[0018] The embodiments of the present application have the following beneficial effects:
[0019] By encoding the audio into a first linguistic feature and a first paralinguistic feature, the features of the two branches can be decoupled from the audio; then, the first linguistic feature is subjected to feature denoising of the corresponding paralinguistic dimension to obtain the second linguistic feature, and at the same time, the first paralinguistic feature is subjected to feature denoising of the corresponding linguistic dimension to obtain the second paralinguistic feature. By setting different vector dimensions for feature denoising, the paralinguistic information and the linguistic information can be effectively decoupled, reducing the information redundancy between the two features; finally, the second paralinguistic feature, the second linguistic feature, and the first prompt word for the first audio are used to generate a reply text, which can comprehensively consider the linguistic information, paralinguistic information, and instruction prompts in the audio, thereby improving the accuracy and anthropomorphism of the generated reply text. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 1 is a schematic diagram comparing generating a reply dialogue with and without considering paralinguistic information, provided by an embodiment of the present application;
[0021] Figure 2 1 is a schematic diagram of the architecture of the interactive processing system 100 provided in an embodiment of the present application;
[0022] Figure 3 is a structural diagram of an electronic device 500 provided in an embodiment of the present application;
[0023] Figure 4 This is a flow chart of the interactive processing method provided in the embodiment of the present application;
[0024] Figure 5 This is a flow chart of the interactive processing method provided in the embodiment of the present application;
[0025] Figure 6 This is a schematic diagram of the structure of the speech language large model provided in the embodiment of the present application;
[0026] Figure 7 is a schematic diagram of different adapter training provided in an embodiment of the present application;
[0027] Figure 8 This is a schematic diagram of the randomization of the order provided in the embodiment of the present application;
[0028] Figure 9 Schematic diagram of different adapter training after randomized combination provided in an embodiment of the present application;
[0029] Figure 10 This is a training diagram of style-aware behavior alignment provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0031] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0032] It is understandable that in the embodiments of the present application, when user information and other related data are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.
[0033] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0034] In the following description, the terms "first\second\..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second\..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0036] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0037] 1) Large Language Model (LLM): Large language models are deep learning-based models specifically designed for processing and generating text data. These models are typically trained on large text datasets and are able to understand complex language structure and semantics, resulting in excellent performance in a variety of natural language processing tasks, such as text generation, translation, question answering, and summarization. Key features of large language models include large-scale training, high-quality text generation, multi-task processing capabilities, and contextual understanding.
[0038] 2) Multimodal Large Language Model (MLM): A multimodal large language model (MLM) is an extension of the large language model. It can process and generate a variety of data types, including text, images, audio, and video. By integrating data from different modalities, these models provide richer and more comprehensive information processing capabilities, enabling them to be effective in a wider range of scenarios. The multimodal large language model not only inherits the text processing capabilities of the large language model but also expands its processing capabilities for data in other modalities.
[0039] 3) Speech Language Model: Speech language models are an important type of multimodal large language model that focuses on processing and understanding speech data. They primarily process speech signals and use deep learning techniques to extract and understand information contained in speech. These models can be applied to a variety of fields, including speech recognition, speech synthesis, sentiment analysis, and speaker recognition.
[0040] 4) Linguistic Information: Linguistic information refers to the linguistic content directly conveyed by speech, including vocabulary, grammar, and semantics. This information is expressed through speech units such as phonemes, syllables, words, and sentences, and is the basis for communication. For example, the linguistic information corresponding to the phrase "The weather is great today" is the speaker's positive assessment of the weather.
[0041] 5) Paralinguistic Information: Paralinguistic information does not directly relate to the linguistic content but is conveyed through non-verbal features of speech, such as pitch, volume, rate, and pauses. Paralinguistic information can provide additional information about the speaker's emotional state, attitude, and social context. For example, if someone says "I'm fine" in a low voice or slowly, this may convey that they are not actually well or are experiencing other emotional issues.
[0042] 6) Acoustic-Content Disentanglement (ACD): Acoustic content disentanglement is a technique that aims to separate the acoustic features and content features in speech signals.
[0043] 7) Acoustic characteristics: Acoustic characteristics include physical properties such as pitch, loudness, and timbre, which are related to the speaker's vocal organs, emotional state, speaking speed, etc.
[0044] 8) Content features: Content features include linguistic information such as vocabulary, grammar, and semantics, and are related to the specific content expressed by the speaker.
[0045] 9) Linguistic-Paralinguistic Information Disentanglement: Linguistic-paralinguistic information decoupling refers to separating the linguistic information and paralinguistic information in the speech signal.
[0046] 10) Instruction Tuning: Instruction tuning is an optimization method for large pre-trained language models. It aims to fine-tune the model through a specific set of instructions, enabling it to better understand and perform specific tasks. Instruction tuning is typically performed on a large-scale pre-trained model. By providing a series of instructions and corresponding output examples, the model learns how to generate the correct output based on the instructions. The main purpose of instruction tuning is to improve the model's performance on a specific task while maintaining its versatility on other tasks.
[0047] 11) Direct Preference Optimization (DPO): Preference alignment is a technique used to train large language models, aiming to align model outputs more closely with human preferences. DPO directly optimizes the model, leveraging human feedback to guide the learning process and ensuring that the content generated by the model is more aligned with human values and expectations. The primary goal of preference alignment is to align model-generated content more closely with human preferences, thereby improving the model's acceptability and safety.
[0048] In information interaction systems, the primary focus is often on the content of the interaction, while the importance of emotional factors in the communication process is often overlooked. To enhance user experience, emotion understanding has become a key research direction for next-generation interactive systems. Information interaction systems with emotion understanding capabilities are crucial for achieving efficient and natural human-computer interaction. The key to developing emotional interaction systems lies in understanding historical interaction content and user expressions. In related technologies, interaction systems with emotion understanding capabilities typically rely on text-based large language models (LLMs). By leveraging the text processing capabilities of LLMs, they achieve multi-round understanding of human-computer interactions in information interaction scenarios and generate coherent responses. Interaction systems based on LLMs still suffer from the problem of misinterpreting the interaction partner's intent. This is primarily because LLMs use text-based language as input and primarily consider the interaction partner's content, i.e., linguistic information. They largely fail to consider the paralinguistic information of the interaction content (e.g., intonation, speech rate, and other speaking styles). Paralinguistic information is crucial for understanding the emotion and intent of the interaction partner. Understanding paralinguistic information helps information interaction systems generate more accurate and appropriate responses. Without an understanding of the paralinguistic information in information interaction content, the information interaction system may misjudge the emotions or intentions of the interaction partner, thereby generating response content that does not match the user's needs.
[0049] For example, see Figure 1 , taking the dialogue system as an example, Figure 1 is a comparative diagram of generating a reply dialogue with and without considering paralinguistic information, as provided in an embodiment of the present application. Figure 1As shown, for a dialogue system receiving a high-pitched, cheerful female voice saying "I'm moving soon," if only the speaker's linguistic information (i.e., the content of the speech) is considered while ignoring paralinguistic information (i.e., the way it speaks), the system might misjudge the user's emotional state (for example, misclassifying excitement as anxiety) and thus give an inappropriate response: "Don't worry, moving is a hassle. You can ask some friends for help and it will go more smoothly." In contrast, considering both linguistic and paralinguistic information helps the dialogue system accurately understand the speaker's emotional state and generate a more appropriate response: "That's so exciting! How's the new place?" This highlights the importance of incorporating the speaker's voice as input in affective dialogue systems. The speaker's voice can be used to perceive not only the content of the speech but also the way it is spoken, leading to a more comprehensive understanding of the speaker's expression and the generation of more appropriate responses. Therefore, a key focus in the development of affective interaction systems is to build a large multimodal model that integrates speech and language, namely a large speech-language model.
[0050] However, building a large speech and language model is a challenging task, with two main technical approaches. One is to develop a speech-text based model that can natively process and understand speech and text. Although this approach is very effective, it requires large multimodal datasets, high computing resources, and advanced training techniques, which to some extent limits its widespread application. The other, more feasible technical approach is to expand the large text-based language model to enable it to understand speech input while inheriting the large language model's conversational context understanding capabilities and knowledge reserves. This approach usually involves integrating a speech encoder with a large language model and connecting the two through an adapter module to build a large speech and language model. The technical approach of expanding the large text-based language model has attracted widespread attention due to its advantages such as being able to leverage the capabilities of existing pre-trained models and having controllable training costs.
[0051] However, although the speech language model built by extending the text-based large language model has shown potential in some tasks, it still has shortcomings when applied to emotional interaction systems. Two major problems are:
[0052] (1) Insufficient understanding of paralinguistic information: Although the speech language model can effectively understand the linguistic information in the interactive audio by using the interactive audio as the model input, it is difficult to fully understand the paralinguistic information in the interactive audio;
[0053] (2) Insufficient understanding of contextual content: Compared with the large language model, the large speech language model obtained through extended training has a decreased ability to understand instructions and historical interaction content.
[0054] Due to the above two major problems, the information interaction system is prone to misjudge the user's interaction intention and generate inappropriate reply content.
[0055] In view of this, embodiments of the present application provide an interactive processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can effectively improve the accuracy of generated reply text. The electronic device provided in embodiments of the present application can be implemented as a server, or can be implemented collaboratively by a server and a terminal. The following description uses the interactive processing method provided in embodiments of the present application collaboratively implemented by a server and a terminal as an example.
[0056] For example, see Figure 2 , Figure 2 This is a schematic diagram of the architecture of the interactive processing system 100 provided in an embodiment of the present application, which is used to support an interactive processing application, such as Figure 2 As shown, the interactive processing system 100 includes: a server 200, a network 300, and a terminal 400. The terminal 400 is connected to the server 200 via the network 300. The network 300 can be a local area network or a wide area network, or a combination of the two.
[0057] In some embodiments, the user inputs the first audio and the first prompt word for the first audio through the terminal 400, and transmits the first audio and the first prompt word for the first audio to the server 200 through the network 300; then, the server 200 generates a reply text based on the first audio and the first prompt word for the first audio; finally, the server 200 transmits the reply text to the terminal 400 through the network 300 and displays it on the terminal 400.
[0058] The interaction processing method provided in the embodiments of the present application can be applied to various scenarios requiring voice interaction processing, for example, it can be applied to the following scenarios:
[0059] 1) Intelligent customer service scenarios. In customer service, timely and accurate handling of customer issues and emotions is crucial. By recognizing the emotional information in customer voices, intelligent customer service can provide more considerate and effective service, thereby improving customer satisfaction. For example, when a customer calls an e-commerce platform to complain about product quality issues and sounds angry and agitated, the intelligent customer service system will recognize the customer's anger and respond in a gentle, soothing tone, such as "I fully understand how you feel right now. Anyone would be angry if they encountered such a problem. Please calm down first, and we will definitely help you handle this issue properly." If the customer calmly inquires about the product return and exchange process, the intelligent customer service will provide relevant information in a professional and concise tone, such as "Hello, you can find the order on our app, click on the return and exchange application, and then follow the prompts."
[0060] 2) Educational scenarios: In online education, intelligent voice tutoring can provide personalized learning guidance based on students' emotional state, enhancing their learning motivation and effectiveness. For example, when a student encounters a difficult problem in learning mathematics and asks the intelligent learning assistant by voice, their tone reveals anxiety and frustration. After the intelligent learning assistant recognizes this emotion, it will first offer encouragement, such as "Don't worry, many students will encounter difficulties with this knowledge point, which is very normal. Let's take a look at this problem together and analyze it slowly." It will then gradually guide the student to solve the problem; if the student shows excitement and confidence during the learning process, the intelligent learning assistant will give affirmation and further challenges, such as "Your analysis is very good! It seems that you have mastered this knowledge point, so let's try a more difficult question."
[0061] 3) In the field of mental health counseling, intelligent voice-activated psychological assistants can provide companionship and support at all times for people with psychological distress, providing targeted psychological counseling by identifying users' emotional information. For example, when a user confides in an intelligent psychological assistant that they have recently felt depressed and helpless, the intelligent psychological assistant will recognize the user's negative emotions and respond with a warm and caring tone, such as "I can feel that you are feeling very upset right now, but this low mood is only temporary. Many people go through this stage, and you are not alone. You can tell me more about what has happened recently, and we can find a solution together." It will also provide some simple relaxation and mood regulation suggestions, such as deep breathing and listening to music.
[0062] 4) In the smart home sector, intelligent voice assistants can provide personalized home services based on the user's emotional state, enhancing their living experience. For example, when a user returns home exhausted and asks a smart speaker to play music, their tone of voice may be noticeably tired. Recognizing this emotion, the smart speaker will automatically play soothing, relaxing music, such as, "I sense you may be feeling a bit tired today. I'm playing some gentle music to help you relax." If the user is in a good mood and asks the smart speaker to play music, it will play upbeat, dynamic music.
[0063] 5) In the fields of virtual reality (VR) and augmented reality (AR), users can interact with characters or objects in the virtual or augmented reality environment through voice input. The text responses generated by the model can be used for dialogue between characters, providing a more realistic and immersive experience. For example, in a VR game, users can use voice to communicate with non-player characters (NPCs), and the model generates corresponding text responses based on the user's voice input, enriching the game plot.
[0064] 6) In the production of movies, games, and animations, the model can generate corresponding text responses based on audio input, which can be used to create sound effects and enrich the expressiveness of the work. For example, in animation production, the voice input of a voice actor can be used to generate corresponding text responses through the model to produce the character's dialogue, improving the quality of the work.
[0065] For example, Figure 2 The server 200 in the example can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, car terminal, etc., but is not limited to these. The terminal 400 and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0066] The following continues to describe the structure of the electronic device provided by the embodiment of the present application. Take the electronic device as an example, see Figure 3 , Figure 3 is a structural diagram of an electronic device 500 provided in an embodiment of the present application, Figure 3 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 540 is not shown in FIG. Figure 3 Various buses are labeled as bus system 540 .
[0067] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0068] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0069] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 510.
[0070] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0071] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0072] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0073] A network communication module 552 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB);
[0074] a presentation module 553 for enabling presentation of information via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0075] The input processing module 554 is configured to detect one or more user inputs or interactions from one of the one or more input devices 532 and to translate the detected inputs or interactions.
[0076] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 3 The interactive processing device 555 stored in the memory 550 is shown. It can be software in the form of a program or plug-in, and includes the following software modules: an encoding module 5551, a mapping module 5552, a generation module 5553, and a training module 5554. These modules are logical and can be arbitrarily combined or further divided according to the functions implemented. It should be noted that in Figure 3 For the sake of convenience, all the above modules are shown at once, but it should not be considered that the interactive processing device 555 excludes the implementation of only including the encoding module 5551, the mapping module 5552 and the generation module 5553. The functions of each module will be explained below.
[0077] The interactive processing method provided in the embodiment of the present application will be described in detail below in conjunction with the exemplary application and implementation of the terminal provided in the embodiment of the present application.
[0078] See also Figure 4 , Figure 4 This is a flow chart of the interactive processing method provided by the embodiment of the present application, which will be combined with Figure 4 The steps shown are explained.
[0079] It should be noted that Figure 4 The illustrated methods can be executed by various computer programs running on a terminal, not limited to a client. For example, they can also be executed by the operating system, software module, script, and applet described above. Therefore, the examples below using a client should not be considered as limiting the embodiments of the present application. In addition, for ease of expression, the following description does not specifically distinguish between a terminal and a client running on the terminal.
[0080] In step 101, a first speech encoding is performed on a first audio to obtain a first linguistic feature and a first paralinguistic feature.
[0081] It should be noted that the linguistic and paralinguistic features of the first audio can be extracted through a machine learning model, for example, using a Gaussian mixture model (GMM) to model different audio features, thereby achieving classification and recognition of different audio features, and using a support vector machine (SVM) to classify and process different audio features; deep learning methods can also be used to extract the linguistic and paralinguistic features of the first audio, for example, using pre-trained models (Wav2Vec 2.0, HuBERT, etc.) and neural network structures (convolutional neural networks CNN, recurrent neural networks RNN, Transformer, etc.) to extract the linguistic and paralinguistic features of audio, which are not specifically limited here.
[0082] In some embodiments, step 101 can be implemented as follows: performing first speech encoding on the first audio through multiple cascaded speech coding layers to obtain first audio features output by each of the speech coding layers; determining the first audio features output by the i-th speech coding layer in the multiple cascaded speech coding layers as the first paralinguistic features; and determining the first audio features output by the j-th speech coding layer in the multiple cascaded speech coding layers as the first linguistic features, where i and j are both positive integers, and the value of i is not greater than the value of j. When the values of i and j are different, the multiple cascaded speech coding layers can extract features at different scales and levels from the input first audio. The first paralinguistic features and the first linguistic features are independent to a certain extent. By combining the first paralinguistic features and the first linguistic features, a more comprehensive and accurate understanding of the audio information can be achieved, thereby improving the accuracy of the interactive processing.
[0083] It should be noted that, through a large amount of experimental data verification, in multiple cascaded speech coding layers, the higher the speech coding layer, the more the output features tend to represent linguistic information. Therefore, in the i-th speech coding layer corresponding to the first paralinguistic feature and the j-th speech coding layer corresponding to the first linguistic feature, where i and j are both positive integers, the value of i is not greater than the value of j. Assuming that there are N speech coding layers, where N is a positive integer greater than or equal to 2, then When i=j, the first paralinguistic feature extracted is the same as the first linguistic feature, and the second paralinguistic feature and the second linguistic feature are obtained by performing feature denoising of the corresponding linguistic dimension and feature denoising of the corresponding paralinguistic dimension respectively; and when The rounded down value of , that is, when N is an odd number, the value of i is When N is an even number, the value of i is When j=N, the first linguistic feature and the first paralinguistic feature extracted are used to train the model. When the loss function converges, the loss value corresponding to the loss function is smaller than the loss value obtained when the loss function converges when the model training is performed using features extracted from other speech coding layers. That is, when When j=N, the model generated by training the extracted first linguistic features and the first paralinguistic features has the best effect. In addition, the first audio features output by the first i speech coding layers can be fused, and the fused features can be used as the first paralinguistic features of the first audio. The model generated by such training has a better effect.
[0084] As an example, assuming that the speech encoder includes 6 cascaded speech coding layers, the first audio can be first speech encoded by each speech coding layer, and the first audio feature E output by each speech coding layer can be obtained. k, where E is used to characterize the audio features output by the speech coding layer, k is used to indicate the k-th speech coding layer, k is a positive integer, and the first audio feature output by the i-th speech coding layer can be used as the first paralinguistic feature, and the first audio feature output by the j-th speech coding layer can be used as the first linguistic feature. i and j are both positive integers, and the value of i is not greater than the value of j, thereby obtaining the first linguistic feature and the first paralinguistic feature corresponding to the first audio.
[0085] In addition, step 101 can also be implemented using only one speech coding layer, that is, using a single speech coding layer to perform first speech encoding on the first audio. In this case, the first linguistic feature and the first paralinguistic feature obtained are actually the same. Subsequently, feature denoising of the corresponding linguistic dimension and feature denoising of the corresponding paralinguistic dimension are performed respectively to obtain the second paralinguistic feature and the second linguistic feature, similar to the above case where i=j.
[0086] In step 102 , feature denoising corresponding to the paralinguistic dimension is performed on the first linguistic feature to obtain a second linguistic feature.
[0087] It should be noted that feature denoising corresponding to the paralinguistic dimension can be achieved through linear mapping, for example, determining the mapping matrix corresponding to the first linguistic feature, mapping the first linguistic feature through the mapping matrix, and obtaining the second linguistic feature; it can also be achieved through nonlinear mapping, for example, constructing a neural network model (multi-layer perceptron MLP, MLP consists of an input layer, a hidden layer, and an output layer, and there can be multiple hidden layers), and using the trained neural network model to obtain the corresponding second linguistic feature through forward propagation. No specific limitation is made here.
[0088] In some embodiments, step 102 can be implemented as follows: performing a dimensionality transformation on the first linguistic feature to obtain a first dimensionality transformation result; and performing a first linear transformation on the first dimensionality transformation result to obtain the second linguistic feature, wherein the first linear transformation is used to perform feature denoising corresponding to the paralinguistic dimension. In this manner, the dimensionality transformation process can recombin and filter the features, removing redundant or unimportant features while highlighting key features, thereby reducing the complexity of the feature space and improving the quality and expressiveness of the features. Furthermore, the first linear transformation can further perform feature denoising on the paralinguistic dimension based on the dimensionality transformation, thereby enhancing the expressiveness of the feature linguistic dimension and enabling the second linguistic feature to more accurately describe the semantic information of the audio.
[0089] As an example, suppose the first linguistic feature is E SL, SL (abbreviation of Speech Linguistic) is used to indicate that the feature is a linguistic feature, through E SL Every k adjacent elements i and k are both positive integers, Splicing into a single compact vector along the feature dimension, thereby achieving the dimensional change of the first linguistic feature, and obtaining the first dimension transformation result H, where the dimension of feature H is (n / k)×k, and n is feature E SL sequence length; then, the first dimension transformation result is linearly transformed for the first time through the first linear layer to obtain the corresponding linear transformation result Linear(H), wherein Linear(H) refers to the result after the linear transformation of feature H; then, the linear transformation result Linear(H) is activated by the ReLU activation function to obtain the corresponding activation result ReLU(Linear(H)), wherein ReLU(Linear(H)) refers to the result after the linear transformation result Linear(H) is activated; finally, the activation result ReLU(Linear(H)) is linearly transformed for the second time through the second linear layer to obtain the corresponding linear transformation result Linear(ReLU(Linear(H))), that is, the second linguistic feature is obtained, wherein Linear(ReLU(Linear(H))) refers to the result after the linear transformation of the activation result ReLU(Linear(H)).
[0090] In step 103 , feature denoising of the corresponding linguistic dimension is performed on the first paralinguistic feature to obtain a second paralinguistic feature.
[0091] It should be noted that feature denoising corresponding to the linguistic dimension can be achieved through linear mapping, for example, determining the mapping matrix corresponding to the first paralinguistic feature, mapping the first paralinguistic feature through the mapping matrix, and obtaining the second paralinguistic feature; it can also be achieved through nonlinear mapping, for example, constructing a neural network model (multi-layer perceptron MLP, MLP consists of an input layer, a hidden layer and an output layer, and there can be multiple hidden layers), and using the trained neural network model to obtain the second paralinguistic feature through forward propagation. No specific limitation is made here.
[0092] In some embodiments, step 103 can be implemented as follows: performing a second speech encoding on the first paralinguistic feature to obtain a first paralinguistic feature that has undergone feature denoising, wherein the second speech encoding is used to perform feature denoising on a corresponding linguistic dimension; pooling the first paralinguistic feature that has undergone feature denoising to a set length to obtain a pooling result that conforms to the set length; and performing a second linear transformation on the pooling result that conforms to the set length to obtain a second paralinguistic feature that conforms to a set vector dimension. Thus, by performing the second speech encoding on the first paralinguistic feature that has undergone feature denoising on the linguistic dimension to obtain the first paralinguistic feature that has undergone feature denoising, the paralinguistic information contained in the feature denoising is purer and the expression of the paralinguistic information is richer and more accurate. By fixing the feature length to the set length through the pooling operation, the length of the features of the sample can be kept consistent. Performing a second linear transformation on the pooling result that conforms to the set length to adjust the feature dimension to the set vector dimension can reduce feature redundancy and noise, allowing the model to focus more on the effective information in the features.
[0093] As an example, suppose the first paralinguistic feature is E SP , SP (Speech Paralinguistic) is used to indicate that the feature is a paralinguistic feature, and the E SP The second speech coding process can be performed to obtain the first paralinguistic feature Transformer (E SP ), where the Transformer module includes a multi-head attention layer and a regularization layer. Transformer (E SP ) refers to the feature E SP The result after encoding processing; then, the first linguistic feature Transformer (E SP ) Adaptively pool to a set length m, where m is a positive integer, and obtain the pooling result Pool(Transformer(E SP )), where Pool(Transformer(E SP )) refers to the feature Transformer (E SP ) is pooled; then, the pooled result that meets the set length is processed by the linear layer Pool(Transformer(E SP )) performs linear transformation processing, transforms the feature to the set vector dimension t, t is a positive integer, and obtains the linear transformation result Linear(Pool(Transformer(E SP))), that is, the second paralinguistic feature that meets the set vector dimension is obtained, where Linear(Pool(Transformer(E SP ))) refers to the pooling result Pool(Transformer(E SP ))The result after linear transformation.
[0094] In step 104 , a reply text is generated based on the second paralinguistic feature, the second linguistic feature, and the first prompt word for the first audio.
[0095] Here, the reply text is used to interact with the first audio according to the instruction of the first prompt word.
[0096] In some embodiments, step 104 can be implemented by encoding the first prompt word to obtain a first text feature; and generating a response text based on the second paralinguistic feature, the second linguistic feature, and the first prompt word for the first audio. In this way, by encoding the text features and combining them with the paralinguistic features, the meaning and context of the input text can be better understood, thereby generating a more accurate response text.
[0097] Here, text encoding can be achieved by processing the word segmenter and embedding layer separately.
[0098] As an example, the first prompt word is X T , where X represents the input content and T (abbreviation of Text) indicates that the input content is text. Therefore, using X T Indicates the first prompt word that needs to be input, and the first prompt word X is T Perform word segmentation processing to obtain the word segmentation result Tokenizer(X T ), where Tokenizer(X T ) refers to the first prompt word X T The result of word segmentation is then processed using the embedding layer to Tokenize the word segmentation result (X T ) is embedded and the corresponding embedding representation result is obtained. T )), that is, the first text feature is obtained, wherein, Embedlayer(Tokenizer(X T )) refers to the word segmentation result Tokenizer(X T )The result of performing embedding processing.
[0099] It should be noted that the generated reply text can be generated according to a preset reply strategy, or obtained by analyzing and processing the second paralinguistic features, the second linguistic features and the first prompt word through a large language model, and no specific limitation is made here.
[0100] As an example, suppose in an intelligent customer service scenario, the first prompt word for the first audio is "The mobile phone I bought broke down after only three days of use. You have to fix it for me." After text encoding, the first text feature E0 is obtained, and the first audio is processed to obtain the second linguistic feature E1 and the second paralinguistic feature E2; the first text feature E0, the second linguistic feature E1, and the second paralinguistic feature E2 are input into the model for processing, and the model (such as a large language model) outputs the corresponding reply text based on the input features.
[0101] In some embodiments, see Figure 5 , Figure 5 This is a flow chart of the interactive processing method provided in the embodiment of the present application. Figure 5 As shown, when executing Figure 4 Before step 101 shown, you can also perform Figure 5 Steps 105 to 108 shown will be combined with Figure 5 The steps shown are explained.
[0102] In step 105 , the initialized model is subjected to a first-stage training based on paralinguistic tasks and linguistic tasks to obtain a model that has undergone the first-stage training.
[0103] It should be noted that the first stage of training is to enable the model to generate responses based on audio input. During the first stage of training, only the parameters of the two adapters in the model are updated.
[0104] In some embodiments, the initialized model includes an initialized first adapter and an initialized second adapter. Step 105 can be implemented as follows: performing first speech encoding on the second audio to obtain a third linguistic feature and a third paralinguistic feature; performing feature denoising corresponding to the paralinguistic dimension on the third linguistic feature using the initialized first adapter to obtain a fourth linguistic feature, and performing feature denoising corresponding to the linguistic dimension on the third paralinguistic feature using the initialized second adapter to obtain a fourth paralinguistic feature; performing text encoding on the second prompt word to obtain a second text feature, wherein the second prompt word includes an instruction corresponding to the paralinguistic task and an instruction corresponding to the linguistic task; generating a first text prediction probability based on the fourth paralinguistic feature, the fourth linguistic feature, and the second text feature; generating a first text true probability based on the marked response text corresponding to the second audio and the second prompt word; and updating the initialized first adapter and the initialized second adapter in the initialized model based on the first text prediction probability and the first text true probability to obtain a model trained in the first stage.
[0105] In this way, by forward propagating the second audio and the second prompt word in the initialized model, the first text prediction probability corresponding to the second audio and the second prompt word is obtained. According to the first text prediction probability and the first text true probability, only the parameters of the first adapter and the second adapter in the initialized model are updated. There is no need to make large-scale adjustments to the entire model, which can reduce the complexity and computational cost of model training. By comparing the first text prediction probability and the first text true probability for training, the model can gradually learn a more accurate feature mapping method, thereby reducing the risk of overfitting of the model on the training data, and thus improving the generalization ability of the model.
[0106] It should be noted that the first text prediction probability is a probability matrix. Assuming that the marked reply text has 5 word positions, there are 10 candidate words in the first text prediction probability, and each word corresponds to a prediction probability, then the first text prediction probability is a 5×10 probability matrix; the first text true probability is also a probability matrix. The true probability of each word corresponding to the marked reply text in the probability matrix is 1, and the true probability of other words in the probability matrix is 0; the training data corresponding to the first stage training includes the second audio, the second prompt word, and the marked reply text corresponding to the second audio and the second prompt word.
[0107] As an example, taking the intelligent customer service scenario as an example, the second audio is first voice-encoded to obtain a third linguistic feature related to the audio semantics and a third paralinguistic feature that can reflect the customer's emotions, tone, etc.; then, the initialized first adapter performs feature denoising of the corresponding paralinguistic dimension on the third linguistic feature, and converts the third linguistic feature into a form more suitable for model understanding to obtain a fourth linguistic feature. At the same time, the initialized second adapter performs feature denoising of the corresponding linguistic dimension on the third paralinguistic feature, and converts the paralinguistic features reflecting the customer's emotions, tone, etc. into a format that can be processed by the model to obtain the fourth paralinguistic feature; then, the second prompt word of the second audio is text-encoded. The second prompt word contains instructions for paralinguistic tasks and linguistic tasks, and the instructions can be converted into The model converts the fourth paralinguistic feature, the fourth linguistic feature and the second text feature into a feature form that can be understood by the computer to obtain the second text feature; then, based on the fourth paralinguistic feature, the fourth linguistic feature and the second text feature, the model combines the customer's semantic information (fourth linguistic feature), emotional state (fourth paralinguistic feature) and task instructions (second text feature) for comprehensive analysis and processing, predicts possible reply texts, and generates a first text prediction probability; at the same time, based on the marked reply text of the second audio and the second prompt word, the first text true probability is generated, that is, the true probability of each word in the marked reply text is 1, and the obtained first text true probability is a probability matrix; finally, based on the first text prediction probability and the first text true probability, the initialized first adapter and second adapter are updated to obtain the model trained in the first stage.
[0108] In step 106 , the model trained in the first stage is subjected to second stage training based on the paralinguistic task to obtain a model trained in the second stage.
[0109] It should be noted that paralinguistic tasks refer to tasks that mainly rely on paralinguistic information (little or no linguistic information). The main purpose is to enable the model to more accurately understand the paralinguistic information of the audio by training the paralinguistic adapter (i.e., the second adapter).
[0110] In some embodiments, the model trained in the first stage includes a first adapter trained in the first stage and a second adapter trained in the first stage; the above-mentioned step 106 can be implemented by: performing first speech encoding on the third audio to obtain a fifth linguistic feature and a fifth paralinguistic feature; performing feature denoising of the corresponding paralinguistic dimension on the fifth linguistic feature using the first adapter trained in the first stage to obtain a sixth linguistic feature, and performing feature denoising of the corresponding linguistic dimension on the fifth paralinguistic feature using the second adapter trained in the first stage to obtain a sixth paralinguistic feature; performing text encoding on the third prompt word to obtain a third text feature, and performing feature denoising on the third audio The method comprises the steps of: encoding the text content corresponding to the frequency to obtain a fourth text feature, wherein the third prompt word includes an instruction corresponding to the paralinguistic task; performing random sampling processing on the fourth text feature, the sixth linguistic feature, and the missing feature to obtain a seventh linguistic feature; generating a second text prediction probability based on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature; generating a second text true probability based on the marked reply text corresponding to the third audio and the third prompt word; and updating the second adapter trained in the first stage in the model trained in the first stage based on the second text prediction probability and the second text true probability to obtain a model trained in the second stage.
[0111] In this way, based on the predicted probability of the second text and the true probability of the second text, only the second adapter trained in the first stage is updated. Since the second adapter is responsible for feature denoising in the linguistic dimension, adjusting its parameters can directly optimize the input conversion process from paralinguistic features to the model. Compared with large-scale adjustments to the entire model, it can effectively reduce the complexity and computational cost of model training, while effectively avoiding unnecessary interference with other already trained components. In addition, by randomly sampling the fourth text feature, the sixth linguistic feature, and the missing feature, the feature diversity is increased, avoiding the model's over-reliance on certain specific feature combinations, which helps to improve the model's generalization ability. Random sampling can simulate different feature distributions, enabling the model to perform better when facing various possible inputs.
[0112] It should be noted that the training data corresponding to the second stage of training includes the third audio, the third prompt word, and the marked response text corresponding to the third audio and the third prompt word; missing features represent missing or non-existent values. When the random sampling selects a missing feature, the seventh linguistic feature obtained by sampling is equivalent to an empty feature, and only the sixth paralinguistic feature and the third text feature are subsequently input into the model.
[0113] As an example, taking the intelligent customer service scenario as an example, the third audio is subjected to the first speech encoding to obtain the fifth linguistic feature related to the audio semantics and the fifth paralinguistic feature that can reflect the customer's emotions, tone, etc.; then, the first adapter trained in the first stage performs feature denoising of the corresponding paralinguistic dimension on the fifth linguistic feature, and converts the fifth linguistic feature into a form more suitable for model understanding to obtain the sixth linguistic feature. At the same time, the second adapter trained in the first stage performs feature denoising of the corresponding linguistic dimension on the fifth paralinguistic feature, and converts the paralinguistic features reflecting the customer's emotions, tone, etc. into a format that can be processed by the model to obtain the sixth paralinguistic feature; then, the third prompt word is text-encoded, and the third prompt word contains instructions for the paralinguistic task, and the instructions are converted into a special form that the computer can understand. Formula, obtain the third text feature, and perform text encoding on the text content corresponding to the third audio to obtain the fourth text feature; then, perform random sampling on the fourth text feature, the sixth linguistic feature and the missing feature to obtain the seventh linguistic feature, and perform comprehensive analysis based on the sixth paralinguistic feature, the seventh linguistic feature and the third text feature to predict the possible reply text and generate the second text prediction probability; at the same time, based on the marked reply text of the third audio and the third prompt word, generate the second text true probability, that is, the true probability of each word corresponding to the marked reply text is 1, and the obtained second text true probability is a probability matrix; finally, based on the second text prediction probability and the second text true probability, the second adapter trained in the first stage is updated to obtain the model trained in the second stage.
[0114] In some embodiments, generating the second text prediction probability corresponding to the third audio and the third prompt word based on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature can be achieved by: performing a sequence-randomized concatenation process on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature to obtain a first input feature; and then mapping the first input feature using the model to obtain the second text prediction probability. In this way, the sequence-randomized concatenation method breaks the inherent order of feature concatenation, enabling the model to learn a wider range of feature patterns when faced with diverse data, thereby improving the model's generalization ability. Furthermore, the randomized order requires the model to pay more attention to the intrinsic information of the features, thereby reducing the risk of overfitting and improving the model's stability and reliability.
[0115] As an example, assuming that the sixth paralinguistic feature is E1, the seventh linguistic feature is E2, and the third text feature is E3, the above features are concatenated based on sequential randomization, and the first input feature can be obtained as [E3, E1, E2] or [E3, E2, E1], and the first input feature is mapped through the model to obtain the second text prediction probability.
[0116] In step 107, the model trained in the second stage is subjected to a third stage training based on the linguistic task to obtain a model trained in the third stage.
[0117] It should be noted that a linguistic task refers to a task that mainly relies on linguistic information (with little or no reliance on paralinguistic information). The main purpose is to enable the model to accurately understand the linguistic information of the audio by training the linguistic adapter (i.e., the first adapter).
[0118] In some embodiments, the model trained in the second stage includes a first adapter trained in the first stage and a second adapter trained in the second stage; the above-mentioned step 107 can be implemented in the following manner: performing first speech encoding on the fourth audio to obtain an eighth linguistic feature and a seventh paralinguistic feature; performing feature denoising of the corresponding paralinguistic dimension on the eighth linguistic feature through the first adapter trained in the first stage to obtain a ninth linguistic feature, and performing feature denoising of the corresponding linguistic dimension on the seventh paralinguistic feature through the second adapter trained in the second stage to obtain an eighth paralinguistic feature; performing text encoding on the fourth prompt word to obtain a fifth text feature, and performing text encoding on the style description corresponding to the fourth audio to obtain a sixth text feature, wherein the fourth prompt word includes the corresponding The method comprises the steps of: encoding a fourth prompt word to obtain a fifth text feature, and encoding a style description corresponding to the fourth audio to obtain a sixth text feature, wherein the fourth prompt word includes an instruction corresponding to the linguistic task; performing random sampling processing on the sixth text feature, the eighth paralinguistic feature, and the missing feature to obtain a ninth paralinguistic feature; generating a third text prediction probability based on the ninth paralinguistic feature, the ninth linguistic feature, and the fifth text feature; generating a third text true probability based on the fourth audio and the marked reply text of the fourth prompt word; and updating the first adapter trained in the first stage in the second-stage trained model based on the third text prediction probability and the third text true probability to obtain a third-stage trained model.
[0119] In this way, based on the predicted probability of the third text and the true probability of the third text, only the first adapter trained in the first stage is updated. Since the first adapter is responsible for feature denoising in the paralinguistic dimension, adjusting its parameters can directly optimize the input conversion process from linguistic features to the speech model. Compared with large-scale adjustments to the entire model, it can effectively reduce the complexity and computational cost of model training, while effectively avoiding unnecessary interference with other already trained components. In addition, by randomly sampling the fourth text feature, the eighth paralinguistic feature, and the missing feature, the feature diversity is increased, avoiding the model's over-reliance on certain specific feature combinations, which helps to improve the model's generalization ability. Random sampling can simulate different feature distributions, enabling the model to perform better when faced with various possible inputs.
[0120] It should be noted that the training data corresponding to the third stage of training includes the fourth audio, the fourth prompt word, and the marked response text corresponding to the fourth audio and the fourth prompt word; missing features represent missing or non-existent values. When the random sampling selects a missing feature, the ninth paralinguistic feature obtained by sampling is equivalent to an empty feature. Only the ninth linguistic feature and the fifth text feature are subsequently input into the model.
[0121] As an example, taking the intelligent customer service scenario as an example, the fourth audio is subjected to the first speech encoding to obtain the eighth linguistic feature related to the audio semantics and the seventh paralinguistic feature that can reflect the customer's emotions, tone, etc.; then, the first adapter trained in the first stage performs feature denoising of the corresponding paralinguistic dimension on the eighth linguistic feature, and converts the eighth linguistic feature into a form more suitable for model understanding to obtain the ninth linguistic feature. At the same time, the second adapter trained in the second stage performs feature denoising of the corresponding linguistic dimension on the seventh paralinguistic feature, and converts the paralinguistic features reflecting the customer's emotions, tone, etc. into a format that can be processed by the model to obtain the eighth paralinguistic feature; then, the fourth prompt word is text-encoded, and the fourth prompt word contains instructions for the linguistic task, and the instructions are converted into a special form that the computer can understand. , obtain the fifth text feature, and perform text encoding on the style description corresponding to the fourth audio to obtain the sixth text feature; then, perform random sampling on the sixth text feature, the eighth paralinguistic feature, and the missing feature to obtain the ninth paralinguistic feature, and perform a comprehensive analysis based on the ninth linguistic feature, the ninth paralinguistic feature, and the fifth text feature to predict the possible reply text and generate a third text prediction probability; at the same time, based on the fourth audio and the marked reply text of the fourth prompt word, generate the third text true probability, that is, the true probability of each word corresponding to the marked reply text is 1, and the obtained third text true probability is a probability matrix; finally, based on the third text prediction probability and the third text true probability, the first adapter trained in the first stage is updated to obtain a model trained in the third stage.
[0122] In some embodiments, generating the third text prediction probability corresponding to the fourth audio and the fourth prompt word based on the ninth paralinguistic feature, the ninth linguistic feature, and the fourth text feature can be achieved by: performing a randomized sequential concatenation process on the ninth paralinguistic feature, the ninth linguistic feature, and the fifth text feature to obtain a second input feature; and performing a mapping process on the second input feature to obtain the third text prediction probability. This randomized concatenation breaks the inherent order of feature concatenation, enabling the model to learn a wider range of feature patterns when presented with diverse data, thereby improving the model's generalization ability. Furthermore, the randomized sequence forces the model to pay more attention to the intrinsic information of the features, thereby reducing the risk of overfitting and improving the model's stability and reliability.
[0123] It should be noted that the above-mentioned step of generating the third text prediction probability corresponding to the fourth audio and the fourth prompt word based on the ninth paralinguistic feature, the ninth linguistic feature, and the fourth text feature is similar to the above-mentioned step of generating the second text prediction probability corresponding to the third audio and the third prompt word based on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature. For details, please refer to the above-mentioned step of generating the second text prediction probability corresponding to the third audio and the third prompt word based on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature, which will not be repeated here.
[0124] In step 108, the model trained in the third stage is subjected to a fourth stage training based on the alignment task to obtain a model trained in the fourth stage.
[0125] Here, the model trained in the fourth stage is used to be called to perform feature denoising on the first linguistic feature and the first paralinguistic feature.
[0126] In some embodiments, the model trained in the third stage includes a first adapter trained in the third stage and a second adapter trained in the second stage. The above step 108 can be implemented by: performing first speech encoding on the fifth audio to obtain a tenth linguistic feature and a tenth paralinguistic feature; performing feature denoising of the corresponding paralinguistic dimension on the tenth linguistic feature using the first adapter trained in the third stage to obtain an eleventh linguistic feature; and performing feature denoising of the corresponding linguistic dimension on the tenth paralinguistic feature using the second adapter trained in the second stage to obtain the eleventh paralinguistic feature; performing text encoding on the fifth prompt word, Obtain a seventh text feature, and perform text encoding on the style description and text content corresponding to the fifth audio to obtain an eighth text feature, wherein the seventh prompt word includes instructions corresponding to the dual-information task; generate a fourth text prediction probability based on the eleventh paralinguistic feature, the eleventh linguistic feature, and the seventh text feature; generate a fourth text true probability based on the fifth audio and the seventh prompt word; and update the first adapter trained in the third stage and the second adapter trained in the second stage in the model trained in the third stage based on the fourth text prediction probability and the fourth text true probability to obtain a model trained in the fourth stage.
[0127] In this way, based on the eleventh paralinguistic feature, the eleventh linguistic feature and the seventh text feature, the fourth text prediction probability is generated, and based on the fifth audio and the seventh prompt word, the fourth text true probability is generated as a reference standard for evaluating the accuracy of the model prediction. By comparing the fourth text prediction probability and the fourth text true probability, the difference between the self-prediction and the actual situation can be found, and the parameters of the first adapter in the model trained in the third stage and the second adapter trained in the second stage are updated according to the difference, so that the model can predict the text content more accurately in subsequent processing, thereby improving the performance and generalization ability of the model.
[0128] It should be noted that the training data corresponding to the fourth stage of training includes the fifth audio, the fifth prompt word, and the marked reply text corresponding to the fifth audio and the fifth prompt word; the audio (including the second audio, the third audio, the fourth audio, and the fifth audio) used in the training process at different stages can be the same or different, and there is no specific limitation here; the prompt words and marked reply text in the training data at different stages of training are different.
[0129] As an example, taking the intelligent customer service scenario as an example, the fifth audio is first voice-encoded to obtain the tenth linguistic feature related to the audio semantics and the tenth paralinguistic feature that can reflect the customer's emotions, tone, etc.; then, the first adapter trained in the third stage performs feature denoising of the corresponding paralinguistic dimension on the tenth linguistic feature, and converts the tenth linguistic feature into a form more suitable for model understanding to obtain the eleventh linguistic feature. At the same time, the second adapter trained in the second stage performs feature denoising of the corresponding linguistic dimension on the tenth paralinguistic feature, and converts the paralinguistic features reflecting the customer's emotions, tone, etc. into a format that can be processed by the model to obtain the eleventh paralinguistic feature; then, the fifth prompt word is text-encoded, and the fifth prompt word contains instructions for the dual-information task, and the instructions are converted into a form that the computer can understand. , and obtain the seventh text feature; then, based on the eleventh paralinguistic feature, the eleventh linguistic feature and the seventh text feature, combined with the customer's semantic information (eleventh linguistic feature), emotional state (eleventh paralinguistic feature) and task instructions (seventh text feature), a comprehensive analysis and processing is performed to predict the possible reply text and generate the fourth text prediction probability; at the same time, according to the fifth audio and the seventh prompt word, the fourth text true probability is generated by the pre-trained behavior alignment model, that is, the true probability of each word is 1, and the obtained fourth text true probability is a probability matrix; finally, according to the fourth text prediction probability and the fourth text true probability, the first adapter trained in the third stage and the second adapter trained in the second stage are updated to obtain the model trained in the fourth stage.
[0130] In some embodiments, after executing step 108, preference alignment training may be performed on the model trained in the fourth stage to obtain a preference alignment trained model.
[0131] Preference alignment training is a machine learning method designed to adjust model outputs to align with human values, preferences, or ethical standards. Its core purpose is to use human feedback or preference data to guide the model to make more predictable decisions or generate more appropriate responses in complex scenarios. Human feedback can be used to build reward models and optimize strategies through reinforcement learning. Model parameters can also be adjusted directly through comparative data to improve efficiency and stability. Furthermore, adversarial examples can be generated to train the model to recognize and reject inputs that do not conform to preferences. These methods are not specifically defined here.
[0132] Thus, first, in the first phase, simultaneous training on paralinguistic and linguistic tasks allows the model to fully engage with the diverse information dimensions in the audio and extract a rich variety of features from it. Subsequently, building on the foundation of the first phase, the second phase focuses on training on linguistic tasks, further deepening the model's understanding and processing of speech semantics. This allows the model to learn more complex semantic structures, contextual relationships, and semantic reasoning rules, thereby improving the accuracy and reliability of linguistic-related tasks such as semantic analysis and information extraction. Through specialized training on linguistic tasks, the model can better grasp the expressive patterns and conventions of audio, thereby generating more natural, fluent, and accurate verbal responses or text output. Subsequently, the third phase focuses on training on paralinguistic tasks, enabling the model to more keenly perceive paralinguistic information such as emotion and tone in the audio. This allows the model to better understand the corresponding emotional state and intent in the audio when processing it, thereby generating more appropriate responses. Finally, training on the alignment task in the fourth phase ensures good consistency and coordination between different features when processing audio. Through training in this fourth phase, the model can achieve better balance and synergy between various tasks, further improving overall performance and stability. In summary, the phased training method allows the model to gradually learn and optimize different aspects of its capabilities, avoiding the confusion and inefficiency caused by processing too many complex tasks at one time.
[0133] The following describes an exemplary application of the embodiment of the present application in a practical application scenario, which describes the specific implementation process of the interactive processing method in a speech recognition scenario.
[0134] In the prior art, dialogue systems typically focus primarily on the content of conversations, often overlooking the importance of emotional factors in communication. To enhance user experience, emotion understanding has become a key research direction for next-generation dialogue systems. Dialogue systems with emotion understanding capabilities are crucial for achieving efficient and natural human-computer interaction. The key to developing emotional dialogue systems lies in understanding the conversation context and user expressions. In the prior art, dialogue systems with emotion understanding capabilities typically rely on text-based large language models (LLMs). By leveraging the text processing capabilities of LLMs, they achieve multi-turn understanding of human-computer interaction in multi-turn dialogue scenarios, thereby generating coherent responses. However, dialogue systems based on LLMs still suffer from the problem of misunderstanding user expressions. This is primarily because LLMs use text-based language as input and primarily consider the speaker's content, i.e., linguistic information, while largely ignoring the speaker's paralinguistic information, such as intonation and speaking speed. Paralinguistic information is crucial for understanding the speaker's emotions and intentions, and understanding this information helps dialogue systems generate more accurate and appropriate responses. Without understanding paralinguistic information, the dialogue system may misjudge the speaker's emotions or intentions, resulting in inappropriate responses.
[0135] As an example, Figure 1 As shown, for a dialogue system receiving a vibrant, cheerful female voice saying "I'm moving soon" in a high-pitched and rapid voice, if only the speaker's linguistic information (i.e., the content of the speech) is considered while ignoring paralinguistic information (i.e., the way it speaks), the dialogue system may misjudge the user's emotional state (for example, misclassifying excitement as anxiety) and thus give an inappropriate response: "Don't worry, moving is a hassle. You can ask some friends for help and it will go more smoothly." In contrast, considering both linguistic and paralinguistic information helps the dialogue system accurately understand the speaker's emotional state and generate a more appropriate response: "That's so exciting! How's the new place?" This highlights the importance of incorporating the speaker's voice as input in affective dialogue systems. The speaker's voice can be used to perceive not only the content of the speech but also the way it is spoken, leading to a more comprehensive understanding of the speaker's expression and the generation of more appropriate responses. Therefore, a key focus in the development of affective dialogue systems is to build a large multimodal model that integrates speech and language, namely a large speech-language model.
[0136] However, building a large speech and language model is a challenging task, with two main technical approaches. One feasible technical approach is to develop a speech-text based model that can natively process and understand speech and text. Although this approach is very effective, it requires large multimodal datasets, high computing resources, and advanced training techniques, which to some extent limits its widespread application. Another more feasible technical approach is to expand the large text-based language model to enable it to understand speech input while inheriting the large language model's conversational context understanding capabilities and knowledge reserves. This approach usually involves integrating a speech encoder with a large language model and connecting the two through an adapter module to build a large speech and language model. The technical approach of expanding the large text-based language model has attracted widespread attention due to its advantages such as being able to leverage the capabilities of existing pre-trained models and having controllable training costs.
[0137] However, although the speech language model built by extending the text-based large language model has shown potential in some tasks, it still has shortcomings when applied to emotional dialogue systems. Two major problems are:
[0138] (1) Insufficient understanding of paralinguistic information: Although the speech-language model uses the speaker's speech as input and shows effective understanding of the linguistic information in the speech, it fails to fully understand the paralinguistic information in the speech;
[0139] (2) Insufficient understanding of context: Compared with the large language model, the large speech language model obtained through extended training has a decreased ability to understand instructions and the context of multi-round dialogues.
[0140] On the one hand, existing technologies construct large speech-language models by expanding large language models. However, these technologies employ only a single speech feature extraction model and mapping layer (i.e., an adapter) to generate a sequence of embeddings representing both linguistic and paralinguistic information. Furthermore, no measures are taken to prevent the embeddings generated by the adapter from degenerating into task-specific vectors. Consequently, the trained large speech-language models suffer from insufficient understanding of paralinguistic information and insufficient understanding of context, making them unsuitable for use in emotional dialogue systems. On the other hand, existing technologies pursue the most thorough decoupling possible to achieve speech conversion. To this end, technical solutions are designed to minimize the correlation between the content, rhythm, and pitch representations extracted by different encoders.
[0141] However, in order to enable the large speech and language model to obtain relevant information from different adapters, this application explicitly ensures, through model structure design and training strategies, that the embedding vectors output by the paralinguistic adapter can represent paralinguistic information as comprehensively as possible, while also ensuring that the embedding vectors output by the linguistic adapter can represent linguistic information as comprehensively as possible. Furthermore, to ensure low information redundancy between the information represented by the embedding vectors output by the two heterogeneous adapters, this application does not pursue the goal of completely eliminating linguistic information from the embedding vectors output by the paralinguistic adapter, nor the opposite. Instead, it indirectly achieves decoupling between the two by implicitly reducing information redundancy (i.e., information duplication) between the paralinguistic and linguistic embedding vectors through downstream training task loss functions and training strategies (e.g., information shortcuts). For building a large model of speech and language, representation decoupling only focuses on the performance of the intermediate state. The interactive processing method provided by this application ultimately focuses on the performance of the whole (i.e., the large model uses information to complete the task). Since most tasks rely on both paralinguistic information and linguistic information to complete, but the proportion of the two types of information depends on different tasks is different, therefore, more emphasis is placed on the coordination of the two types of information rather than the degree of decoupling; in addition, perfect decoupling is not realistic, so the paralinguistic-linguistic information decoupling designed by this application will not achieve perfect decoupling. For example, the dimension of the paralinguistic embedding vector is fixed, but some paralinguistic information (such as changes in pronunciation) is time-dependent and is more likely to be captured by the time-dependent linguistic embedding vector. If the linguistic embedding vector is forced to not capture the paralinguistic information, then this part of the information is likely to be lost, which will affect the performance of some downstream tasks. Therefore, this application leaves freedom in choice and does not pursue perfect decoupling, but rather pursues reducing information redundancy.
[0142] During the implementation of this application, it was found that "information coupling" and "inappropriate training strategies" are the causes of the two problems of "insufficient understanding of paralinguistic information" and "insufficient contextual understanding". In the related art, a single adapter model is usually used to generate speech embedding vectors (Embeddings) that simultaneously represent linguistic information and paralinguistic information. When the speech embedding vector is projected into the input embedding space of the pre-trained large language model, since the space is obtained based on text corpus training, the linguistic information part of the speech embedding vector that is closer to the text will be retained first, while the paralinguistic information part of the speech embedding vector that is dissimilar to the text will be ignored; in addition, in the related art, the speech language large model is usually trained by instruction fine-tuning. Since the task types used in instruction fine-tuning are limited and there is a lack of measures to avoid the speech embedding vector generated by the adapter from degenerating into a task-specific vector, it may cause the speech language large model to overfit to the training task, thereby reducing the contextual understanding ability of the speech language large model.
[0143] Based on this, this application proposes an interactive processing method. In order to solve the problem of "insufficient understanding of paralinguistic information", this application separates linguistic information and paralinguistic information through two technical means: "heterogeneous adapter model architecture" and "weakly supervised training strategy", so that the large speech language model can perceive the two types of information separately through different mechanisms.
[0144] It should be noted that the heterogeneity of heterogeneous adapters in this application includes the following three meanings:
[0145] (1) The model structures of the two adapters are different;
[0146] (2) The dimensions of the embedding vectors input to the two adapters are different, and this dimension difference takes into account the characteristics of the information to be represented;
[0147] (3) The two embedding vectors are understood by the large language model in different ways, and this difference takes into account the understanding mechanism of the large language model.
[0148] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the structure of the speech language model provided in the embodiment of the present application. Figure 6 As shown, this application uses a heterogeneous adapter model architecture comprising two heterogeneous adapters, each outputting a paralinguistic embedding vector representing paralinguistic information and a linguistic embedding vector representing linguistic information. Due to the different inherent characteristics of the information to be represented, the dimensions of the two embedding vectors are designed differently. The paralinguistic embedding vector primarily represents paralinguistic information. Since each sentence in a conversation is short, typically less than 30 seconds, paralinguistic information can be considered relatively consistent throughout the entire sentence, possessing global characteristics and being insensitive to time. Therefore, a fixed number of paralinguistic embedding vectors generated by the paralinguistic adapter are used for representation. That is, the dimensionality of the paralinguistic embedding vectors is independent of the input speech length and is fixed. The linguistic embedding vector primarily represents linguistic information. Linguistic information is related to speech duration and also has relatively high requirements for time alignment. Therefore, the linguistic adapter generates a number of linguistic embedding vectors proportional to the actual speech duration. That is, the dimensionality of the linguistic embedding vectors is proportional to the input speech length.
[0149] Because the two embedding vectors have different dimensions, they can more easily represent corresponding aspects of information while having difficulty representing other aspects. For example, an embedding vector with a fixed dimension has difficulty representing linguistic information related to duration, while an embedding vector with a dimension proportional to the length of the speech can more easily represent linguistic information related to duration. This difference allows the two heterogeneous adapters to prioritize representing corresponding information from different perspectives, resulting in different functionalities.
[0150] It should be noted that although different embedding vector dimensions help different adapters represent different aspects of information, structural heterogeneity of adapters alone cannot guarantee the following three points:
[0151] (1) The embedding vector output by the paralinguistic adapter can represent paralinguistic information as comprehensively as possible;
[0152] (2) The embedding vector output by the linguistic adapter can represent the linguistic information as comprehensively as possible;
[0153] (3) The information redundancy between the information represented by the embedding vectors output by the two heterogeneous adapters is low.
[0154] To achieve the above three points, in addition to the structural heterogeneity of the adapter, this application also adopts a "weakly supervised training strategy". The core of "weakly supervised training" is "training task classification" and "equivalence replacement regularization (ERR)". Training task classification refers to dividing training tasks into three categories: "paralinguistic tasks", "linguistic tasks" and "dual-information tasks" based on the information that the tasks mainly rely on. Equivalence replacement regularization guides the speech language large model to generate responses based on instructions and related embedding vectors containing corresponding information. The design of equivalence replacement regularization is mainly based on one main principle, that is, for linguistic information tasks, that is, tasks that mainly rely on linguistic information, regardless of the existence of paralinguistic embedding vectors or how the source of paralinguistic embedding vectors changes, both instructions and linguistic embedding vectors should be sufficient to enable the speech language large model to maintain stable performance on linguistic information tasks. Therefore, when using equivalent substitution regularization to train the linguistic adapter to handle linguistic tasks, the paralinguistic adapter keeps its parameters frozen, and the linguistic embedding vectors are randomly combined with paralinguistic embedding vectors from text, speech, or missing data for training. Similarly, the same principles apply to training the paralinguistic adapter to handle paralinguistic tasks and will not be repeated here.
[0155] In some embodiments, different training tasks are designed according to the above principles to train different adapters to extract embedding vectors representing corresponding information. Paralinguistic tasks are used to train paralinguistic adapters to learn and generate embedding vectors representing paralinguistic information. Paralinguistic tasks are not completely independent of linguistic information. If the embedding vector representation contains some linguistic information, it can reduce the loss function to a certain extent. Therefore, during training, the paralinguistic embedding vectors are encouraged to also represent a portion of the linguistic vectors. To ensure that paralinguistic task training focuses on features representing paralinguistic information, vectors representing linguistic information from other sources are provided during training, so that the loss function does not need to obtain the required linguistic information from the paralinguistic embedding vectors. Linguistic tasks are used to train linguistic adapters to learn and generate embedding vectors representing linguistic information. Furthermore, the equivalent substitution regularization of the weakly supervised training strategy uses information shortcuts to prevent the embedding vectors output by the adapter from representing unexpected information, thereby reducing information redundancy. For example, during training, when training a paralinguistic adapter using a paralinguistic task, vectors representing linguistic information from other sources are provided to reduce the need and opportunity for the paralinguistic embedding vector to represent some linguistic information. Similarly, when training a linguistic adapter using a linguistic task, vectors representing paralinguistic information from other sources are provided to reduce the need and opportunity for the linguistic embedding vector to represent some linguistic information. This allows the embedding vectors output by the two adapters to each focus on representing corresponding information and reduces information redundancy between them, thereby jointly providing comprehensive features of the speech information. Here, the paralinguistic task is not completely independent of linguistic information. If the embedding vector representation contains some linguistic information, it can reduce the loss function to a certain extent. Therefore, during training, the paralinguistic embedding vector is encouraged to also represent some linguistic information. Furthermore, to ensure that the paralinguistic task training focuses on the characteristics of paralinguistic information, vectors representing linguistic information from other sources are provided during training, so that the loss function does not need to obtain the required linguistic information from the paralinguistic embedding vector. The implementation of the linguistic task is similar to that of the paralinguistic task and will not be further described here.
[0156] It should be noted that this application does not implement any punitive rules. This means that paralinguistic embeddings do not represent linguistic information at all, nor does it require linguistic embeddings to represent paralinguistic information at all. Instead, the model is adaptively adjusted during training by optimizing the training objective. This is primarily due to the fact that, when designing embedding dimensions, most paralinguistic information is insensitive to time. However, some paralinguistic information, such as the emphasis of each word, is sensitive to time and correlated with speech duration. Therefore, this paralinguistic information is more likely to be represented by appropriately structured linguistic embeddings. Furthermore, the interactive processing scheme provided by this application allows linguistic embeddings to represent some additional paralinguistic information in addition to linguistic information.
[0157] In some embodiments, paralinguistic and linguistic embeddings influence the behavior of the speech language model through an equivalent replacement text content embedding mechanism and soft prompts. For linguistic embeddings, since the linguistic information in speech and the linguistic information in text are highly aligned, the linguistic embeddings can directly replace the corresponding text content embeddings in the original large language model. For paralinguistic embeddings, since the paralinguistic information in speech lacks a corresponding counterpart in the original large language model, this application uses a soft prompt mechanism to guide the speech language model to perceive the paralinguistic information in the paralinguistic embeddings.
[0158] In some embodiments, to address the problem of "lack of context understanding", this application introduces three randomization schemes including position randomization, combination randomization, and order randomization by optimizing the training strategy. By preventing the embedding vector output by the adapter from degenerating into a task-specific vector, the speech language large model retains the context understanding ability of the original large language model. For position randomization, it is achieved by using multi-round dialogue data and dynamic context length during the training process, such as Figure 6 As shown, the context embedding vector is placed at the beginning of the speech language model input, before the paralinguistic embedding vector and the linguistic embedding vector. Therefore, the position of the paralinguistic embedding vector and the linguistic embedding vector varies significantly with the length of the context; for combined randomization, see Figure 7 , Figure 7 is a schematic diagram of different adapter training provided in the embodiment of the present application, such as Figure 7As shown in , combinatorial randomization refers to the possible combinations of paralinguistic embedding vectors and linguistic embedding vectors (which can be any one of the linguistic embedding vectors from speech, linguistic embedding vectors from text, and space) generated by probability sampling during the equivalent replacement regularization sampling process; and the possible combinations of linguistic embedding vectors and paralinguistic embedding vectors (which can be any one of the paralinguistic embedding vectors from speech, paralinguistic embedding vectors from text, and space) generated by probability sampling; for sequential randomization, see Figure 8 , Figure 8 This is a schematic diagram of the random order provided in the embodiment of the present application, such as Figure 8 As shown in the figure, order randomization means that the order relationship between the paralinguistic embedding vector and the linguistic embedding vector is random during training. The paralinguistic embedding vector is composed of <style>< / style> The corresponding embedding vector of the tag is wrapped, and the linguistic embedding vector is composed of <content>< / content> The embedding vectors corresponding to the tags are wrapped. Sequence randomization enables the speech and language model to distinguish between the two based on the differences in tags. Thus, through the combined effect of position randomization, combination randomization, and sequence randomization, randomization is introduced during the training of the speech and language model to avoid the emergence of fixed patterns. This prevents the speech and language model from overfitting to the limited training tasks and resulting in task-specific vectors, which in turn prevents a reduction in contextual understanding.
[0159] In some embodiments, as Figure 6 As shown, the speech language model in this application mainly includes four modules: speech encoder, paralinguistic adapter, linguistic adapter and large language model (tokenizer and embedding layer are considered as part of the large language model). The input content of the speech language model includes text prompt X T and speech signal X S Two parts, where X represents the input content, T (abbreviation of Text) indicates that the input content is text, and S (abbreviation of Sound) indicates that the input text is voice, so using X T Indicates a text prompt that needs to be entered, using X S Indicates the speech signal that needs to be input. The specific mathematical form of the feature below is a vector, so the two represent the same meaning below. Text embedding vector E T ∈R n×d, where R is a real number, meaning that each element in the embedding vector H is a real number, T is used to refer to the feature as a text embedding vector, n is the length of the text embedding vector sequence, and d is the feature dimension of the text embedding vector, which is generated by the word segmenter and embedding layer of the large language model and can be calculated using formula (1).
[0160] E T =Embedlayer(Tokenizer(X T )) (1)
[0161] Among them, Embedlayer refers to the embedding layer of the large language model, and Tokenizer refers to the word segmenter of the large language model.
[0162] This application obtains paralinguistic speech embedding vectors from speech and linguistic speech embedding vectors Where R is a real number, meaning that each element in the embedding vector H is a real number. SP (SpeechParalinguistic) is used to indicate that the feature is a paralinguistic feature, and SL (abbreviation of Speech Linguistic) is used to indicate that the feature is a linguistic feature. n1 and n2 are the lengths of the paralinguistic speech embedding vector sequence and the linguistic speech embedding vector sequence, respectively. d1 and d2 are the feature dimensions of the paralinguistic speech embedding vector and the linguistic speech embedding vector, respectively. Both are generated by the i-th layer and the j-th layer of the speech encoder, respectively. i and j are both positive integers, and the value of i is not greater than the value of j. They can be calculated according to formulas (2) and (3), respectively.
[0163] E SP =SpeechEncoder_i(X S ) (2)
[0164] E SL =SpeechEncoder_j(X S ) (3)
[0165] Among them, SpeechEncoder_i refers to the speech encoding layer of the i-th layer, SpeechEncoder_j refers to the speech encoding layer of the j-th layer, E SP The input of the paralinguistic adapter should retain relatively complete paralinguistic information. SP (SpeechParalinguistic) is used to indicate that the feature is a paralinguistic feature; E SLThe input of the linguistic adapter should retain relatively complete linguistic information. SL (abbreviation of Speech Linguistic) is used to indicate that the feature is a linguistic feature. Considering that most language encoders, such as Whisper-V3 Encoder, have higher-level output features that tend to represent linguistic information, we extract E from the middle i-th layer. SP , where SP (Speech Paralinguistic) is used to indicate that the feature is a paralinguistic feature. That is, usually the number of layers i <= j, i and j are both positive integers. E SP It can come not only from the i-th layer, but also from the feature fusion results of multiple layers of the speech encoder.
[0166] The two speech embedding vectors E obtained SP and E SL , where SP (Speech Paralinguistic) is used to indicate that the feature is a paralinguistic feature, and SL (abbreviation of Speech Linguistic) is used to indicate that the feature is a linguistic feature. After passing through the paralinguistic adapter A and the linguistic adapter C respectively, the corresponding paralinguistic embedding vector can be calculated according to formula (4) and formula (5): and linguistic embedding vectors Where R is a real number, which means that every element in the embedding vector H is a real number, A is used to indicate that the embedding vector is a paralinguistic embedding vector, C is used to indicate that the embedding vector is a linguistic embedding vector, n3 and n4 are the lengths of the paralinguistic embedding vector sequence and the linguistic embedding vector sequence, respectively, and d is the feature dimension of the embedding vector.
[0167] E A =ParalingusticAdapter(E SP ) (4)
[0168] E C =LingusticAdapter(E SL ) (5)
[0169] Among them, ParalingusticAdapter refers to the paralinguistic adapter, and LingusticAdapter refers to the linguistic adapter.
[0170] Finally, the obtained vectors are concatenated along the sequence dimension and input into the large language model to generate the response R, which can be calculated according to formula (6).
[0171] [E T ,E A ,E C ]→LLM→R (6)
[0172] Among them, E T Refers to the embedding vector corresponding to the prompt text, E A refers to the paralinguistic embedding vector processed by the paralinguistic adapter A, E C It refers to the linguistic embedding vector processed by the linguistic adapter C.
[0173] In some embodiments, the paralinguistic adapter is designed to capture paralinguistic information in speech based on the paralinguistic speech embedding vector E SP Generate paralinguistic embedding vector E A , where SP (Speech Paralinguistic) is used to indicate that the embedding vector is a paralinguistic embedding vector, E A Refers to the paralinguistic embedding vector processed by the paralinguistic adapter A. This application can use a compact transformer module to implement the paralinguistic adapter, which includes multiple transformer layers with multi-head self-attention and random dropout. The processed sequence is adaptively pooled in the sequence dimension to a preset fixed length n, where n is a positive integer; then, a linear layer is used to transform the pooled embedding vector to the feature dimension of the input space of the large language model to generate the paralinguistic embedding vector, which can be calculated according to formula (7).
[0174] E A =Linear(Pool(Transformer(E SP ))) (7)
[0175] Among them, Linear refers to the linear layer, Pool refers to the pooling layer, and Transformer refers to the conversion layer.
[0176] In some embodiments, for a linguistic adapter, the linguistic adapter is designed to capture linguistic information in speech, based on a linguistic speech embedding vector E SL Generate linguistic embedding vector E C , where SL (abbreviation for SpeechLinguistic) is used to indicate that the embedding vector is a linguistic embedding vector, E C Refers to the linguistic embedding vector processed by the linguistic adapter C. Given the redundancy of the speech embedding vector relative to the text embedding vector, the linguistic speech embedding vector E SL This problem is solved by converting it into a more compact vector representation. Specifically, by embedding the linguistic speech into a vector E SL Every k adjacent vector elements ESL i, i and k are all positive integers, Concatenate along the feature dimension into a single compact vector H, which can be E SL Converted to a compact embedding vector H∈R (n / k)×(d·k) , where R is a real number, meaning that each element in the embedding vector H is a real number, and n is the linguistic speech embedding vector E SL The length of the embedding vector sequence, d is the linguistic speech embedding vector E SL The feature dimension of , then, the compact embedding vector H is processed by two linear layers with an intermediate ReLU activation function and a hidden layer dimension of d0 to generate a linguistic embedding vector, which can be calculated according to formula (8).
[0177] E C =Linear(ReLU(Linear(H))) (8)
[0178] Among them, Linear refers to the linear layer and ReLU is the activation layer.
[0179] It should be noted that during the training process, in order to maintain the capabilities of the pre-trained model, the text-to-speech encoder and large language model parameters remain frozen, and only the parameters of the two adapters are updated.
[0180] In some embodiments, the linguistic tasks in this application refer to tasks that rely primarily on linguistic information (with little or no reliance on paralinguistic information). Such tasks are intended to enable the speech language model to better understand the linguistic information. This application uses automatic speech recognition (ASR) tasks as linguistic tasks for training. In automatic speech recognition tasks, the speech language model is instructed to generate a transcription of the input speech. For example, an instruction example may be: "Please write down what you heard word for word." This application generates diverse instructions by rephrasing and constructing training data.
[0181] In some embodiments, the paralinguistic tasks in this application refer to tasks that rely primarily on paralinguistic information (little or no linguistic information). Such tasks are intended to enable large speech language models to understand paralinguistic information. This application uses various speech attribute classification tasks, including gender, pitch, speed, energy, and emotion recognition, as paralinguistic tasks for training. In order to have a deterministic response, this application converts the task into multiple options. For example, the instruction for gender classification can be: "Identify the gender of the speaker. Select from the following options: female or male." This application generates diverse instructions by re-expressing and randomizing the order of options to construct training data.
[0182] In some embodiments, the dual-information task in this application refers to a task that relies on both paralinguistic information and linguistic information. This type of task is designed to enable the speech language model to adaptively utilize both types of information. This application uses the style-aware behavior alignment task for training. See Figure 10 , Figure 10 This is a schematic diagram of the training of style-aware behavior alignment provided by an embodiment of the present application, such as Figure 10 As shown in Figure 2, the style-aware behavior alignment task involves aligning the responses of a large speech language model with the responses of the large language model it employs given equivalent input (typically a speech input and its stylized transcription). For example, a piece of speech can be roughly equivalent to a description of the speaking style (style) plus the content (content): <style>一个充满活力和欢乐的尖细女声,语速很快。< / style> <content> I'm moving< / content> ", the response generated by the large language model based on this can be used as a supervision signal for the large speech language model.
[0183] In some embodiments, for subsequent preference learning, the large language model will generate two replies, one of which is generated based on the speaking style description and content, denoted as R1, for example, " <style>一个充满活力和欢乐的尖细女声,语速很快。< / style> <content> I'm moving< / content> "; the other sentence is generated based on pure content (i.e., without speaking style description), recorded as R2, for example, " <content> I'm moving< / content> ”. Pairs of (R1, R2) will be used for preference alignment training. For emotional dialogue, R1 is a better response than R2 considering the speaking style description.
[0184] It should be noted that for dual-information tasks, this application utilizes multi-turn conversation data. When training a large speech and language model, it simultaneously leverages conversation context, linguistics, and paralinguistics. Multi-turn conversations can be supervised by responses from different turns, constructing multiple samples with varying historical conversational turns. These samples have varying conversational context lengths due to the varying number of historical conversational turns, and this diversity is used for training. For example, given a three-turn conversation sample, denoted as [A1: Person A's speech text 1; B1: Person B's speech text 1; A2: Person A's speech text 2; B2: Person B's speech text 2; A3: Person A's speech text 3; B3: Person B's speech text 3], A1 represents Person A in the first turn, and Person A's speech text 1 represents the speaking style and content of Person A in the first turn. Similar examples exist for Person B1, A2, and so on, and are not further detailed. At the same time, the following samples can be multiplied: The speech and language model plays the role of A, the dialogue history is [A1: A's speech text 1; B1: B's speech text 1;], and the prediction target is "A's speech text 2"; the dialogue history is [A1: A's speech text 1; B1: B's speech text 1; A2: A's speech text 2; B2: B's speech text 2;], and the prediction target is "A's speech text 3". The speech and language model plays the role of B, the dialogue history is [A1: A's speech text 1;], and the prediction target is "B's speech text 1"; the dialogue history is [A1: A's speech text 1; B1: B's speech text 1; A2: A's speech text 2;], and the prediction target is "B's speech text 2"; the dialogue history is [A1: A's speech text 1; B1: B's speech text 1; A2: A's speech text 2; B2: B's speech text 2; A3: A's speech text 3;], and the prediction target is "B's speech text 3". Based on the above-mentioned multiplied samples, the speech language model's ability to communicate in different rounds can be trained. The vector can also be fully randomly embedded in the prompt word position through diverse conversation context lengths, thereby reducing the risk of the above-mentioned embedding vector degenerating into a task-specific vector.
[0185] In some embodiments, the weakly supervised training strategy proposed in this application mainly includes three stages, each of which uses corresponding training tasks, as follows:
[0186] (1) The first training phase. The focus of the first training phase is to enable the speech language model to generate reply text based on speech input. Therefore, the first training phase uses a combination of linguistic tasks and paralinguistic tasks to train the speech language model through instruction fine-tuning. It should be noted that during the training process, only the parameters of the two adapters of the speech language model are updated, and the speech encoder and the large language model keep the pre-trained parameters frozen. Although the speech language model trained in the first training phase has the ability to reply, its understanding of the context is limited, and the two types of information are still coupled together. Therefore, further training is required in the later stages.
[0187] (2) The second training phase. The focus of the second training phase is to ensure that the embedding vectors generated by the two adapters meet the following three requirements: the embedding vector output by the paralinguistic adapter can represent the paralinguistic information as comprehensively as possible; the embedding vector output by the linguistic adapter can represent the linguistic information as comprehensively as possible; and the information redundancy between the information represented by the embedding vectors output by the two heterogeneous adapters is low. Figure 9 , Figure 9 Schematic diagram of different adapter training after randomized combination provided in the embodiment of the present application, such as Figure 7 (a) and Figure 9 As shown in (a) in the figure, when training the linguistic adapter to handle linguistic tasks, the paralinguistic adapter remains frozen, and the linguistic embedding vector is randomly combined with the paralinguistic embedding vector from the speech style description text (Speech Caption), the paralinguistic embedding vector of the speech, or the missing (None) paralinguistic embedding vector, where the probabilities can be the same or different. Since the linguistic task mainly relies on linguistic information, the source of the paralinguistic information, that is, whether it comes from the speech style description text, speech, or missing, should not significantly affect the results. Figure 7 (b) and Figure 9 As shown in (b) in the figure, when training the paralinguistic adapter to handle the paralinguistic task, the linguistic adapter remains frozen, and the paralinguistic embedding vector is randomly combined with one of the linguistic embedding vectors from the speech content, the linguistic embedding vector from the speech, or the missing (None) linguistic embedding vector, where the probabilities can be equal or different. Since the paralinguistic task mainly relies on paralinguistic information, the source of the linguistic information, that is, whether it comes from the speech content description text, speech, or missing, should not significantly affect the results.
[0188] It should be noted that the equivalent substitution regularization adopted in this application does not strictly train the two adapters independently. For the training of paralinguistic adapters, this application adopts paralinguistic tasks. Linguistic information can also assist paralinguistic tasks, and linguistic embedding vectors can also represent part of the paralinguistic information. For example, speech content may be helpful for emotion recognition. When training a paralinguistic adapter, this application provides linguistic embedding vectors probabilistically to encourage the paralinguistic adapter to focus on capturing paralinguistic information without having to capture linguistic information components that are beneficial to completing paralinguistic tasks. For the training of linguistic adapters, this application adopts linguistic tasks. Paralinguistic information can also assist linguistic tasks, and paralinguistic embedding vectors can also represent part of the linguistic features. When training a linguistic adapter, this application provides paralinguistic embedding vectors probabilistically to ensure that it focuses on capturing linguistic information without having to take into account capturing paralinguistic information components that are beneficial to completing linguistic tasks.
[0189] (3) The third training phase. The third training phase focuses on fine-tuning the overall speech and language model to ensure that it can adaptively use both linguistic and paralinguistic information in context while avoiding generating task-specific vectors. This phase mainly includes two steps: style-aware behavior alignment training and preference alignment training.
[0190] Continue to see Figure 10 ,like Figure 10 As shown, dual-information task training data is used, and some paralinguistic task and linguistic task training data are added as auxiliary, and the large speech language model is trained with instruction fine-tuning; among them, the dual-information task supervision signal uses R1, where R1 refers to the response generated based on the speaking style description and content.
[0191] Subsequently, the dual-information task training data is used to train the speech language model through the preference alignment algorithm Direct Preference Optimization (DPO), so that the probability of the speech language model outputting a better response R1 is higher than R2, and the final speech language model is obtained.
[0192] The following continues to describe the exemplary structure of the interactive processing device 555 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 3 As shown, the software modules stored in the interaction processing device 555 of the memory 550 may include: an encoding module 5551 , a mapping module 5552 and a generating module 5553 .
[0193] The encoding module 5551 is used to perform a first speech encoding on the first audio to obtain a first linguistic feature and a first paralinguistic feature; the mapping module 5552 is used to perform feature denoising of the corresponding paralinguistic dimension on the first linguistic feature to obtain a second linguistic feature, and perform feature denoising of the corresponding linguistic dimension on the first paralinguistic feature to obtain a second paralinguistic feature; the generation module 5553 is used to generate a reply text based on the second paralinguistic feature, the second linguistic feature, and a first prompt word for the first audio, wherein the reply text is used to interact with the first audio according to the instruction of the first prompt word.
[0194] In some embodiments, the encoding module 5551 is further used to perform first speech encoding on the first audio through multiple cascaded speech coding layers to obtain a first audio feature output by each of the speech coding layers; determine the first audio feature output by the i-th speech coding layer in the multiple cascaded speech coding layers as the first paralinguistic feature; and determine the first audio feature output by the j-th speech coding layer in the multiple cascaded speech coding layers as the first linguistic feature, where i and j are both positive integers, and the value of i is not greater than the value of j.
[0195] In some embodiments, the mapping module 5552 is further configured to perform a dimensionality transformation on the first linguistic feature to obtain a first dimensionality transformation result; and perform a first linear transformation on the first dimensionality transformation result for paralinguistic dimension denoising to obtain the second linguistic feature, wherein the first linear transformation is used to perform feature denoising corresponding to the paralinguistic dimension.
[0196] In some embodiments, the mapping module 5552 is further configured to perform a second speech encoding on the first paralinguistic feature to obtain a first paralinguistic feature that has undergone feature denoising, wherein the second speech encoding is used to perform feature denoising of a corresponding linguistic dimension; pool the first paralinguistic feature that has undergone feature denoising to a set length to obtain a pooling result that meets the set length; and perform a second linear transformation on the pooling result that meets the set length corresponding to a set vector dimension to obtain a second paralinguistic feature that meets the set vector dimension.
[0197] In some embodiments, the above-mentioned interaction processing device 555 also includes a training module 5554.
[0198] The training module 5554 is configured to perform a first-stage training based on a paralinguistic task and a linguistic task on the initialized model to obtain a first-stage trained model; perform a second-stage training based on the paralinguistic task on the first-stage trained model to obtain a second-stage trained model; perform a third-stage training based on the linguistic task on the second-stage trained model to obtain a third-stage trained model; and perform a fourth-stage training based on an alignment task on the third-stage trained model to obtain a fourth-stage trained model. The fourth-stage trained model is configured to be called to perform feature denoising on the first linguistic feature and the first paralinguistic feature.
[0199] In some embodiments, the initialized model includes an initialized first adapter and an initialized second adapter; the training module 5554 is further configured to perform a first speech encoding on the second audio to obtain a third linguistic feature and a third paralinguistic feature; perform feature denoising of the corresponding paralinguistic dimension on the third linguistic feature using the initialized first adapter to obtain a fourth linguistic feature, and perform feature denoising of the corresponding linguistic dimension on the third paralinguistic feature using the initialized second adapter to obtain a fourth paralinguistic feature; perform text encoding on the second prompt word to obtain a second text feature, wherein the second prompt word includes an instruction corresponding to the paralinguistic task and an instruction corresponding to the linguistic task; generate a first text prediction probability based on the fourth paralinguistic feature, the fourth linguistic feature, and the second text feature; generate a first text true probability based on the marked reply text corresponding to the second audio and the second prompt word; and update the initialized first adapter and the initialized second adapter in the initialized model based on the first text prediction probability and the first text true probability to obtain a model trained in the first stage.
[0200] In some embodiments, the model trained in the first stage includes a first adapter trained in the first stage and a second adapter trained in the first stage; the training module 5554 is further used to perform a first speech encoding on the third audio to obtain a fifth linguistic feature and a fifth paralinguistic feature; perform feature denoising of the corresponding paralinguistic dimension on the fifth linguistic feature through the first adapter trained in the first stage to obtain a sixth linguistic feature, and perform feature denoising of the corresponding linguistic dimension on the fifth paralinguistic feature through the second adapter trained in the first stage to obtain a sixth paralinguistic feature; perform text encoding on the third prompt word to obtain a third text feature, and perform feature denoising of the corresponding linguistic dimension on the third audio. The text content is encoded to obtain a fourth text feature, wherein the third prompt word includes an instruction corresponding to the paralinguistic task; the fourth text feature, the sixth linguistic feature, and the missing feature are randomly sampled to obtain a seventh linguistic feature; a second text prediction probability is generated based on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature; a second text true probability is generated based on a marked response text corresponding to the third audio and the third prompt word; and a second adapter trained in the first stage in the model trained in the first stage is updated based on the second text prediction probability and the second text true probability to obtain a model trained in the second stage.
[0201] In some embodiments, the training module 5554 is further configured to perform sequential randomization-based concatenation processing on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature to obtain a first input feature; and perform mapping processing on the first input feature to obtain a second text prediction probability.
[0202] In some embodiments, the model trained in the second stage includes a first adapter trained in the first stage and a second adapter trained in the second stage; the training module 5554 is further used to perform a first speech encoding on the fourth audio to obtain an eighth linguistic feature and a seventh paralinguistic feature; perform feature denoising of the corresponding paralinguistic dimension on the eighth linguistic feature through the first adapter trained in the first stage to obtain a ninth linguistic feature, and perform feature denoising of the corresponding linguistic dimension on the seventh paralinguistic feature through the second adapter trained in the second stage to obtain an eighth paralinguistic feature; perform text encoding on the fourth prompt word to obtain a fifth text feature, and perform feature denoising of the corresponding linguistic dimension on the fourth audio. The style description is text-encoded to obtain a sixth text feature, wherein the fourth prompt word includes an instruction corresponding to the linguistic task; the sixth text feature, the eighth paralinguistic feature, and the missing feature are randomly sampled to obtain a ninth paralinguistic feature; a third text prediction probability is generated based on the ninth paralinguistic feature, the ninth linguistic feature, and the fifth text feature; a third text true probability is generated based on the fourth audio and the marked response text of the fourth prompt word; and a first adapter trained in the first stage in the model trained in the second stage is updated based on the third text prediction probability and the third text true probability to obtain a model trained in the third stage.
[0203] In some embodiments, the training module 5554 is further configured to perform random sequential concatenation processing on the ninth paralinguistic feature, the ninth linguistic feature, and the fifth text feature to obtain a second input feature; and perform mapping processing on the second input feature to obtain a third text prediction probability.
[0204] It should be noted that the description of the device in the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, so it will not be repeated here. Figure 4 ,or Figure 5 The present invention should be understood by referring to the description of any one of the accompanying drawings.
[0205] The present invention provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the interaction processing method described in the present invention.
[0206] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the interactive processing method provided in the embodiment of the present application, for example, Figure 4 ,or Figure 5 The interactive processing method shown.
[0207] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or compact disc read-only memory (CD-ROM); or various devices including one or any combination of the above memories.
[0208] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0209] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0210] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0211] In summary, by encoding the audio into a first linguistic feature and a first paralinguistic feature through the embodiment of the present application, the features of the two branches can be decoupled from the audio; then, the first linguistic feature is subjected to feature denoising of the corresponding paralinguistic dimension to obtain the second linguistic feature, and at the same time, the first paralinguistic feature is subjected to feature denoising of the corresponding linguistic dimension to obtain the second paralinguistic feature. By setting different vector dimensions for feature denoising processing, the paralinguistic information and the linguistic information can be effectively decoupled, and the information redundancy between the two features can be reduced; finally, the second paralinguistic feature, the second linguistic feature and the first prompt word for the first audio are used to generate a reply text, which can comprehensively consider the linguistic information, paralinguistic information and instruction prompts in the audio, thereby improving the accuracy and degree of anthropomorphism of the generated reply text.
[0212] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. An interactive processing method, characterized in that: The method comprises: Performing a first speech encoding on the first audio to obtain a first linguistic feature and a first paralinguistic feature; performing feature denoising of a corresponding paralinguistic dimension on the first linguistic feature to obtain a second linguistic feature, and performing feature denoising of a corresponding linguistic dimension on the first paralinguistic feature to obtain a second paralinguistic feature; A reply text is generated based on the second paralinguistic feature, the second linguistic feature, and a first prompt word for the first audio, wherein the reply text is used to interact with the first audio according to an instruction of the first prompt word.
2. The method according to claim 1, characterized in that The performing the first speech encoding on the first audio to obtain the first linguistic feature and the first paralinguistic feature includes: Performing first speech encoding on the first audio through multiple cascaded speech encoding layers to obtain first audio features output by each of the speech encoding layers; Determining a first audio feature output by an i-th speech coding layer among the multiple cascaded speech coding layers as the first paralinguistic feature; The first audio feature output by the j-th speech coding layer in the multiple cascaded speech coding layers is determined as the first linguistic feature, where i and j are both positive integers, and the value of i is not greater than the value of j.
3. The method according to claim 1, characterized in that The performing feature denoising corresponding to the paralinguistic dimension on the first linguistic feature to obtain the second linguistic feature includes: Performing dimension transformation processing on the first linguistic feature to obtain a first dimension transformation result; A first linear transformation is performed on the first dimensional transformation result to obtain the second linguistic feature, wherein the first linear transformation is used to perform feature denoising corresponding to the paralinguistic dimension.
4. The method according to claim 1, wherein The performing feature denoising of the corresponding linguistic dimension on the first paralinguistic feature to obtain the second paralinguistic feature includes: performing a second speech encoding on the first paralinguistic feature to obtain a first paralinguistic feature that has undergone feature denoising, wherein the second speech encoding is used to perform feature denoising on a corresponding linguistic dimension; Pooling the first paralinguistic feature after feature denoising into a set length to obtain a pooling result that meets the set length; A second linear transformation corresponding to the set vector dimension is performed on the pooling result that meets the set length to obtain a second paralinguistic feature that meets the set vector dimension.
5. The method according to claim 1, wherein The method further comprises: Performing the first phase training on the initialized model based on paralinguistic tasks and linguistic tasks to obtain a model that has undergone the first phase training; Performing a second-stage training based on the paralinguistic task on the model trained in the first stage to obtain a model trained in the second stage; Performing a third-stage training based on the linguistic task on the model trained in the second stage to obtain a model trained in the third stage; Performing a fourth-stage training based on the alignment task on the model trained in the third stage to obtain a fourth-stage trained model; The model trained in the fourth stage is used to be called to perform the feature denoising on the first linguistic feature and the first paralinguistic feature.
6. The method according to claim 5, characterized in that The initialized model includes an initialized first adapter and an initialized second adapter; The first stage training of the initialized model based on the paralinguistic task and the linguistic task to obtain the first stage trained model includes: performing a first speech encoding on the second audio to obtain a third linguistic feature and a third paralinguistic feature; performing feature denoising of the corresponding paralinguistic dimension on the third linguistic feature using the initialized first adapter to obtain a fourth linguistic feature, and performing feature denoising of the corresponding linguistic dimension on the third paralinguistic feature using the initialized second adapter to obtain a fourth paralinguistic feature; Performing text encoding on the second prompt word to obtain a second text feature, wherein the second prompt word includes an instruction corresponding to the paralinguistic task and an instruction corresponding to the linguistic task; generating a first text prediction probability based on the fourth paralinguistic feature, the fourth linguistic feature, and the second text feature; Generating a true probability of the first text based on the marked reply text corresponding to the second audio and the second prompt word; Based on the first text prediction probability and the first text true probability, the initialized first adapter and the initialized second adapter in the initialized model are updated to obtain a model trained in the first stage.
7. The method according to claim 5, characterized in that The first-stage trained model includes a first adapter trained in the first stage and a second adapter trained in the first stage; The performing of a second-stage training based on the paralinguistic task on the model trained in the first stage to obtain a model trained in the second stage includes: performing a first speech encoding on the third audio to obtain a fifth linguistic feature and a fifth paralinguistic feature; performing feature denoising of the corresponding paralinguistic dimension on the fifth linguistic feature using the first adapter trained in the first stage to obtain a sixth linguistic feature, and performing feature denoising of the corresponding linguistic dimension on the fifth paralinguistic feature using the second adapter trained in the first stage to obtain the sixth paralinguistic feature; performing text encoding on the third prompt word to obtain a third text feature, and performing text encoding on the text content corresponding to the third audio to obtain a fourth text feature, wherein the third prompt word includes an instruction corresponding to the paralinguistic task; Performing random sampling on the fourth text feature, the sixth linguistic feature, and the missing feature to obtain a seventh linguistic feature; generating a second text prediction probability based on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature; generating a second text true probability based on the marked reply text corresponding to the third audio and the third prompt word; Based on the second text prediction probability and the second text true probability, the second adapter trained in the first stage in the model trained in the first stage is updated to obtain a model trained in the second stage.
8. The method according to claim 7, characterized in that Generating a second text prediction probability based on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature includes: performing a sequence-randomized concatenation process on the sixth paralinguistic feature, the seventh linguistic feature, and the third text feature to obtain a first input feature; Mapping is performed on the first input feature to obtain a second text prediction probability.
9. The method according to claim 5, characterized in that The model trained in the second stage includes a first adapter trained in the first stage and a second adapter trained in the second stage; The performing of the third stage training based on the linguistic task on the model trained in the second stage to obtain the model trained in the third stage includes: performing a first speech encoding on the fourth audio to obtain an eighth linguistic feature and a seventh paralinguistic feature; performing feature denoising of the corresponding paralinguistic dimension on the eighth linguistic feature using the first adapter trained in the first stage to obtain a ninth linguistic feature, and performing feature denoising of the corresponding linguistic dimension on the seventh paralinguistic feature using the second adapter trained in the second stage to obtain an eighth paralinguistic feature; Performing text encoding on the fourth prompt word to obtain a fifth text feature, and performing text encoding on the style description corresponding to the fourth audio to obtain a sixth text feature, wherein the fourth prompt word includes an instruction corresponding to the linguistic task; performing random sampling processing on the sixth text feature, the eighth paralinguistic feature, and the missing feature to obtain a ninth paralinguistic feature; generating a third text prediction probability based on the ninth paralinguistic feature, the ninth linguistic feature, and the fifth text feature; Generating a third text true probability based on the fourth audio and the marked reply text of the fourth prompt word; Based on the third text prediction probability and the third text true probability, the first adapter trained in the first stage in the model trained in the second stage is updated to obtain a model trained in the third stage.
10. The method according to claim 9, characterized in that Generating a third text prediction probability based on the ninth paralinguistic feature, the ninth linguistic feature, and the fifth text feature includes: performing random sequential concatenation processing on the ninth paralinguistic feature, the ninth linguistic feature, and the fifth text feature to obtain a second input feature; Mapping processing is performed on the second input feature to obtain a third text prediction probability.
11. An interactive processing device, characterized in that: The device comprises: an encoding module, configured to perform a first speech encoding on the first audio to obtain a first linguistic feature and a first paralinguistic feature; a mapping module configured to perform feature denoising of a corresponding paralinguistic dimension on the first linguistic feature to obtain a second linguistic feature, and perform feature denoising of the corresponding linguistic dimension on the first paralinguistic feature to obtain a second paralinguistic feature having a different vector dimension from the second linguistic feature; A generation module is configured to generate a response text based on the second paralinguistic feature, the second linguistic feature, and a first prompt word for the first audio, wherein the response text is used to interact with the first audio according to an instruction of the first prompt word.
12. An electronic device, characterized in that: include: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the interactive processing method according to any one of claims 1 to 10 when executing the computer-executable instructions or computer programs stored in the memory.
13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the interactive processing method according to any one of claims 1 to 10 is implemented.
14. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the interactive processing method according to any one of claims 1 to 10 is implemented.