Dialogue information processing method and system, electronic equipment and storage medium
Through large-model technology, the parallel generation of text and audio unit sets is solved, and the problem of low efficiency in dialogue information processing in the existing technology is achieved, real-time and natural voice interaction is achieved.
Patent Information
- Application Number
- CN202510329989.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has insufficient real-time and natural fluency of voice interactions in dialogue information processing, and relying on external components leads to high delays and low processing efficiency.
The big model technology is adopted to convert the conversation input information into text tensors and audio tensors through parallel generation strategies, and use the information processing model trained by multimodal information samples to generate text unit sets and audio unit sets in real time to avoid the process of regenerating speech by text.
Reduces interaction delay, improves the efficiency of dialogue information processing, and achieves a more natural and smooth voice interaction experience.
Smart Images

Figure CN120256569A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of large model technology and natural language processing technology. Specifically, it relates to a method, system, electronic device, and storage medium for processing dialogue information. Background Art
[0002] Currently, when processing dialogue information, the Automatic Speech Recognition (ASR) technology is usually used to convert the dialogue information into text, and then the Text-to-Speech (TTS) technology is used to convert the text into speech to output the speech.
[0003] However, the above method requires generating text first and then converting the text into speech, resulting in insufficient real-time performance and natural fluency of voice interaction, and high latency due to dependence on external components, thus having the problem of low processing efficiency of dialogue information.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a method, system, electronic device, and storage medium for processing dialogue information to at least solve the technical problem of low processing efficiency of dialogue information.
[0006] According to one aspect of the embodiments of this application, a method for processing dialogue information is provided. The method may include: obtaining dialogue input information; converting the dialogue input information into a text tensor and an audio tensor, where the text tensor and the audio tensor match an information processing model, and the information processing model is trained using multimodal information samples; inputting the text tensor and the audio tensor into the information processing model, and using the information processing model to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; converting the text unit set into text output information that matches the dialogue input information, and converting the audio unit set into audio output information that matches the dialogue input information.
[0007] According to another aspect of the embodiments of the present application, a method for determining a model is provided. The method may include: obtaining a multi-modal information sample; training an information processing model by using the multi-modal information sample, where the information processing model is used to analyze a text tensor and an audio tensor corresponding to dialogue input information, and generate a text unit set and an audio unit set in parallel, the text tensor and the audio tensor match the information processing model, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, and the text unit set is used to be converted into a text output information that matches the dialogue input information, and the audio unit set is used to be converted into an audio output information that matches the dialogue input information.
[0008] According to another aspect of the embodiments of the present application, a method for processing dialogue information is provided. The method may include: obtaining voice input information; converting the voice input information into a text tensor and a voice tensor, where the text tensor and the voice tensor match the information processing model, and the information processing model is trained by using the multi-modal information sample; inputting the text tensor and the voice tensor into the information processing model, and using the information processing model to analyze the text tensor and the voice tensor, and generate a text unit set and a voice unit set in parallel, where the voice content corresponding to the voice unit set matches the text content corresponding to the text unit set; converting the text unit set into a text output information that matches the voice input information, and converting the voice unit set into a voice output information that matches the voice input information.
[0009] According to another aspect of the embodiments of the present application, a method for processing dialogue information is provided. The method may include: obtaining dialogue query information for interacting with a virtual avatar; converting the dialogue query information into a text tensor and an audio tensor, where the text tensor and the audio tensor match the information processing model, and the information processing model is trained by using the multi-modal information sample; inputting the text tensor and the audio tensor into the information processing model, and using the information processing model to analyze the text tensor and the audio tensor, and generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; converting the text unit set into a text output information that matches the dialogue query information, and converting the audio unit set into an audio output information that matches the dialogue query information; controlling the virtual avatar to execute an interaction behavior according to the text output information and the audio output information, so as to output dialogue reply information that matches the dialogue query information.
[0010] According to another aspect of the embodiments of the present application, a method for processing dialogue information is provided. The method may include: in response to an input operation instruction acting on an operation interface, inputting dialogue input information; in response to a reply operation instruction acting on the operation interface, outputting text output information that matches the dialogue input information, and audio output information that matches the dialogue input information, where the text output information is obtained by converting a text unit set, the audio output information is obtained by converting an audio unit set, the text unit set and the audio unit set are generated in parallel by analyzing a text tensor and an audio tensor using an information processing model, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, the text tensor and the audio tensor are obtained by converting the dialogue input information, and the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples.
[0011] According to another aspect of the embodiments of the present application, a dialogue information processing system is provided, which is deployed on the mobile device side. The system may include: an information input end for obtaining dialogue input information; an information processing end for converting the dialogue input information into a text tensor and an audio tensor, where the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples; inputting the text tensor and the audio tensor into the information processing model, and using the information processing model to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; converting the text unit set into text output information that matches the dialogue input information, and converting the audio unit set into audio output information that matches the dialogue input information; and an information output end for outputting the text output information and the audio output information.
[0012] According to another aspect of the embodiments of the present application, a computing device is further provided, including: a memory storing an executable program; and a processor for running the program, where when the program runs, it executes the methods in the various embodiments of the present application.
[0013] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory storing an executable program; and a processor connected to the memory through a bus for running the program, where when the program runs, it executes the methods in the various embodiments of the present application.
[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present application.
[0015] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program which, when executed by a processor, implements the methods in the various embodiments of the present application.
[0016] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a non-volatile computer-readable storage medium storing a computer program which, when executed by a processor, implements the methods in the various embodiments of the present application.
[0017] According to another aspect of the embodiments of the present application, there is also provided a computer program which, when executed by a processor, implements the methods in the various embodiments of the present application.
[0018] In the embodiments of the present application, dialogue input information is obtained; the dialogue input information is converted into a text tensor and an audio tensor, where the text tensor and the audio tensor match an information processing model, and the information processing model is trained using multimodal information samples; the text tensor and the audio tensor are input into the information processing model, and the information processing model analyzes the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; the text unit set is converted into text output information matching the dialogue input information, and the audio unit set is converted into audio output information matching the dialogue input information. That is to say, after converting the dialogue input information into a text tensor and an audio tensor in the embodiments of the present application, through a parallel generation strategy, the information processing model is used to generate a text unit set and an audio unit set in parallel, and then the text unit set is converted into text output information, and the audio unit set is converted into audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0019] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application and do not constitute a limitation on the present application. Description of the Drawings
[0020] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0021] Figure 1 is a schematic diagram of an application scenario of a method for processing dialogue information according to an embodiment of the present application;
[0022] Figure 2It is a flowchart of a method for processing conversation information according to an embodiment of the present application;
[0023] Figure 3 It is a flowchart of a method for determining a model according to an embodiment of the present application;
[0024] Figure 4 It is a flowchart of another method for processing conversation information according to an embodiment of the present application;
[0025] Figure 5 It is a flowchart of another method for processing conversation information according to an embodiment of the present application;
[0026] Figure 6 It is a flowchart of yet another method for processing conversation information according to an embodiment of the present application;
[0027] Figure 7 It is a schematic diagram of a system for processing conversation information according to an embodiment of the present application;
[0028] Figure 8 It is a schematic diagram of a Speech to Speech conversation system running locally on a device side according to an embodiment of the present application;
[0029] Figure 9 It is a schematic diagram of a device for processing conversation information according to an embodiment of the present application;
[0030] Figure 10 It is a schematic diagram of a device for determining a model according to an embodiment of the present application;
[0031] Figure 11 It is a schematic diagram of another device for processing conversation information according to an embodiment of the present application;
[0032] Figure 12 It is a schematic diagram of another device for processing conversation information according to an embodiment of the present application;
[0033] Figure 13 It is a schematic diagram of yet another device for processing conversation information according to an embodiment of the present application;
[0034] Figure 14 It is a block diagram of the structure of a computing device according to an embodiment of the present application;
[0035] Figure 15 It is a block diagram of the structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0036] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0037] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0038] The technical solution provided by this application is mainly implemented by using large model technology. Here, the large model refers to a deep learning model with a large number of model parameters, which usually can include hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one hundred trillion model parameters. The large model can also be called the Foundation Model. Through large-scale pre-training of the large model with unlabeled corpus, a pre-trained model with more than hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0039] It should be noted that when the large model is actually applied, the pre-trained model can be fine-tuned with a small number of samples so that the large model can be applied to different tasks. For example, the large model can be widely applied in fields such as Natural Language Processing (NLP), computer vision, and speech processing. Specifically, it can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Caption (IC), and image generation. It can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0040] First, some nouns or terms that appear in the process of describing the embodiments of this application are applicable to the following explanations:
[0041] Text-to-Speech (TTS) is a technology that converts text into speech output. By analyzing the input text, including word segmentation, grammar analysis, semantic understanding, etc., and then, based on the analysis results, each word or phrase is converted into a corresponding phoneme sequence. Considering the pronunciation rules and intonation information of the language, a speech waveform is further generated based on the phoneme sequence. Finally, post-processing is performed on the generated speech waveform, such as volume adjustment, noise cancellation, etc., to output speech information.
[0042] Automatic Speech Recognition (ASR) is a technology that converts speech signals into text. By extracting features from the speech signals, the speech signals are converted into digital feature vectors. Further, through an identification algorithm, the extracted feature vectors are compared with known speech patterns to identify the corresponding phonemes or words. Finally, post-processing is performed on the recognized text, such as spelling correction, context understanding, etc., to output the finally generated text.
[0043] According to the embodiments of this application, a method for processing dialogue information is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0044] The above method provided by the embodiments of this application can be applied to, for example Figure 1 the application scenarios shown, but not limited to this. Figure 1 is a schematic diagram of the application scenario of a method for processing dialogue information according to the embodiments of this application. In Figure 1In the application scenario shown, the large model is deployed on the client device without relying on a server or network connection. The client device here may include, but is not limited to: smartphones, tablets, laptops, PDAs, personal computers, smart home devices, in-vehicle devices, etc.
[0045] On the graphical user interface of the client device, an operation interface can be deployed. In response to an input operation instruction acting on the operation interface, dialogue input information can be input. At this time, the client device can perform the following steps: Step S102, obtain the dialogue input information; Step S104, convert the dialogue input information into a text tensor and an audio tensor, where the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples; Step S106, input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor, and generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; Step S108, convert the text unit set into text output information that matches the dialogue input information, and convert the audio unit set into audio output information that matches the dialogue input information.
[0046] Among them, after the client device obtains the above-mentioned dialogue input information, the preprocessing module on the client device can convert the dialogue input information into a text tensor and an audio tensor, and the text tensor and the audio tensor are in a data format that the information processing model can understand and process. Further input the text tensor and the audio tensor into the information processing model deployed on the client device, and use the information processing model to analyze the text tensor and the audio tensor, and generate a text unit set and an audio unit set in parallel. Finally, the conversion module on the client device converts the text unit set into text output information, and converts the audio unit set into audio output information in parallel. The text output information and the audio output information can be directly provided to the user through the output channel of the client device. For example, the text output information is displayed on the screen, and the audio output information is played through the speaker.
[0047] Considering that the number of model parameters of the large model is huge and the computing resources of the mobile terminal are limited, in the Figure 1 application scenario shown, the large model can also be deployed in the server. The server can connect to one or more client devices through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks.
[0048] On the graphical user interface of the client device, an operation interface for obtaining dialogue input information can be deployed. The client device can interact with the user through the graphical user interface to call the large model, thereby implementing the method provided in the embodiments of the present application. In the embodiments of the present application, the system composed of the client device and the server can perform the following steps: corresponding operations can be performed in the operation interface on the client device to obtain dialogue input information. The client device can obtain the dialogue input information and send it to the server through the network. After receiving the dialogue input information, the server can execute the above steps S102 to step S108. It should be noted that with the rapid development of high-performance computing units, in other application scenarios, the above method provided in the embodiments of the present application can also be applied to a model all-in-one machine. In an alternative embodiment, multiple models are built into the model all-in-one machine, and one model can be selected for adjustment as needed. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the adjusted model to execute the above method provided in the embodiments of the present application. In another alternative embodiment, a trained model is built into the large model all-in-one machine. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the model to execute the above method provided in the embodiments of the present application.
[0049] Further, when it is necessary to train a model in a target scenario, the client can also upload its own dataset. The dataset is sent from the client to the server, enabling the server to use the dataset to adjust the pre-trained model to obtain the model in the target scenario, and then deploy it to the production environment. To facilitate the adjustment of the model, the server can provide complete adjustment tools, development frameworks, and processes, and support multiple adjustment strategies, enabling the adjusted model to better adapt to different field applications and achieve high customization.
[0050] Under the above operating environment, the present application provides a Figure 2 processing method for dialogue information as shown. Figure 2 It is a flowchart of a processing method for dialogue information according to an embodiment of the present application, as Figure 2 shown, and the method may include the following steps:
[0051] Step S202, obtain dialogue input information.
[0052] In the technical solution provided in step S202 of the present application, dialogue input information can be obtained. Among them, the dialogue input information can be the original input data provided by the user or other dialogue participants in the dialogue system, and the information type of the dialogue input information can include text, audio, image, video, etc.
[0053] In this embodiment, in a dialogue system based on artificial intelligence (AI), dialogue input information may mainly include text input information and audio input information. Among them, text input information may be obtained through a text input box, a chat interface, an application programming interface (API) call, and other methods. Audio input information may be obtained through a microphone, an audio acquisition device, an audio input device, and other methods. It should be noted that the above is only an example, and does not specifically limit the method for obtaining text input information and audio input information.
[0054] Step S204: converting the dialogue input information into a text tensor and an audio tensor.
[0055] In the technical solution provided in the above step S204 of the present application, the text tensor and the audio tensor are matched with the information processing model, and the information processing model is trained using multimodal information samples.
[0056] In this embodiment, after acquiring the dialogue input information, the acquired dialogue input information may be converted into a text tensor and an audio tensor. The text tensor may be a tensor obtained by converting the text input information, and the text tensor may include a text vector. The audio tensor may be a tensor obtained by converting the audio input information, and the audio tensor may include an audio vector.
[0057] In this embodiment, the above-mentioned text tensor and audio tensor are matched with an information processing model, and the information processing model can be obtained by training a large model using multimodal information samples. Among them, the multimodal information samples can be data samples containing multiple different information types, for example, the multimodal information samples can be samples containing multiple information types such as text, audio, image, video, etc. The large model can be a basic model that is pre-trained on a large scale on multiple text and language understanding tasks, and can also be called a multimodal pre-trained model, a base large model, etc.
[0058] Optionally, after obtaining the dialogue input information including the text input information and the audio input information, the text input information can be directly converted into a text tensor with the same input dimension as the information processing model. For the audio input information, the audio input information can be feature extracted by an encoder, and the extracted audio features can be further converted into an audio tensor with the same input dimension as the information processing model by an audio adapter.
[0059] Step S206: input the text tensor and the audio tensor into the information processing model, analyze the text tensor and the audio tensor using the information processing model, and generate a text unit set and an audio unit set in parallel.
[0060] In the technical solution provided in step S206 of the present application, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set.
[0061] In this embodiment, after converting the dialogue input information into a text tensor and an audio tensor, the converted text tensor and audio tensor can be input into an information processing model. Further, the information processing model can be used to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel. Among them, the text unit set can include multiple text units, and the text unit can also be called a text token (which can be called a text token), and the text token can be the basic unit after word segmentation of the text tensor and can be used to represent the text content. The audio unit set can include multiple audio units, and the audio unit can also be called an audio token (which can be called an audio token), and the audio token can be an abstract representation of the audio tensor, which can be composed of audio features, and the audio token can be used to represent the audio content.
[0062] Optionally, each audio token in the audio unit set can be an audio feature in the audio tensor, such as spectral information, pitch, rhythm, etc. Each text token in the text unit set can be a basic semantic unit in the text tensor, such as words, phrases, etc. When using the information processing model to analyze the text tensor and the audio tensor, a parallel generation strategy can be adopted to generate the text unit set and the audio unit set, that is, while generating the text unit set, the corresponding audio unit set is generated in parallel, so as to ensure that the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, so as to avoid the delay caused by first generating the text and then converting the text into speech through the TTS technology.
[0063] It should be noted that the information processing model will refer to the generated text tokens when generating audio tokens, that is, the text content guides the audio content, so as to ensure the accuracy and natural fluency of the voice output. Through the above parallel generation strategy, this embodiment can realize real-time voice dialogue, and at the same time ensure the consistency and accuracy of the voice content and the text content, so as to provide a more natural and fluent interaction experience.
[0064] For example, this embodiment can use the information processing model to analyze the text tensor and the audio tensor and generate 8-way tokens in parallel. Among them, the 8-way tokens can include 1-way text tokens and 7-way audio tokens. The 1-way text tokens can form a text unit set, and the 7-way audio tokens can form an audio unit set.
[0065] It should be noted that when the information processing model analyzes the text tensor and the audio tensor and generates the text unit set and the audio unit set in parallel in this embodiment, it does not necessarily output a single modality based on a single modality. For example, the information processing model can be used to analyze the text tensor to generate the audio unit set, or the information processing model can be used to analyze the audio tensor to generate the text unit set and the audio unit set. That is, the information processing model is used to analyze a single tensor (text tensor or audio tensor) to generate a single unit set (text unit set or audio unit set); the information processing model is used to analyze a single tensor to generate a multi-unit set (text unit set and audio unit set); the information processing model is used to analyze multiple tensors (text tensor and audio tensor) to generate a single unit set; the information processing model is used to analyze multiple tensors to generate a multi-unit set.
[0066] Step S208: Convert the text unit set into text output information that matches the dialogue input information, and convert the audio unit set into audio output information that matches the dialogue input information.
[0067] In the technical solution provided in step S208 of the present application, after using the information processing model to analyze the text tensor and the audio tensor and generating the text unit set and the audio unit set in parallel, the generated text unit set can be decoded to obtain text output information that matches the dialogue input information, and the generated audio unit set can be decoded to obtain audio output information that matches the dialogue input information. Among them, the text output information can be the text reply information of the dialogue input information expressed in text form, and the text output information can include a text token sequence. The audio output information can be the audio reply information of the dialogue input information expressed in audio form, and the audio output information can include an audio stream.
[0068] Optionally, after using the information processing model to analyze the text tensor and the audio tensor and generating the text unit set and the audio unit set in parallel, the text decoder can be used to decode the text unit set to obtain text output information that matches the dialogue input information, and the audio decoder can be used to decode the audio unit set to obtain audio output information that matches the dialogue input information.
[0069] Optionally, during the process of analyzing the text tensor and the audio tensor using the information processing model and generating the text unit set and the audio unit set in parallel, the information processing model can determine the structure and content of the output (text unit set and audio unit set) according to the type of the dialogue input information (such as questions, instructions, statements, etc.). For example, when the dialogue input information is a query message, the information processing model can generate in parallel a text unit set and an audio unit set containing the reply information corresponding to the query message. Further, when using the text decoder to decode the text unit set containing the reply information, the obtained text output information includes the reply information that directly answers or responds to the question raised by the dialogue input information, so as to match the dialogue input information. When using the audio decoder to decode the audio unit set containing the reply information, the obtained audio output information is consistent with the text output information, includes the reply information corresponding to the dialogue input information, and the obtained audio output information includes the same sentences as the dialogue input information (such as happy, worried, etc.), so as to match the dialogue input information in terms of semantics, tone, etc.
[0070] Optionally, when using the text decoder to decode the text unit set, the generated text token sequence can be directly output as the text output information. When using the audio decoder to decode the audio unit set, the generated audio tokens can be converted into an audio stream through text-to-speech synthesis technology, and then the audio stream can be output as the audio output information.
[0071] Through the above steps S202 to S208 of the present application, the dialogue input information is obtained; the dialogue input information is converted into a text tensor and an audio tensor, wherein the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multi-modal information samples; the text tensor and the audio tensor are input into the information processing model, and the information processing model analyzes the text tensor and the audio tensor to generate in parallel a text unit set and an audio unit set, wherein the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; the text unit set is converted into text output information that matches the dialogue input information, and the audio unit set is converted into audio output information that matches the dialogue input information. That is to say, after converting the dialogue input information into a text tensor and an audio tensor in the embodiments of the present application, through a parallel generation strategy, the information processing model is used to generate in parallel a text unit set and an audio unit set, and then the text unit set is converted into text output information, and the audio unit set is converted into audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0072] The method of analyzing the text tensor and the audio tensor using the information processing model in the above embodiment and generating the text unit set and the audio unit set in parallel will be further introduced below.
[0073] As an alternative implementation, in step S206, input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor, and generate the text unit set and the audio unit set in parallel, including: concatenating the text tensor and the audio tensor to obtain a target tensor, where the target tensor matches the information processing model; input the target tensor into the information processing model, and use the information processing model to analyze the target tensor, and generate the text unit set and the audio unit set in parallel.
[0074] In this embodiment, after converting the dialogue input information into a text tensor and an audio tensor, the text tensor and the audio tensor can be concatenated to obtain a target tensor. Further, the target tensor can be input into the information processing model, and then the information processing model is used to analyze the target tensor, and the text unit set and the audio unit set are generated in parallel. Among them, the target tensor matches the information processing model.
[0075] Optionally, after extracting the audio feature from the audio input information through the encoder and then converting the extracted audio feature into an audio tensor with the same input dimension as the information processing model through the audio adapter, the audio tensor can be concatenated with the text tensor converted from the text input information to obtain a target tensor, and the target tensor is the tensor dimension that the large model can receive.
[0076] In an alternative example, usually, speech recognition technology is used to convert the audio input information into text, and then the text is converted into audio through TTS technology. However, in the embodiment of the present application, through the parallel generation strategy, the intermediate text conversion step can be skipped, and the information processing model is directly used to analyze the target tensor, and the text unit set and the audio unit set are generated in parallel, which can significantly reduce the latency from receiving the input to generating the output, thereby improving the real-time performance of the interaction.
[0077] In this embodiment, by concatenating the text tensor and the audio tensor to obtain a target tensor, that is, through a unified tensor dimension, the information processing model can consider both text information and audio information at the same time, thereby providing a richer and more natural dialogue experience.
[0078] The method of analyzing the target tensor using the information processing model in the above embodiment and generating the text unit set and the audio unit set in parallel will be further introduced below.
[0079] As an alternative implementation, an information processing model is used to analyze a target tensor and generate a text unit set and an audio unit set in parallel, including: in the process of using the language modeling component of the information processing model to sample text units from the target tensor, the obtained text units are used to determine the audio parameters of the audio units to be generated, where the audio parameters are used to characterize the quality level of the audio units to be generated; in the process of using the language modeling component to sample text units other than the obtained text units from the target tensor, audio units that meet the audio parameters are generated to obtain an audio unit set, and the collected text units are generated into a text unit set.
[0080] In this embodiment, after concatenating the text tensor and the audio tensor to obtain the target tensor, in the process of using the language modeling component of the information processing model to sample text units from the target tensor, the obtained text units can be used to determine the audio parameters of the audio units to be generated. After determining the audio parameters of the audio units to be generated, further in the process of using the language modeling component to sample text units other than the obtained text units from the target tensor, audio units that meet the audio parameters are generated to obtain an audio unit set, and the collected text units are generated into a text unit set. Among them, the language modeling component can be used to generate text units and audio units. The audio parameters can be used to characterize the quality level of the audio units to be generated. For example, the audio parameters can include parameters such as pitch, tone, rhythm, and timbre. This is only an example here, and no specific restrictions are imposed on the content of the audio parameters.
[0081] Optionally, since the generation of text units is usually faster than that of audio units, the language modeling component can be used to first generate some text units, and then the obtained text units are used to determine the audio parameters of the audio units to be generated. These audio parameters can be used to guide the generation of audio units (i.e., using text units to guide the generation of audio units) to improve the output quality of audio units. Further, in the process of using the language modeling component to generate the remaining ungenerated text units, that is, while generating text units, audio units that meet the audio parameters can be generated to obtain an audio unit set and a text unit set.
[0082] Optionally, a text token can be sampled from the target tensor using a language modeling component. While generating the text token, the audio parameters of the audio token to be generated can be determined using the generated text token. Since the determination of the audio parameters is affected by the text token, it can be ensured that the generated audio is semantically consistent with the text and has good audio quality. After determining the audio parameters, the language modeling component can be further used to generate audio tokens that meet the audio parameters, and the generation of the text token and the audio token is parallel, that is, while generating the set of text units, the corresponding set of audio units is also determined and generated. Among them, the audio token is generated through text-to-speech synthesis technology to ensure real-time output.
[0083] In this embodiment, by using the text unit to determine the audio parameters and then guiding the generation of the audio unit, it can be ensured that the generated audio unit is semantically consistent with the text unit, and the selection of the audio parameters directly affects the generation quality of the audio unit and the speech synthesis effect. By optimizing the audio parameters based on the text unit, higher-quality audio units can be generated, thereby improving the naturalness and clarity of the speech synthesis.
[0084] Next, a further introduction is made to the method of generating the audio unit that meets the audio parameters and obtaining the set of audio units in this embodiment.
[0085] As an alternative implementation, generating the audio unit that meets the audio parameters and obtaining the set of audio units includes: sequentially generating multi-layer audio units that meet the audio parameters to obtain the set of audio units, where the multi-layer audio units are in a parallel relationship.
[0086] In this embodiment, after using the obtained text unit to determine the audio parameters of the audio unit to be generated, during the process of using the language modeling component to sample text units other than the obtained text unit from the target tensor, multi-layer audio units that meet the audio parameters can be sequentially generated to obtain the set of audio units. Among them, the multi-layer audio units can be used to represent multiple audio token layers, and the multi-layer audio units are in a parallel relationship. For example, the number of layers of the multi-layer audio units can be seven. This is only an example and does not specifically limit the number of layers of the multi-layer audio units.
[0087] Optionally, after determining the audio parameters using the text token, in order to accelerate the generation of the audio token, a parallel decoding strategy can be adopted to simultaneously generate multi-layer audio units (i.e., multiple audio token layers) that meet the audio parameters, thereby obtaining the set of audio units.
[0088] For example, assume that seven audio token layers are generated, and audio tokens are generated sequentially from the first layer to the seventh layer, that is, seven audio tokens are output. The text token is output first, that is, one text token is output. It should be noted that the seven tokens output through seven layers are used by an audio decoder to generate one phoneme, and multiple phonemes can generate a sequence of audio.
[0089] This embodiment adopts a parallel decoding strategy to generate audio tokens on multiple audio token layers. Compared with the method of generating each audio token sequentially, the total time for generating audio tokens is significantly reduced. Since audio token generation is usually more time-consuming than text token generation, the above strategy can effectively improve the response speed of the interaction, especially the delay of audio output.
[0090] Next, the method of converting the text unit set into text output information matching the dialogue input information and converting the audio unit set into audio output information matching the dialogue input information in this embodiment will be further introduced.
[0091] As an alternative embodiment, in step S208, converting the text unit set into text output information matching the dialogue input information and converting the audio unit set into audio output information matching the dialogue input information includes: performing parallel decoding on the text unit set and the audio unit set to obtain text output information corresponding to the text unit set and audio output information corresponding to the audio unit set.
[0092] In this embodiment, after inputting the text tensor and the audio tensor into the information processing model and using the information processing model to analyze the text tensor and the audio tensor to generate the text unit set and the audio unit set in parallel, the generated text unit set and audio unit set can be decoded in parallel to obtain text output information corresponding to the text unit set and audio output information corresponding to the audio unit set.
[0093] Optionally, after generating the text unit set and the audio unit set in parallel, the generated text unit set and audio unit set can be decoded in parallel, that is, the decoding processes of the text unit set and the audio unit set are carried out simultaneously, rather than waiting sequentially for one process to complete before starting the next process. This parallel decoding strategy reduces the latency time for generating the entire response.
[0094] Optionally, when decoding the text unit set, the text unit set can be decoded into readable text output information, such as text on a screen or any other form of text display. When decoding the audio unit set, the audio unit set can be decoded into a continuous audio stream through TTS or similar speech synthesis technology, that is, audio output information. This process converts audio tokens into actual sounds, which can be a voice response played through a speaker.
[0095] This embodiment adopts a parallel decoding strategy so that the text output information and the audio output information can be generated simultaneously, thereby shortening the time from receiving the input to outputting the response, improving the real-time response ability of the dialogue system, and making the dialogue smoother.
[0096] The following further introduces the method of parallelly decoding the above-mentioned text unit set and audio unit set in this embodiment to obtain the text output information corresponding to the text unit set and the audio output information corresponding to the audio unit set.
[0097] As an optional implementation manner, the text output information includes a text unit sequence, and the audio output information includes an audio stream. Parallelly decoding the text unit set and the audio unit set to obtain the text output information corresponding to the text unit set and the audio output information corresponding to the audio unit set includes: using the text decoder of the information processing model to decode the text unit set to obtain the text unit sequence corresponding to the text unit set, and simultaneously using the audio decoder of the information processing model to decode the audio unit set to obtain the audio stream corresponding to the audio unit set.
[0098] In this embodiment, the text output information may include a text unit sequence, and the audio output information may include an audio stream. After inputting the text tensor and the audio tensor into the information processing model and using the information processing model to analyze the text tensor and the audio tensor to generate the text unit set and the audio unit set in parallel, the text decoder of the information processing model can be used to decode the text unit set to obtain the text unit sequence corresponding to the text unit set, and simultaneously the audio decoder of the information processing model can be used to decode the audio unit set to obtain the audio stream corresponding to the audio unit set.
[0099] Optionally, using the text decoder in the information processing model, the generated text unit set can be decoded into a text unit sequence (i.e., a text token sequence). Among them, the text decoder can be used to convert the generated text tokens into a readable text format, that is, a text unit sequence, so as to ensure that the user can receive and understand the text response generated by the dialogue system visually.
[0100] Optionally, using the audio decoder in the information processing model, the generated audio unit set can be decoded, and the generated audio tokens can be converted into an audio stream output through text-to-speech synthesis technology. Among them, the audio decoder can be used to convert the audio token sequence into a continuous audio signal, that is, an audio stream, so as to ensure that the user can receive the voice response generated by the dialogue system auditorily.
[0101] In this embodiment, the text decoder decodes the text unit set, and at the same time, the audio decoder decodes the audio unit set, which can significantly reduce the output latency of the text unit sequence and the audio stream, thereby improving the user experience.
[0102] The method of converting the dialogue input information into a text tensor and an audio tensor in the above embodiment will be further introduced below.
[0103] As an alternative implementation, the method further includes: determining the target data dimension of the input data of the information processing model; step S204, converting the dialogue input information into a text tensor and an audio tensor, including: using the text encoder of the information processing model to extract initial text units from the dialogue input information, and in the embedding layer of the information processing model, converting the initial text units into text tensors of the target data dimension; using the audio encoder of the information processing model to extract initial audio features from the dialogue input information, and using the audio adapter of the information processing model to convert the initial audio features into audio tensors of the target data dimension.
[0104] In this embodiment, the target data dimension of the input data of the information processing model can be determined, where the target data dimension can be the dimension of the input data set in advance according to the actual situation. For example, the target data dimension can be 896 dimensions. Here is only an example, and no specific limit is imposed on the value of the target data dimension. After obtaining the dialogue input information, the initial text units can be extracted from the dialogue input information by using the text encoder of the information processing model, and the initial text units are converted into text tensors of the target data dimension in the embedding layer of the information processing model. The initial audio features can be extracted from the dialogue input information by using the audio encoder of the information processing model, and the initial audio features are converted into audio tensors of the target data dimension by using the audio adapter of the information processing model. Among them, the initial text units can be units (such as tokens or words) that form the basic components of the text by splitting the text input information. The initial audio features can be the preliminary digital and vectorized representations of the audio input information extracted from the audio input information.
[0105] Optionally, after obtaining the text input information, the text input information can be converted into initial text unit tokens through a tokenizer, and then the initial text units are converted into text tensors of the target data dimension through an embedding layer. After obtaining the audio input information, the audio input information can be extracted into initial audio features through an encoder, and then the initial audio features are converted into audio tensors of the target data dimension through an audio adapter.
[0106] For example, it can be determined that the target data dimension of the input data of the information processing model is 896 dimensions. After obtaining the text input information, the text input information is converted into initial text tokens through a tokenizer, and then the initial text tokens are converted into an 896-dimensional text tensor through an embedding layer. After obtaining the audio input information, the audio input information can be extracted into initial audio features through an encoder. For example, it is extracted into a 768-dimensional vector, and then the 768-dimensional vector is converted into an 896-dimensional audio tensor through an adapter. Further, the text tensor and the audio tensor are superimposed to form a final target vector and input to the information processing model for inference to obtain text output information and audio output information.
[0107] In this embodiment, by encoding the speech input information and the text input information into a text tensor and an audio tensor with the same target data dimension, the information processing model can effectively fuse the information of the two modalities, thereby realizing a more natural and coherent multi-modal dialogue.
[0108] Next, the method of converting the dialogue input information into a text tensor and an audio tensor in this embodiment will be further introduced.
[0109] As an alternative implementation, the dialogue input information at least includes first dialogue input information and second dialogue input information. Step S204, converting the dialogue input information into a text tensor and an audio tensor, includes: converting the first dialogue input information into a first text tensor and an audio tensor, and converting the second dialogue input information into a second text tensor, where the text quality of the second text tensor is higher than that of the first text tensor.
[0110] In this embodiment, the dialogue input information can at least include first dialogue input information and second dialogue input information. After obtaining the dialogue input information, the first dialogue input information can be converted into a first text tensor and an audio tensor, and the second dialogue input information can be converted into a second text tensor. Among them, the first dialogue input information can be a dialogue input sample including audio input information and text input information, which can be used to generate text responses and audio responses simultaneously. The second dialogue input information can be a dialogue input sample including only text input information, which can be used to generate only text responses. The first text tensor can be a text tensor converted from the first dialogue input information, and the second text tensor can be a text tensor converted from the second dialogue input information, and the text quality of the second text tensor is higher than that of the first text tensor.
[0111] Optionally, expand a single dialogue input message into batch processing. The dialogue input message can at least include a first dialogue input message and a second dialogue input message. Among them, the first dialogue input message can be used to generate a text response and an audio response, the second dialogue input message is only used to generate a text response, and the audio response of the first dialogue input message is streamed based on the text response content of the second dialogue input message, which can ensure the consistency of the audio response and the text response in semantics and information content.
[0112] This embodiment expands a single input into batch processing and adopts a parallel generation strategy, significantly accelerating the inference speed of the information processing model, thereby improving the efficiency of real-time interaction.
[0113] Next, a further introduction will be given to the method of analyzing text tensors and audio tensors using the information processing model in this embodiment and parallelly generating a text unit set and an audio unit set.
[0114] As an alternative implementation manner, step S206, analyzing text tensors and audio tensors using the information processing model and parallelly generating a text unit set and an audio unit set, includes: analyzing a first text tensor, an audio tensor, and a second text tensor using the information processing model, and parallelly generating a first text unit set and an audio unit set corresponding to the first dialogue input message, and a second text unit set corresponding to the second dialogue input message.
[0115] In this embodiment, after converting the first dialogue input message into a first text tensor and an audio tensor, and converting the second dialogue input message into a second text tensor, the information processing model can be used to analyze the first text tensor, the audio tensor, and the second text tensor, and parallelly generate a first text unit set and an audio unit set corresponding to the first dialogue input message, and a second text unit set corresponding to the second dialogue input message. Among them, the first text unit set can be a text unit set generated from the first text tensor and the audio tensor, and the second text unit set can be a text unit set generated only from the second text tensor.
[0116] For example, an information processing model can be used to analyze the first text tensor and the audio tensor, and generate 8 tokens in parallel. Among them, the 8 tokens can include 1 text token and 7 audio tokens. The 1 text token is the first text unit set, and the 7 audio tokens can form an audio unit set. The information processing model can be used to analyze the second text tensor and generate 1 text token, which is the second text unit set. Since outputting both text and audio requires outputting 8 tokens, and there is only 1 text token, compared with only outputting 1 text token, the text quality of the second text unit set is higher than that of the first text unit set.
[0117] This embodiment generates a personalized response based on more comprehensive information (i.e., the text unit set and the audio unit set). At the same time, the high-quality text response of the second text unit set can be used to enhance the situational awareness and make the conversation more intelligent and user-friendly.
[0118] As an alternative implementation, since the text quality of the second text unit set is higher than that of the first text unit set, the method further includes: replacing the first text unit set with the second text unit set.
[0119] In this embodiment, since the text quality of the second text unit set is higher than that of the first text unit set, the first text unit set can be replaced with the second text unit set.
[0120] Optionally, since the text quality of the second text unit set is higher than that of the first text unit set, the first text unit set can be discarded from the first conversation input information, and the second text unit set can be embedded into the corresponding text token position of the first conversation input information. That is, the high-quality second text unit set is used to replace the first text unit set, thereby improving the overall response quality of the information processing model.
[0121] Next, the method of converting the first conversation input information into the first text tensor and the audio tensor, and converting the second conversation input information into the second text tensor in this embodiment will be further introduced.
[0122] As an alternative implementation, converting the first conversation input information into the first text tensor and the audio tensor includes: converting the text input information in the first conversation input information into the first text tensor, and converting the audio input information in the first conversation input information into the audio tensor; converting the second conversation input information into the second text tensor includes: converting the text input information in the second conversation input information into the second text tensor.
[0123] In this embodiment, after obtaining the first conversation input information, the text input information in the first conversation input information can be converted into a first text tensor, and the audio input information in the first conversation input information can be converted into an audio tensor. After obtaining the second conversation input information, the text input information in the second conversation input information can be converted into a second text tensor.
[0124] Optionally, for the text input information in the first conversation input information and the text input information in the second conversation input information, natural language processing techniques, such as word embedding or an encoder, can be used to convert the text input information in the first conversation input information into a first text tensor and convert the text input information in the second conversation input information into a second text tensor. For the audio input information in the first conversation input information, feature extraction can be performed on the audio input information through an encoder, and the extracted audio features can be further converted into an audio tensor through an audio adapter.
[0125] This embodiment inputs the first text tensor and the audio tensor into an information processing model to generate a text response and an audio response in parallel, realizing a low-latency interaction experience. The second text tensor is separately used as an input to the information processing model to generate a high-quality text response, thereby improving the overall response quality of the information processing model.
[0126] The method for obtaining conversation input information in this embodiment is further introduced below.
[0127] As an optional implementation manner, obtaining conversation input information includes: obtaining conversation input information in different languages.
[0128] In this embodiment, conversation input information in different languages can be obtained. Among them, different languages can at least include Chinese and English. This is only an example here and does not specifically limit the languages of the conversation input information.
[0129] This embodiment can obtain conversation input information in different languages and generate corresponding responses in real time. Users can freely switch languages during the conversation, significantly improving the cross-language processing ability of the conversation system, thereby providing a more personalized, efficient, and cross-cultural user experience.
[0130] In the embodiment of the present application, after converting the conversation input information into a text tensor and an audio tensor, through a parallel generation strategy, an information processing model is used to generate a text unit set and an audio unit set in parallel, and then the text unit set is converted into text output information, and the audio unit set is converted into audio output information in parallel, avoiding generating text first and then generating speech, reducing the latency of the interaction, and thus achieving the technical effect of improving the processing efficiency of conversation information and solving the technical problem of low processing efficiency of conversation information.
[0131] The embodiment of the present application also provides a method for determining a model, Figure 3 which is a flowchart of a method for determining a model according to an embodiment of the present application. As Figure 3 shown, the method may include the following steps:
[0132] Step S302, obtain multimodal information samples.
[0133] In the technical solution provided in step S302 of the present application, multimodal information samples can be obtained. Among them, the multimodal information samples can be data samples containing multiple different information types. For example, the multimodal information samples can be samples containing multiple information types such as text, audio, images, videos, etc.
[0134] Step S304, use the multimodal information samples to train and obtain an information processing model.
[0135] In the technical solution provided in step S304 of the present application, the information processing model is used to analyze the text tensor and audio tensor corresponding to the dialogue input information, and generate a text unit set and an audio unit set in parallel. The text tensor and audio tensor match the information processing model, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, and the text unit set is used to be converted into text output information matching the dialogue input information, and the audio unit set is used to be converted into audio output information matching the dialogue input information.
[0136] In this embodiment, after obtaining the multimodal information samples, the obtained multimodal information samples can be used to train a large model to obtain an information processing model. After obtaining the dialogue input information, and converting the dialogue input information into a text tensor and an audio tensor, the text tensor and audio tensor match the information processing model, and the text tensor and audio tensor can be input into the information processing model, and the information processing model is used to analyze the text tensor and audio tensor, and generate a text unit set and an audio unit set in parallel. Further decode the text unit set to obtain text output information matching the dialogue input information, and decode the audio unit set to obtain audio output information matching the dialogue input information.
[0137] Optionally, after obtaining the dialogue input information, the text input information in the dialogue input information can be directly converted into a text tensor with the same input dimension as the information processing model, and the audio input information in the dialogue input information is subjected to feature extraction by an encoder, and further the extracted audio features are converted into an audio tensor with the same input dimension as the information processing model through an audio adapter.
[0138] Optionally, after analyzing the text tensor and the audio tensor using the information processing model and generating the text unit set and the audio unit set in parallel, the text decoder can be used to decode the text unit set to obtain text output information that matches the dialogue input information, and the audio decoder can be used to decode the audio unit set to obtain audio output information that matches the dialogue input information.
[0139] Through the above steps S302 to S304 of the present application, multi-modal information samples are obtained; an information processing model is trained using the multi-modal information samples, where the information processing model is used to analyze the text tensor and the audio tensor corresponding to the dialogue input information, and generate the text unit set and the audio unit set in parallel. The text tensor and the audio tensor match the information processing model, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, and the text unit set is used to be converted into text output information that matches the dialogue input information, and the audio unit set is used to be converted into audio output information that matches the dialogue input information. That is to say, the embodiment of the present application uses multi-modal information samples to train an information processing model, and uses the information processing model to generate the text unit set and the audio unit set in parallel through a parallel generation strategy, and then converts the text unit set into text output information, and converts the audio unit set into audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0140] The method for obtaining the multi-modal information samples in the above embodiment will be further introduced below.
[0141] As an optional implementation manner, the multi-modal information samples include text dialogue information samples and audio dialogue information samples. Step S302 of obtaining the multi-modal information samples includes: determining different dialogue tasks; obtaining the text dialogue information samples and the audio dialogue information samples under the dialogue tasks.
[0142] In this embodiment, the multi-modal information samples may include text dialogue information samples and audio dialogue information samples. Different dialogue tasks can be determined, and the text dialogue information samples and audio dialogue information samples under the dialogue tasks can be obtained. Among them, the dialogue task can be the function that the dialogue system should execute or the type of problem to be solved. For example, it can be tasks such as Chinese-English automatic speech recognition (ASR), text question answering (Text QA for short), and audio question answering (Audio QA for short). This is only for illustration and does not specifically limit the content of the dialogue task. The text dialogue information sample can be the dialogue data in text form collected under the dialogue task. The audio dialogue information sample can be the voice dialogue data collected under the dialogue task.
[0143] Optionally, the large model is trained using multiple data sets (such as a voice data set and a large corpus). These data sets can cover multiple tasks such as Chinese-English ASR, Text QA, and Audio QA. After determining different dialogue tasks, the text dialogue information samples and audio dialogue information samples under the dialogue tasks can be obtained.
[0144] This embodiment can train the large model by combining data sets of different dialogue tasks, enabling the obtained information processing model to understand and adapt to a wider range of dialogue scenarios and enhancing the diversity of the dialogue of the information processing model.
[0145] Next, the method for obtaining the text dialogue information samples and audio dialogue information samples under the dialogue tasks in this embodiment will be further introduced.
[0146] As an optional implementation manner, obtaining the text dialogue information samples and audio dialogue information samples under the dialogue tasks includes: obtaining the text dialogue information samples in different languages and the audio dialogue information samples in different languages under the dialogue task.
[0147] In this embodiment, after determining different dialogue tasks, the text dialogue information samples in different languages and the audio dialogue information samples in different languages under the dialogue tasks can be obtained. Among them, different languages can at least include Chinese and English. This is only for illustration and does not specifically limit the languages of the dialogue input information.
[0148] Optionally, after determining different dialogue tasks, a dedicated voice assistant data set can be synthesized using a text-to-speech model. Through zero-shot TTS technology, the text data is converted into Chinese-English voice dialogue data to increase the diversity and quantity of the training data, thereby adding the Chinese-English voice interaction ability to the information processing model.
[0149] In this embodiment, speech synthesis technology can be used to generate a large amount of speech data from limited text data, significantly increasing the diversity and quantity of training data. By converting text data into Chinese-English speech dialogue data, the information processing model can learn to process multilingual inputs, enhancing the dialogue ability and speech synthesis quality in different language environments.
[0150] Next, a further introduction will be given to the method of training an information processing model using the above-mentioned multi-modal information samples in this embodiment.
[0151] As an alternative implementation, in step S304, training an information processing model using multi-modal information samples includes: performing modality alignment on text dialogue information samples and audio dialogue information samples; using the aligned text dialogue information samples and the aligned audio dialogue information samples to train the initial model corresponding to the information processing model to obtain the information processing model.
[0152] In this embodiment, after obtaining text dialogue information samples and audio dialogue information samples under a dialogue task, modality alignment (Modality Alignment) can be performed on the obtained text dialogue information samples and audio dialogue information samples, and further, the aligned text dialogue information samples and the aligned audio dialogue information samples are used to train the initial model corresponding to the information processing model to obtain the information processing model.
[0153] Optionally, when performing modality alignment on text dialogue information samples and audio dialogue information samples, speech recognition and speech synthesis data can be used for training, and only the gradients of the audio adapter and the speech generation adapter are allowed to be updated, thereby enhancing the ability of the information processing model to understand and generate speech.
[0154] Through modality alignment training, this embodiment can better learn and understand the characteristics of speech signals, accurately recognize and analyze speech content even in a noisy environment, and improve the accuracy of speech recognition.
[0155] Next, a further introduction will be given to the method of performing modality alignment on the text dialogue information samples and the audio dialogue information samples in this embodiment.
[0156] As an alternative implementation, performing modality alignment on text dialogue information samples and audio dialogue information samples includes: converting the text dialogue information samples into text tensor samples with a target data dimension matching the initial model; extracting initial audio feature samples from the audio dialogue information samples, and converting the initial audio feature samples into audio tensor samples with the target data dimension.
[0157] In this embodiment, after obtaining the text dialogue information sample and the audio dialogue information sample under the dialogue task, the text dialogue information sample can be converted into a text tensor sample with a target data dimension that matches the initial model, and an initial audio feature sample can be extracted from the audio dialogue information sample, and the initial audio feature sample is converted into an audio tensor sample with a target data dimension. Among them, the text tensor sample can be a tensor sample obtained by converting the text dialogue information sample. The audio tensor sample can be a tensor sample obtained by converting the audio dialogue information sample. The initial audio feature sample can be an audio feature sample extracted from the audio dialogue information sample.
[0158] Optionally, after obtaining the text dialogue information sample and the audio dialogue information sample under the dialogue task, the text dialogue information sample can be cleaned and formatted to ensure that the text dialogue information sample is suitable for large model training, and then the text dialogue information sample is converted into a text tensor sample with a target data dimension that matches the large model. The initial audio feature sample can be obtained by extracting features from the audio dialogue information sample through an encoder (that is, using a pre-trained model to extract features and perform audio encoding on the audio dialogue information sample), and then the extracted initial audio feature sample is converted into an audio tensor sample with a target data dimension through an audio adapter.
[0159] In order to adapt to the input format of the large model, this embodiment converts the text dialogue information sample and the audio dialogue information sample into a tensor form that the large model can process, which is convenient for the large model to perform cross-modal joint learning and processing.
[0160] Next, a further introduction will be made to the method of training the initial model with the aligned text dialogue information sample and the aligned audio dialogue information sample in this embodiment to obtain an information processing model.
[0161] As an optional implementation manner, using the aligned text dialogue information sample and the aligned audio dialogue information sample to train the initial model to obtain an information processing model includes: concatenating the text tensor sample and the audio tensor sample to obtain a target tensor sample, where the target tensor sample matches the initial model; using the target tensor sample to train the initial model to obtain an information processing model.
[0162] In this embodiment, after performing modal alignment on the text dialogue information sample and the audio dialogue information sample, the text tensor sample and the audio tensor sample can be concatenated to obtain a target tensor sample, and then the target tensor sample is used to train the initial model to obtain an information processing model. Among them, the target tensor sample matches the initial model.
[0163] Optionally, after extracting features from the audio dialogue information samples through an encoder and converting the extracted initial audio feature samples into audio tensor samples with the target data dimension through an audio adapter, the audio tensor samples can be concatenated with the text tensor samples converted from the text dialogue information samples to obtain target tensor samples, which are the tensor dimensions that the large model can receive.
[0164] In this embodiment, by concatenating the text tensor samples and the audio tensor samples to obtain target tensor samples, the large model can process information in two modalities simultaneously, thereby achieving cross-modal understanding of text and speech.
[0165] Next, a further introduction will be given to the method of training the initial model with the aligned text dialogue information samples and the aligned audio dialogue information samples to obtain an information processing model in this embodiment.
[0166] As an alternative implementation, training the initial model with the aligned text dialogue information samples and the aligned audio dialogue information samples to obtain an information processing model includes: training the text processing model of the initial model with the aligned audio dialogue information samples and the aligned audio dialogue information samples to obtain an information processing model.
[0167] In this embodiment, after aligning the modalities of the text dialogue information samples and the audio dialogue information samples, the text processing model of the initial model can be trained with the aligned audio dialogue information samples and the aligned audio dialogue information samples to obtain an information processing model. Among them, the text processing model can be a model that generates text based on audio input.
[0168] Optionally, after aligning the modalities of the text dialogue information samples and the audio dialogue information samples, the adapter can be trained. Optionally, data such as speech recognition, spoken Q&A, and text response tasks are used for training, so as to train the text generation ability of the large model given an audio input.
[0169] In the adapter training stage of this embodiment, by freezing the audio adapter, the large model can focus on improving and optimizing the text generation ability, ensuring that the original text processing and generation performance will not decline after adding the speech modality, so as to effectively integrate the speech input while maintaining the original text processing ability, realizing more intelligent, natural, and fluent text generation under speech guidance, and providing users with more satisfactory and efficient dialogue services.
[0170] Next, a further introduction will be given to the method of training the initial model with the aligned text dialogue information samples and the aligned audio dialogue information samples to obtain an information processing model in this embodiment.
[0171] As an alternative implementation, the initial model is trained using the aligned text dialogue information samples and the aligned audio dialogue information samples to obtain an information processing model, including: training the initial model using the aligned text dialogue information samples and the aligned audio dialogue information samples; training the trained initial model using a target data sample including at least speech data samples to obtain an information processing model.
[0172] In this embodiment, after performing modality alignment on the text dialogue information samples and the audio dialogue information samples, the initial model can be trained using the aligned text dialogue information samples and the aligned audio dialogue information samples. Further, the trained initial model is trained using a target data sample including at least speech data samples to obtain an information processing model. Among them, the target data sample can be comprehensive data including speech data samples (speech modality).
[0173] Optionally, after training the large model using the aligned text dialogue information samples and the aligned audio dialogue information samples, the trained large model can be fine-tuned in multiple modalities. Optionally, the trained large model is trained using comprehensive data including speech modality, that is, the trained large model is fine-tuned using comprehensive data to obtain an information processing model.
[0174] Through multi-modal fine-tuning of the trained large model in this embodiment, the large model can learn more extensive and complex interaction patterns, improve the generalization ability of the large model in multiple scenarios, and better handle diverse and complex problems that users may pose in actual conversations.
[0175] In the embodiment of the present application, an information processing model is trained using multi-modal information samples. The information processing model parallelly generates a text unit set and an audio unit set through a parallel generation strategy, and then converts the text unit set into text output information, and parallelly converts the audio unit set into audio output information, avoiding generating text first and then generating speech, reducing the interaction latency, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0176] The embodiment of the present application also provides a method for processing dialogue information, Figure 4 which is a flowchart of another method for processing dialogue information according to the embodiment of the present application, as Figure 4 shown. The method may include the following steps:
[0177] Step S402, obtaining voice input information.
[0178] In the technical solution provided in step S402 of the present application, voice input information can be obtained. Among them, the voice input information can be obtained through various means such as microphones, voice collection devices, voice input devices, etc. This is only for illustrative purposes and does not specifically limit the way of obtaining voice input information.
[0179] Step S404: Convert the voice input information into a text tensor and a voice tensor.
[0180] In the technical solution provided in step S404 of the present application, the text tensor and the voice tensor match the information processing model, and the information processing model is trained using multimodal information samples.
[0181] In this embodiment, after obtaining the voice input information, the obtained voice input information can be converted into a text tensor and a voice tensor.
[0182] Optionally, after obtaining the voice input information, the voice input information can be feature-extracted through an encoder to obtain audio features. Further, the audio features are converted into a text sequence through the Chinese-English automatic speech recognition (ASR) technology, and then the text sequence is converted into a text tensor that can be processed by the information processing model. The extracted audio features can be converted into an audio tensor with the same input dimension as the information processing model through an audio adapter.
[0183] Step S406: Input the text tensor and the voice tensor into the information processing model, and use the information processing model to analyze the text tensor and the voice tensor to generate a text unit set and a voice unit set in parallel.
[0184] In the technical solution provided in step S406 of the present application, the voice content corresponding to the voice unit set matches the text content corresponding to the text unit set.
[0185] In this embodiment, after converting the voice input information into a text tensor and a voice tensor, the text tensor and the voice tensor can be input into the information processing model, and the information processing model is used to analyze the text tensor and the voice tensor to generate a text unit set and a voice unit set in parallel.
[0186] For example, the information processing model can be used to analyze the text tensor and the audio tensor to generate 8-way tokens in parallel. Among them, the 8-way tokens can include 1-way text token and 7-way audio tokens. The 1-way text token can form the text unit set, and the 7-way audio tokens can form the audio unit set.
[0187] Step S408: Convert the text unit set into text output information that matches the voice input information, and convert the voice unit set into voice output information that matches the voice input information.
[0188] In the technical solution provided in step S408 of the present application, after analyzing the text tensor and the speech tensor by using the information processing model and generating the text unit set and the speech unit set in parallel, the text unit set can be decoded to obtain the text output information matching the speech input information, and the speech unit set can be decoded to obtain the speech output information matching the speech input information.
[0189] Optionally, after analyzing the text tensor and the audio tensor by using the information processing model and generating the text unit set and the audio unit set in parallel, the text decoder can be used to decode the text unit set to obtain the text output information matching the dialogue input information, and the audio decoder can be used to decode the audio unit set to obtain the audio output information matching the dialogue input information.
[0190] Through steps S402 to S408 of the present application, the speech input information is obtained; the speech input information is converted into a text tensor and a speech tensor, wherein the text tensor and the speech tensor match the information processing model, and the information processing model is trained by using multi-modal information samples; the text tensor and the speech tensor are input into the information processing model, and the information processing model is used to analyze the text tensor and the speech tensor and generate the text unit set and the speech unit set in parallel, wherein the speech content corresponding to the speech unit set matches the text content corresponding to the text unit set; the text unit set is converted into the text output information matching the speech input information, and the speech unit set is converted into the speech output information matching the speech input information. That is to say, after converting the speech input information into a text tensor and an audio tensor in the embodiment of the present application, through the parallel generation strategy, the information processing model is used to generate the text unit set and the audio unit set in parallel, and then the text unit set is converted into the text output information, and the audio unit set is converted into the audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and further achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0191] The embodiment of the present application also provides a method for processing dialogue information. Figure 5 It is a flowchart of another method for processing dialogue information according to the embodiment of the present application, as Figure 5 shown, and the method may include the following steps:
[0192] Step S502, obtain the dialogue query information for interacting with the virtual avatar.
[0193] In the technical solution provided in step S502 of the present application, dialogue query information for interacting with the virtual image can be obtained. Among them, the virtual image can be a digital human image created through three-dimensional (3D) modeling and rendering technology. The dialogue query information can be inquiries or instructions in oral or text form issued by the user during the interaction between the user and the virtual image, and the dialogue query information can include text input information and audio input information.
[0194] Step S504: Convert the dialogue query information into a text tensor and an audio tensor.
[0195] In the technical solution provided in step S504 of the present application, the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multi-modal information samples.
[0196] In this embodiment, after obtaining the dialogue query information for interacting with the virtual image, the obtained dialogue query information can be converted into a text tensor and an audio tensor.
[0197] Optionally, after obtaining the dialogue query information for interacting with the virtual image, the text input information in the dialogue query information can be directly converted into a text tensor with the same input dimension as the information processing model. The audio input information in the dialogue query information can be feature-extracted through an encoder, and further, the extracted audio features can be converted into an audio tensor with the same input dimension as the information processing model through an audio adapter.
[0198] Step S506: Input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor, and generate a text unit set and an audio unit set in parallel.
[0199] In the technical solution provided in step S506 of the present application, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set.
[0200] In this embodiment, after converting the dialogue query information into a text tensor and an audio tensor, the text tensor and the audio tensor can be input into the information processing model, and the information processing model is used to analyze the text tensor and the audio tensor, and a text unit set and an audio unit set are generated in parallel.
[0201] For example, the information processing model can be used to analyze the text tensor and the audio tensor, and 8-way tokens are generated in parallel. Among them, the 8-way tokens can include 1-way text token and 7-way audio tokens. The 1-way text token can form the text unit set, and the 7-way audio tokens can form the audio unit set.
[0202] Step S508: Convert the text unit set into text output information that matches the dialogue query information, and convert the audio unit set into audio output information that matches the dialogue query information.
[0203] In the technical solution provided in step S508 of the present application, after analyzing the text tensor and the audio tensor using the information processing model and generating the text unit set and the audio unit set in parallel, the text unit set can be decoded to obtain text output information that matches the dialogue query information, and the audio unit set can be decoded to obtain audio output information that matches the dialogue query information.
[0204] Optionally, after analyzing the text tensor and the audio tensor using the information processing model and generating the text unit set and the audio unit set in parallel, the text unit set can be decoded using a text decoder to obtain text output information that matches the dialogue input information, and the audio unit set can be decoded using an audio decoder to obtain audio output information that matches the dialogue input information.
[0205] Step S510: Control the virtual character to perform an interaction behavior according to the text output information and the audio output information, so as to output dialogue reply information that matches the dialogue query information.
[0206] In the technical solution provided in step S510 of the present application, after decoding the text unit set to obtain text output information that matches the dialogue query information, and decoding the audio unit set to obtain audio output information that matches the dialogue query information, the virtual character can be controlled to perform an interaction behavior according to the text output information and the audio output information, so as to output dialogue reply information that matches the dialogue query information.
[0207] For example, after obtaining the text output information and the audio output information, the virtual character can be controlled to perform an interaction behavior according to the text output information and the audio output information. For instance, according to the emotion or tone in the audio output information, the facial expression of the virtual character can be adjusted to enhance the emotional expression of the dialogue; if the text output information contains specific instructions or auxiliary actions are needed in the conversation of the virtual character, the limb actions of the virtual character will be adjusted accordingly, such as nodding, waving, or pointing in a certain direction; when playing the audio output information, the lip movement of the virtual character will be synchronized with the speech to achieve a more realistic and natural speech output performance.
[0208] Optionally, the virtual avatar combines the text output information and the audio output information through visual and auditory responses to form a unified dialogue reply information, including the playback of voice, the display of text, and corresponding limb and facial movements. This multimodal integrated output method provides users with a richer and more immersive interaction experience, making the virtual avatar's reply audible and visible, enhancing the authenticity and naturalness of the interaction.
[0209] Through the above steps S502 to S510 of the present application, obtain the dialogue query information for interacting with the virtual avatar; convert the dialogue query information into a text tensor and an audio tensor, where the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples; input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; convert the text unit set into text output information that matches the dialogue query information, and convert the audio unit set into audio output information that matches the dialogue query information; control the virtual avatar to execute an interaction behavior according to the text output information and the audio output information to output dialogue reply information that matches the dialogue query information. That is to say, after converting the dialogue query information for interacting with the virtual avatar into a text tensor and an audio tensor in the embodiments of the present application, through a parallel generation strategy, the information processing model is used to generate a text unit set and an audio unit set in parallel, and then the text unit set is converted into text output information, and the audio unit set is converted into audio output information in parallel, avoiding generating text first and then generating voice, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0210] The embodiments of the present application also provide a method for processing dialogue information, Figure 6 which is a flowchart of another method for processing dialogue information according to the embodiments of the present application, as Figure 6 shown, and the method may include the following steps:
[0211] Step S602, in response to an input operation instruction acting on the operation interface, input dialogue input information.
[0212] In the technical solution provided in step S602 of the present application above, the operation interface may be an operation interface displayed on a client device. For example, it may be an interface for interaction between a user and a computer or software system, and may be a graphical user interface, a command line interface, etc. This is only an example here, and no specific limitation is imposed on the form of the operation interface.
[0213] In this embodiment, the input operation instruction may be an instruction for controlling the operation interface to display dialogue input information. For example, it may be an instruction input by the user on the operation interface. The user can input text through the keyboard, or click on a button or option on the operation interface with the mouse, or generate the input operation instruction by dragging the dialogue input information to a specified area.
[0214] In an alternative embodiment, the user can open the operation interface deployed on the client device and input dialogue input information on the operation interface, thereby generating an input operation instruction. According to the generated input operation instruction, the corresponding dialogue input information can be loaded and displayed on the operation interface.
[0215] Step S604: Respond to the reply operation instruction acting on the operation interface, and output text output information matching the dialogue input information and audio output information matching the dialogue input information.
[0216] In the technical solution provided in step S604 of the present application above, the text output information is obtained by converting a text unit set, and the audio output information is obtained by converting an audio unit set. The text unit set and the audio unit set are generated in parallel by analyzing text tensors and audio tensors using an information processing model. The audio content corresponding to the audio unit set matches the text content corresponding to the text unit set. The text tensors and the audio tensors are obtained by converting the dialogue input information, and the text tensors and the audio tensors match the information processing model. The information processing model is trained using multimodal information samples.
[0217] In this embodiment, the reply operation instruction may be an instruction for replying to the dialogue input information displayed on the operation interface. The user can generate the reply operation instruction by clicking on the dialogue input information displayed on the operation interface. This is only an example here and does not specifically limit the reply operation instruction.
[0218] Optionally, after responding to the input operation instruction acting on the operation interface and inputting the dialogue input information, the text input information in the dialogue input information can be directly converted into a text tensor having the same input dimension as the information processing model. Feature extraction can be performed on the audio input information in the dialogue input information through an encoder, and then the extracted audio features can be further converted into an audio tensor having the same input dimension as the information processing model through an audio adapter.
[0219] For example, after converting the dialogue input information into text tensors and audio tensors, an information processing model can be used to analyze the text tensors and audio tensors, and 8 tokens can be generated in parallel. Among them, the 8 tokens can include 1 text token and 7 audio tokens. The 1 text token can form a text unit set, and the 7 audio tokens can form an audio unit set.
[0220] Optionally, after using the information processing model to analyze the text tensors and audio tensors and generating the text unit set and the audio unit set in parallel, a text decoder can be used to decode the text unit set to obtain text output information that matches the dialogue input information, and an audio decoder can be used to decode the audio unit set to obtain audio output information that matches the dialogue input information. In response to a reply operation instruction acting on the operation interface, text output information that matches the dialogue input information and audio output information that matches the dialogue input information can be output.
[0221] Through the above steps S602 to S604 of the present application, in response to an input operation instruction acting on the operation interface, dialogue input information is input; in response to a reply operation instruction acting on the operation interface, text output information that matches the dialogue input information and audio output information that matches the dialogue input information are output. Among them, the text output information is obtained by converting the text unit set, the audio output information is obtained by converting the audio unit set, the text unit set and the audio unit set are generated in parallel by using an information processing model to analyze the text tensors and audio tensors, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, the text tensors and audio tensors are obtained by converting the dialogue input information, and the text tensors and audio tensors match the information processing model, and the information processing model is trained by using multimodal information samples. That is to say, in the embodiment of the present application, in response to a reply operation instruction acting on the operation interface, text output information that matches the dialogue input information and audio output information that matches the dialogue input information are output. Since the text output information and the audio output information are obtained by converting the text unit set and the audio unit set in parallel, and the text unit set and the audio unit set are generated in parallel by using an information processing model through a parallel generation strategy, it avoids generating text first and then generating speech, reduces the interaction delay, and thus achieves the technical effect of improving the processing efficiency of dialogue information and solves the technical problem of low processing efficiency of dialogue information.
[0222] The embodiment of the present application also provides a dialogue information processing system. Figure 7 It is a schematic diagram of a dialogue information processing system according to an embodiment of the present application, as Figure 7As shown, the processing system for the conversation information is deployed on the mobile device side, and the processing system for the conversation information may include: an information input end 702, an information processing end 704, and an information output end 706.
[0223] The information input end 702 is used to obtain conversation input information.
[0224] The information processing end 704 is used to convert the conversation input information into a text tensor and an audio tensor, where the text tensor and the audio tensor match the information processing model, and the information processing model is obtained by training using multimodal information samples; input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor, and generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; convert the text unit set into text output information that matches the conversation input information, and convert the audio unit set into audio output information that matches the conversation input information.
[0225] Optionally, after the information input end 702 obtains the conversation input information, the obtained conversation input information can be transmitted to the information processing end 704. After receiving the conversation input information, the information processing end 704 can directly convert the text input information in the conversation input information into a text tensor with the same input dimension as the information processing model, and can also perform feature extraction on the audio input information in the conversation input information through an encoder, and further convert the extracted audio features into an audio tensor with the same input dimension as the information processing model through an audio adapter.
[0226] For example, after the information processing end 704 converts the conversation input information into a text tensor and an audio tensor, the information processing model can be used to analyze the text tensor and the audio tensor, and 8-way tokens can be generated in parallel, where the 8-way tokens can include 1-way text token and 7-way audio tokens, and the 1-way text token can form the text unit set, and the 7-way audio tokens can form the audio unit set.
[0227] Optionally, after the information processing end 704 uses the information processing model to analyze the text tensor and the audio tensor and generates a text unit set and an audio unit set in parallel, a text decoder can be used to decode the text unit set to obtain text output information that matches the conversation input information, and an audio decoder can be used to decode the audio unit set to obtain audio output information that matches the conversation input information.
[0228] The information output end 706 is used to output the text output information and the audio output information.
[0229] In the processing system of the dialogue information, the dialogue input information is obtained through the information input end 702. The dialogue input information is converted into a text tensor and an audio tensor through the information processing end 704, wherein the text tensor and the audio tensor match the information processing model, and the information processing model is obtained by training with multimodal information samples; the text tensor and the audio tensor are input into the information processing model, and the information processing model is used to analyze the text tensor and the audio tensor, and a text unit set and an audio unit set are generated in parallel, wherein the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; the text unit set is converted into text output information that matches the dialogue input information, and the audio unit set is converted into audio output information that matches the dialogue input information. The text output information and the audio output information are output through the information output end 706. That is to say, after converting the dialogue input information into a text tensor and an audio tensor through the information processing end 704, the processing system of the dialogue information uses the information processing model to generate the text unit set and the audio unit set in parallel through a parallel generation strategy, and then converts the text unit set into text output information, and converts the audio unit set into audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0230] The technical solutions of the embodiments of the present disclosure will be further introduced by way of preferred embodiments below.
[0231] At present, large language models have made remarkable progress in text understanding and generation, but there are deficiencies in real-time speech interaction. Existing large models usually rely on additional TTS systems for speech synthesis, resulting in high latency and affecting the user experience. Although some large models have the ability of real-time multimodal interaction, the technical details are not disclosed, and the closed source limits their wide application.
[0232] In an optional example, the end-to-end speech interaction model needs to generate text first and then generate speech, resulting in a latency problem; some models support speech input but output text, relying on an external TTS system for speech synthesis, with high latency; some models support speech input and text output, also relying on an external TTS system. The above solutions have deficiencies in aspects such as the real-time performance and natural fluency of speech interaction, and relying on external components leads to high latency.
[0233] To solve the above problems, this application proposes a Speech to Speech dialogue system running locally on the device side, which can conduct real-time voice conversations with users. The advantages of this dialogue system are as follows: all modules run on the device side without the need to connect to the network; it adopts voice interaction with a good experience; it has a fast response with a latency within 1 second; the conversation can be in Chinese or English; it can reply with voice and text simultaneously.
[0234] In this embodiment, the Speech to Speech dialogue system includes an audio decoder, a text generator, and an audio adapter, and realizes real-time voice interaction through a parallel generation strategy; a three-stage training method is adopted, including modality alignment, adaptation training, and multi-modal fine-tuning, to ensure that the model has good voice interaction capabilities while maintaining text capabilities. Figure 8 It is a schematic diagram of a Speech to Speech dialogue system running locally on a device side according to an embodiment of the present application, as Figure 8 shown Figure 8 The encoding (encode) part of the model is at the lower left, responsible for converting the input text (for example, the text "Who are you" shown in the figure) and audio into a form that can be received by the large model. Figure 8 The large model base is in the middle, responsible for inference. Figure 8 The decoding (decode) part of the model is at the upper right, responsible for receiving the output of the large model and converting the model output into text and speech tokens for output.
[0235] As Figure 8 shown, the encode part is divided into text and speech. The voice file is extracted into an 868-dimensional vector through an audio encoder, and then converted into an 896-dimensional vector through an adapter. The text is converted into text tokens through a text encoder, and then converted into an 896-dimensional vector through an embedding layer. The text and audio vectors are superimposed (merged embedding) to form the final vector and input to the large model for inference. After the large model infers, the decode part performs token sampling through a sampling layer. The sampling schematic diagram is as Figure 8 shown in the upper left corner. One text token and seven audio tokens will be sampled from the same set of outputs. The text token will generate text (for example, the text "I am XXX" shown in the figure) through a text decoder, and the audio token will generate voice through an audio decoder. Similar to the inference of the large model later, after generating the first column of outputs, subsequent token lists will be generated cyclically, and then serialized text (sentences) and audio (utterances) will be generated. Figure 8 The lower right shows the schematic diagram of the token arrangement during the training process and the inference process. Squares of different colors represent different tokens, which play different roles in training and inference.
[0236] The above method of this embodiment will be further introduced from aspects such as data preparation, model structure, training process, evaluation system, and inference process.
[0237] First, data preparation can include data collection, data preprocessing, data synthesis, data annotation, data partitioning, and data adaptation, etc. For data collection, the Speech to Speech dialogue system uses multiple datasets for training, and the multiple datasets cover various tasks such as Chinese and English Automatic Speech Recognition (ASR), Text Question Answering (Text QA), Audio Question Answering (Audio QA), etc. For data preprocessing, for speech data, a pre-trained model can be used for feature extraction and audio encoding. For text data, the text data can be cleaned and formatted to ensure that the text data is suitable for model training. For data synthesis, a speech synthesis model can be used to synthesize a dedicated voice assistant dataset. Through zero-shot Text-to-Speech (TTS) technology, the text data is converted into Chinese and English voice dialogue data to increase the diversity and quantity of training data. For data annotation, for some tasks, such as Audio QA, accurate text transcription and annotation of speech data are required. For data partitioning, the dataset can be divided into a training set, a validation set, and a test set for model training, parameter tuning, and performance evaluation. For data adaptation, to adapt to the input format of the model, the text and audio data can be converted into a tensor form that the model can process. For audio input, the audio features and input tokens are converted into tensors of the same dimension through an adapter and then concatenated.
[0238] Secondly, the model structure can include a front-end encoder, a middle large model, a back-end decoder, etc. Among them, the front-end encoder part of the model can include a language modeling component, an audio adapter, and an embedding layer. The language modeling component is responsible for generating text and audio tokens, the audio adapter is used to process audio input and output, and the embedding layer is used to convert the input data into a representation that the model can process. The middle large model part of the model can include a pre-trained text large model. The model can simultaneously process the probability distributions of text and speech tokens and generate corresponding outputs. The back-end decoder part of the model can include a text decoder and an audio decoder. These two decoders receive text tokens and audio tokens respectively and decode them to obtain text responses and voice responses.
[0239] Furthermore, the training process may include modality alignment, adapter training, and multi-modal fine-tuning. Among them, modality alignment is trained using speech recognition and speech synthesis data, freezing the core model of the dialogue system and only allowing gradient updates of the two adapters, so as to enhance the text model's ability to understand and generate speech. Adapter training is trained using data from speech recognition, spoken language question answering, and text response tasks. After aligning the new modality with the input of the text model, the adapter is frozen. In this stage, only the text ability of the model is trained, and the audio output is synthesized through text, so as to train the text generation ability of the model given audio input. Multi-modal fine-tuning is trained using comprehensive data containing the speech modality. In this stage, the model weights are unfrozen for comprehensive training, so as to fine-tune the entire model using the comprehensive data.
[0240] Furthermore, the evaluation system can evaluate aspects such as the automatic speech recognition (ASR) ability, text and speech question answering (QA) ability, text-to-speech synthesis (TTS) ability, multi-modal interaction ability, inference speed and efficiency, and model generalization ability. Among them, the evaluation method for the speech recognition ability is to use the ASR test set to evaluate the speech recognition accuracy of the model, and the evaluation metric is the word error rate (WER). The evaluation method for the text and speech question answering ability is to use datasets such as large-scale multi-modal dialogue datasets (Open-Orca) for text and speech question answering tests, and the evaluation metrics are the accuracy and relevance of the answers. The evaluation method for the text-to-speech synthesis ability is to evaluate through the naturalness and intelligibility of the synthesized speech, and the evaluation metrics are naturalness, i.e., the mean opinion score (MOS) and intelligibility. The evaluation method for the multi-modal interaction ability is to use comprehensive data containing speech and text for testing, and the evaluation metric is the performance of the model in multi-modal tasks, such as the joint understanding and generation ability of speech and text. The evaluation method for the inference speed and efficiency is to test the response speed of the model in real-time interaction, and the evaluation metrics are the first-word latency and audio generation speed. The evaluation method for the model generalization ability is to use unseen datasets for testing, and the evaluation metric is the performance of the model under different tasks and different data distributions.
[0241] Finally, the inference process can include multiple processes such as input processing, parallel generation, batch parallel decoding, and output. For input processing, text input can directly enter the model. Audio input first undergoes feature extraction through an encoder, then is converted into a tensor with the same dimension as the model input through an adapter, and is concatenated with the tensor converted from the text input to form a tensor dimension that the large model can receive. For parallel generation, in text-guided audio generation, that is, the model generates text and audio tokens simultaneously. The audio tokens are generated through text-to-speech synthesis to ensure real-time output. To accelerate audio generation, a parallel decoding strategy is adopted to generate multiple audio token layers simultaneously. The audio tokens are generated layer by layer from the first layer to the seventh layer, and the text tokens are output first. For batch parallel decoding, a single input is extended to batch processing. One sample generates text and audio responses, and another sample generates only text responses. The text token output is discarded from the first sample, and the text output from the second sample is embedded into the corresponding text token position of the first sample. At the same time, the audio of the first sample is streamed out through the text response content of the second sample. Finally, the text output is the sequence of directly generated text tokens, and the audio output is the audio stream output obtained by converting the generated audio tokens through text-to-speech synthesis.
[0242] It should be noted that if an independent text-to-speech system is used, end-to-end voice interaction cannot be achieved, and the latency is relatively high. If speech recognition, text generation, and text-to-speech synthesis are connected in series as independent modules, the latency problem still exists.
[0243] Therefore, this embodiment adopts a parallel generation strategy to reduce latency and improve real-time performance by generating text and audio in parallel; adopts text-guided latency parallel decoding to ensure high information density of text output, and the audio output is achieved through text synthesis to improve voice quality; adopts batch parallel decoding to further enhance the inference speed and model performance through batch processing; supports Chinese-English conversations, and by increasing Chinese daily conversation data, a model that can communicate in Chinese and English is realized, thus achieving low-latency, high-quality real-time voice interaction while maintaining the original text processing ability of the model.
[0244] This embodiment supports parallel text and audio generation; generates text and audio simultaneously to reduce latency and improve real-time interaction ability; supports text-guided latency parallel decoding to improve the quality of voice output through text-guided audio generation; supports batch parallel decoding to further enhance the inference speed and model performance through batch processing; the model training method quickly adds Chinese-English voice interaction ability to the model through a small amount of Chinese-English data.
[0245] In an embodiment of the present application, after converting the dialogue input information into a text tensor and an audio tensor, through a parallel generation strategy, an information processing model is used to parallelly generate a set of text units and a set of audio units, and then the set of text units is converted into text output information, and the set of audio units is parallelly converted into audio output information, avoiding generating text first and then generating speech, reducing the interaction latency, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0246] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0247] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0248] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0249] According to an embodiment of the present application, there is also provided a dialogue information processing device for implementing the above Figure 2 processing method of the shown dialogue information.
[0250] Figure 9 is a schematic diagram of a dialogue information processing device according to an embodiment of the present application, as Figure 9As shown in the figure, the processing device 900 for the dialogue information may include: a first acquisition unit 902, a first conversion unit 904, a first processing unit 906, and a second conversion unit 908.
[0251] The first acquisition unit 902 is configured to acquire dialogue input information.
[0252] The first conversion unit 904 is configured to convert the dialogue input information into a text tensor and an audio tensor, where the text tensor and the audio tensor are matched with an information processing model, and the information processing model is obtained by training using multimodal information samples.
[0253] The first processing unit 906 is configured to input the text tensor and the audio tensor into the information processing model, and analyze the text tensor and the audio tensor by using the information processing model to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set is matched with the text content corresponding to the text unit set.
[0254] The second conversion unit 908 is configured to convert the text unit set into text output information that is matched with the dialogue input information, and convert the audio unit set into audio output information that is matched with the dialogue input information.
[0255] It should be noted here that the above first acquisition unit 902, first conversion unit 904, first processing unit 906, and second conversion unit 908 correspond to steps S202 to S208 in the above embodiments. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also be part of a device and may run in the client device provided in the above embodiments.
[0256] In the processing device for the dialogue information, the dialogue input information is obtained by the first obtaining unit 902. The dialogue input information is converted into a text tensor and an audio tensor by the first conversion unit 904, where the text tensor and the audio tensor are matched with the information processing model, and the information processing model is trained using multimodal information samples. The text tensor and the audio tensor are input into the information processing model by the first processing unit 906, and the information processing model is used to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set is matched with the text content corresponding to the text unit set. The text unit set is converted into text output information matched with the dialogue input information, and the audio unit set is converted into audio output information matched with the dialogue input information by the second conversion unit 908. That is to say, after the dialogue input information is converted into a text tensor and an audio tensor, the processing device for the dialogue information uses the information processing model to generate a text unit set and an audio unit set in parallel through a parallel generation strategy, and then converts the text unit set into text output information and converts the audio unit set into audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of the dialogue information and solving the technical problem of low processing efficiency of the dialogue information.
[0257] According to an embodiment of the present application, there is also provided a model determination device for implementing the determination method of the model as described above. Figure 3 The model determination device for implementing the determination method of the model as shown above.
[0258] Figure 10 It is a schematic diagram of a model determination device according to an embodiment of the present application. As Figure 10 shown, the model determination device 1000 may include: a second obtaining unit 1002 and a training unit 1004.
[0259] The second obtaining unit 1002 is configured to obtain multimodal information samples.
[0260] The training unit 1004 is configured to train an information processing model using multimodal information samples, where the information processing model is used to analyze the text tensor and the audio tensor corresponding to the dialogue input information, generate a text unit set and an audio unit set in parallel, the text tensor and the audio tensor are matched with the information processing model, the audio content corresponding to the audio unit set is matched with the text content corresponding to the text unit set, and the text unit set is used to be converted into text output information matched with the dialogue input information, and the audio unit set is used to be converted into audio output information matched with the dialogue input information.
[0261] It should be noted here that the above-mentioned second acquisition unit 1002 and training unit 1004 correspond to steps S302 to S304 in the above-mentioned embodiments. The functions of the two modules are the same as those of the corresponding steps in terms of the implemented examples and application scenarios, but are not limited to the content disclosed in the above-mentioned embodiments. It should be noted that the above-mentioned module or unit can be a hardware component or software component stored in a memory and processed by one or more processors, and the above-mentioned module can also be part of a device and can run in the client device provided in the above-mentioned embodiments.
[0262] In the device for determining the model, the second acquisition unit 1002 acquires multimodal information samples. The training unit 1004 uses the multimodal information samples to train an information processing model. The information processing model is used to analyze the text tensor and audio tensor corresponding to the dialogue input information, and generate a text unit set and an audio unit set in parallel. The text tensor and audio tensor match the information processing model, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, and the text unit set is used to be converted into text output information matching the dialogue input information, and the audio unit set is used to be converted into audio output information matching the dialogue input information. That is to say, the device for determining the model uses multimodal information samples to train an information processing model, and uses the information processing model to generate a text unit set and an audio unit set in parallel through a parallel generation strategy, and then converts the text unit set into text output information, and converts the audio unit set into audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of dialogue information and solving the technical problem of low processing efficiency of dialogue information.
[0263] According to an embodiment of the present application, there is also provided a dialogue information processing device for implementing the Figure 4 dialogue information processing method shown above.
[0264] Figure 11 FIG. is a schematic diagram of another dialogue information processing device according to an embodiment of the present application. As Figure 11 shown, the dialogue information processing device 1100 may include: a third acquisition unit 1102, a third conversion unit 1104, a second processing unit 1106, and a fourth conversion unit 1108.
[0265] The third acquisition unit 1102 is configured to acquire voice input information.
[0266] The third conversion unit 1104 is configured to convert the voice input information into a text tensor and a voice tensor, where the text tensor and the voice tensor match the information processing model, and the information processing model is trained using multimodal information samples.
[0267] A second processing unit 1106, configured to input the text tensor and the speech tensor into an information processing model, analyze the text tensor and the speech tensor by using the information processing model, and generate a text unit set and a speech unit set in parallel, where the speech content corresponding to the speech unit set matches the text content corresponding to the text unit set.
[0268] A fourth conversion unit 1108, configured to convert the text unit set into text output information that matches the speech input information, and convert the speech unit set into speech output information that matches the speech input information.
[0269] It should be noted here that the above third acquisition unit 1102, third conversion unit 1104, second processing unit 1106, and fourth conversion unit 1108 correspond to steps S402 to S408 in the above embodiment. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also be part of a device and may run in the client device provided in the above embodiment.
[0270] In the processing device for the dialogue information, the speech input information is acquired by the third acquisition unit 1102. The speech input information is converted into a text tensor and a speech tensor by the third conversion unit 1104, where the text tensor and the speech tensor match the information processing model, and the information processing model is obtained by training with multimodal information samples. The text tensor and the speech tensor are input into the information processing model by the second processing unit 1106, and the information processing model is used to analyze the text tensor and the speech tensor to generate a text unit set and a speech unit set in parallel, where the speech content corresponding to the speech unit set matches the text content corresponding to the text unit set. The text unit set is converted into text output information that matches the speech input information by the fourth conversion unit 1108, and the speech unit set is converted into speech output information that matches the speech input information. That is to say, after converting the speech input information into a text tensor and an audio tensor, the processing device for the dialogue information uses a parallel generation strategy to generate a text unit set and an audio unit set in parallel by using the information processing model, and then converts the text unit set into text output information, and converts the audio unit set into audio output information in parallel, avoiding generating text first and then generating speech, reducing the interaction delay, and thus achieving the technical effect of improving the processing efficiency of the dialogue information and solving the technical problem of low processing efficiency of the dialogue information.
[0271] According to an embodiment of the present application, there is also provided a processing device for dialogue information for implementing the Figure 5 processing method for dialogue information shown above.
[0272] Figure 12 It is a schematic diagram of another dialogue information processing device according to an embodiment of the present application. As Figure 12 shown, the dialogue information processing device 1200 may include: a fourth acquisition unit 1202, a fifth conversion unit 1204, a third processing unit 1206, a sixth conversion unit 1208, and a control unit 1210.
[0273] The fourth acquisition unit 1202 is configured to acquire dialogue query information for interacting with a virtual avatar.
[0274] The fifth conversion unit 1204 is configured to convert the dialogue query information into a text tensor and an audio tensor, where the text tensor and the audio tensor match an information processing model, and the information processing model is trained using multimodal information samples.
[0275] The third processing unit 1206 is configured to input the text tensor and the audio tensor into the information processing model, analyze the text tensor and the audio tensor using the information processing model, and generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set.
[0276] The sixth conversion unit 1208 is configured to convert the text unit set into text output information that matches the dialogue query information, and convert the audio unit set into audio output information that matches the dialogue query information.
[0277] The control unit 1210 is configured to control the virtual avatar to execute an interaction behavior according to the text output information and the audio output information, so as to output dialogue reply information that matches the dialogue query information.
[0278] It should be noted here that the above-mentioned fourth acquisition unit 1202, fifth conversion unit 1204, third processing unit 1206, sixth conversion unit 1208, and control unit 1210 correspond to steps S502 to S510 in the above embodiment. The five modules have the same implemented examples and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also be part of a device and can run in the client device provided in the above embodiment.
[0279] In the processing device for the dialogue information, the fourth acquisition unit 1202 acquires the dialogue query information for interacting with the virtual avatar. The fifth conversion unit 1204 converts the dialogue query information into a text tensor and an audio tensor, wherein the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples. The third processing unit 1206 inputs the text tensor and the audio tensor into the information processing model, and uses the information processing model to analyze the text tensor and the audio tensor, and parallelly generates a text unit set and an audio unit set, wherein the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set. The sixth conversion unit 1208 converts the text unit set into text output information matching the dialogue query information, and converts the audio unit set into audio output information matching the dialogue query information. The control unit 1210 controls the virtual avatar to execute an interaction behavior according to the text output information and the audio output information, so as to output dialogue reply information matching the dialogue query information. That is to say, after converting the dialogue query information for interacting with the virtual avatar into a text tensor and an audio tensor, the processing device for the dialogue information parallelly generates a text unit set and an audio unit set by using the information processing model through a parallel generation strategy, and then converts the text unit set into text output information, and parallelly converts the audio unit set into audio output information, avoiding generating text first and then generating speech, reducing the interaction latency, and thus achieving the technical effect of improving the processing efficiency of the dialogue information and solving the technical problem of low processing efficiency of the dialogue information.
[0280] According to an embodiment of the present application, there is also provided a processing device for dialogue information for implementing the Figure 6 processing method for dialogue information shown above.
[0281] Figure 13 It is a schematic diagram of another processing device for dialogue information according to an embodiment of the present application. As Figure 13 shown, the processing device 1200 for dialogue information may include: an input unit 1302 and an output unit 1304.
[0282] The input unit 1302 is configured to input dialogue input information in response to an input operation instruction acting on the operation interface.
[0283] An output unit 1304, configured to output text output information matching the dialogue input information and audio output information matching the dialogue input information in response to a reply operation instruction acting on the operation interface, where the text output information is obtained by converting a text unit set, the audio output information is obtained by converting an audio unit set, the text unit set and the audio unit set are generated in parallel by analyzing a text tensor and an audio tensor using an information processing model, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, the text tensor and the audio tensor are obtained by converting the dialogue input information, and the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples.
[0284] It should be noted here that the above input unit 1302 and output unit 1304 correspond to steps S602 to S604 in the above embodiment. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also be part of a device and may run in the client device provided in the above embodiment.
[0285] In the processing device for the dialogue information, the input unit 1302 responds to an input operation instruction acting on the operation interface and inputs the dialogue input information. The output unit 1304 responds to a reply operation instruction acting on the operation interface and outputs text output information matching the dialogue input information and audio output information matching the dialogue input information, where the text output information is obtained by converting a text unit set, the audio output information is obtained by converting an audio unit set, the text unit set and the audio unit set are generated in parallel by analyzing a text tensor and an audio tensor using an information processing model, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, the text tensor and the audio tensor are obtained by converting the dialogue input information, and the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples. That is to say, the processing device for the dialogue information responds to a reply operation instruction acting on the operation interface and outputs text output information matching the dialogue input information and audio output information matching the dialogue input information. Since the text output information and the audio output information are obtained by converting the text unit set and the audio unit set, and the text unit set and the audio unit set are generated in parallel by using an information processing model through a parallel generation strategy, it avoids generating text first and then generating speech, reduces the interaction delay, and thus achieves the technical effect of improving the processing efficiency of the dialogue information and solves the technical problem of low processing efficiency of the dialogue information.
[0286] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in the above embodiments, but are not limited to the schemes provided in the above embodiments.
[0287] Embodiments of the present application can provide a computing device. Figure 14 It is a structural block diagram of a computing device according to an embodiment of the present application, as Figure 14 shown. The computing device 1400 may include: one or more (only one is shown in the figure) processors 1402, a memory 1404, a storage controller, and a peripheral interface.
[0288] The above computing device can be understood as an integrated intelligent terminal, including but not limited to servers, desktop computers, personal computers (abbreviated as PC), model all-in-ones, etc. Moreover, the above-mentioned models in the embodiments of the present application can be pre-installed in the computing device.
[0289] Specifically, the computing device can pre-install various types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi-modal task processing, etc., so as to provide diverse model selections. In different product forms, the computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference and application, etc. In some product forms, the computing device also supports model management, including but not limited to multi-type model management (supporting the management of various types of models such as discriminative and generative), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effect of the model based on model evaluation tools), etc. In other product forms, the computing device can also create applications based on the model, provide API invocation capabilities, and can call the model into the created application through the API interface, while providing application management tools to achieve the control of the application.
[0290] Furthermore, the computing device can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master AI technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.
[0291] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal A through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0292] The processor can call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0293] Embodiments of the present application can provide an electronic device. Figure 15 is a structural block diagram of an electronic device according to an embodiment of the present application, as Figure 15 shown. The electronic device can include: an input / output device 1502; a memory 1504 and a processor 1506, where the processor 1506 is connected to the input / output device 1502 and the memory 1504 through a bus 1508.
[0294] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal A through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0295] The processor can call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0296] Those of ordinary skill in the art can understand that the structure shown in the figure is only illustrative, and the computing device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a personal digital assistant, and mobile Internet devices (abbreviated as MID), a PAD and other terminal devices. The figure does not limit the structure of the above computing device. For example, the computing device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.
[0297] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disc, etc.
[0298] An embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium can be used to store the program code executed by the method provided in the above embodiment.
[0299] Optionally, in this embodiment, the above storage medium may be located in the computing device.
[0300] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method described in any one of the above embodiments.
[0301] An embodiment of the present application also provides a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and the above computer program implements the method provided in the above embodiment when executed by a processor.
[0302] An embodiment of the present application also provides a computer program product. Optionally, the above computer program product may include a non-volatile computer-readable storage medium, and the above non-volatile computer-readable storage medium can be used to store a computer program, and the above computer program implements the method provided in the above embodiment when executed by a processor.
[0303] An embodiment of the present application also provides a computer program. Optionally, in this embodiment, the above computer program implements the method provided in the above embodiment when executed by a processor.
[0304] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0305] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0306] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0307] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0308] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, read-only memory ROM, random access memory RAM, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0309] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for processing dialogue information, characterized in that, Including: Obtain dialogue input information; Convert the dialogue input information into a text tensor and an audio tensor, where the text tensor and the audio tensor match an information processing model, and the information processing model is trained using multimodal information samples; Input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; Convert the text unit set into text output information that matches the dialogue input information, and convert the audio unit set into audio output information that matches the dialogue input information.
2. The method according to claim 1, wherein Input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, including: Concatenate the text tensor and the audio tensor to obtain a target tensor, where the target tensor matches the information processing model; Input the target tensor into the information processing model, and use the information processing model to analyze the target tensor to generate the text unit set and the audio unit set in parallel.
3. The method according to claim 2, wherein Use the information processing model to analyze the target tensor to generate the text unit set and the audio unit set in parallel, including: In the process of sampling text units from the target tensor using the language modeling component of the information processing model, use the obtained text units to determine the audio parameters of the audio units to be generated, where the audio parameters are used to characterize the quality degree of the audio units to be generated; In the process of sampling text units other than the obtained text units from the target tensor using the language modeling component, generate the audio units that meet the audio parameters to obtain the audio unit set, and generate the text unit set from the collected text units.
4. The method according to claim 3, wherein Generate the audio units that meet the audio parameters to obtain the audio unit set, including: Generate multiple layers of the audio units that meet the audio parameters in sequence to obtain the audio unit set, where the multiple layers of the audio units are in a parallel relationship.
5. The method according to claim 1, characterized in that, Convert the text unit set into text output information that matches the dialogue input information, and convert the audio unit set into audio output information that matches the dialogue input information, including: Perform parallel decoding on the text unit set and the audio unit set to obtain the text output information corresponding to the text unit set and the audio output information corresponding to the audio unit set.
6. The method according to claim 5, wherein The text output information includes a text unit sequence, and the audio output information includes an audio stream. Performing parallel decoding on the text unit set and the audio unit set to obtain the text output information corresponding to the text unit set and the audio output information corresponding to the audio unit set, including: The text decoder of the information processing model is used to decode the text unit set to obtain the text unit sequence corresponding to the text unit set. At the same time, the audio decoder of the information processing model is used to decode the audio unit set to obtain the audio stream corresponding to the audio unit set.
7. The method according to claim 1, wherein The method further includes: Determining the target data dimension of the input data of the information processing model; Converting the dialogue input information into a text tensor and an audio tensor, including: using the text encoder of the information processing model to extract initial text units from the dialogue input information, and in the embedding layer of the information processing model, converting the initial text units into text tensors of the target data dimension; using the audio encoder of the information processing model to extract initial audio features from the dialogue input information, and using the audio adapter of the information processing model to convert the initial audio features into audio tensors of the target data dimension.
8. The method according to claim 1, characterized in that, The dialogue input information at least includes first dialogue input information and second dialogue input information. Converting the dialogue input information into a text tensor and an audio tensor includes: Converting the first dialogue input information into a first text tensor and an audio tensor, and converting the second dialogue input information into a second text tensor, where the text quality of the second text tensor is higher than that of the first text tensor.
9. The method according to claim 8, wherein Using the information processing model to analyze the text tensor and the audio tensor, and generating a text unit set and an audio unit set in parallel, including: Using the information processing model to analyze the first text tensor, the audio tensor and the second text tensor, and generating a first text unit set and an audio unit set corresponding to the first dialogue input information, and a second text unit set corresponding to the second dialogue input information in parallel.
10. The method according to claim 9, characterized in that, The text quality of the second text unit set is higher than that of the first text unit set. The method further includes: Replacing the first text unit set with the second text unit set.
11. The method according to claim 8, wherein Converting the first dialogue input information into a first text tensor and an audio tensor includes: Converting the text input information in the first dialogue input information into the first text tensor, and converting the audio input information in the first dialogue input information into the audio tensor; Converting the second dialogue input information into a second text tensor includes: converting the text input information in the second dialogue input information into the second text tensor.
12. The method according to any one of claims 1 to 11, characterized in that, Obtaining dialogue input information includes: Obtaining dialogue input information in different languages.
13. A method for determining a model, characterized in that, Including: Obtaining multi-modal information samples; Using multimodal information samples, an information processing model is trained. Among them, the information processing model is used to analyze the text tensor and audio tensor corresponding to the dialogue input information, and parallelly generate a text unit set and an audio unit set. The text tensor and the audio tensor match the information processing model. The audio content corresponding to the audio unit set matches the text content corresponding to the text unit set. And the text unit set is used to be converted into text output information matching the dialogue input information, and the audio unit set is used to be converted into audio output information matching the dialogue input information.
14. The method according to claim 13, characterized in that, The multimodal information samples include text dialogue information samples and audio dialogue information samples. Obtaining multimodal information samples includes: Determining different dialogue tasks; Obtaining the text dialogue information samples and the audio dialogue information samples under the dialogue task.
15. The method according to claim 14, characterized in that, Obtaining the text dialogue information samples and the audio dialogue information samples under the dialogue task includes: Obtaining the text dialogue information samples in different languages and the audio dialogue information samples in different languages under the dialogue task.
16. The method according to claim 14, characterized in that, Using multimodal information samples to train an information processing model includes: Performing modal alignment on the text dialogue information samples and the audio dialogue information samples; Using the aligned text dialogue information samples and the aligned audio dialogue information samples to train an initial model corresponding to the information processing model to obtain the information processing model.
17. The method according to claim 16, characterized in that Performing modal alignment on the text dialogue information samples and the audio dialogue information samples includes: Converting the text dialogue information samples into text tensor samples with a target data dimension matching the initial model; Extracting initial audio feature samples from the audio dialogue information samples and converting the initial audio feature samples into audio tensor samples with the target data dimension.
18. The method according to claim 17, characterized in that Using the aligned text dialogue information samples and the aligned audio dialogue information samples to train the initial model to obtain the information processing model includes: Concatenating the text tensor samples and the audio tensor samples to obtain target tensor samples, where the target tensor samples match the initial model; Using the target tensor samples to train the initial model to obtain the information processing model.
19. The method according to claim 16, wherein Using the aligned text dialogue information samples and the aligned audio dialogue information samples to train the initial model to obtain the information processing model includes: Using the aligned audio dialogue information samples and the aligned audio dialogue information samples to train the text processing model of the initial model to obtain the information processing model.
20. The method according to claim 16, wherein Using the aligned text dialogue information samples and the aligned audio dialogue information samples to train the initial model to obtain the information processing model includes: Using the aligned text dialogue information samples and the aligned audio dialogue information samples to train the initial model; Training the trained initial model using target data samples that at least include voice data samples to obtain the information processing model.
21. A method for processing conversation information, characterized in that, Including: Obtaining voice input information; Converting the voice input information into a text tensor and a voice tensor, where the text tensor and the voice tensor match the information processing model, and the information processing model is trained using multimodal information samples; Inputting the text tensor and the voice tensor into the information processing model, and using the information processing model to analyze the text tensor and the voice tensor to generate a text unit set and a voice unit set in parallel, where the voice content corresponding to the voice unit set matches the text content corresponding to the text unit set; Converting the text unit set into text output information that matches the voice input information, and converting the voice unit set into voice output information that matches the voice input information.
22. A method for processing conversation information, characterized in that, Including: Obtaining dialogue query information for interacting with the virtual avatar; Converting the dialogue query information into a text tensor and an audio tensor, where the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples; Inputting the text tensor and the audio tensor into the information processing model, and using the information processing model to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, where the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set; Converting the text unit set into text output information that matches the dialogue query information, and converting the audio unit set into audio output information that matches the dialogue query information; Controlling the virtual avatar to execute an interaction behavior according to the text output information and the audio output information to output dialogue reply information that matches the dialogue query information.
23. A method for processing dialogue information, characterized in that, Including: Responding to an input operation instruction on the operation interface and inputting dialogue input information; Responding to a reply operation instruction on the operation interface and outputting text output information that matches the dialogue input information and audio output information that matches the dialogue input information, where the text output information is obtained by converting a text unit set, the audio output information is obtained by converting an audio unit set, the text unit set and the audio unit set are generated in parallel by using the information processing model to analyze a text tensor and an audio tensor, the audio content corresponding to the audio unit set matches the text content corresponding to the text unit set, the text tensor and the audio tensor are obtained by converting the dialogue input information, and the text tensor and the audio tensor match the information processing model, and the information processing model is trained using multimodal information samples.
24. A processing system for dialogue information, characterized in that, The system is deployed on the mobile device side and includes: An information input end for obtaining dialogue input information; An information processing terminal, configured to convert the dialogue input information into a text tensor and an audio tensor, wherein the text tensor and the audio tensor are matched with an information processing model, and the information processing model is trained using multimodal information samples; input the text tensor and the audio tensor into the information processing model, and use the information processing model to analyze the text tensor and the audio tensor to generate a text unit set and an audio unit set in parallel, wherein the audio content corresponding to the audio unit set is matched with the text content corresponding to the text unit set; convert the text unit set into a text output information that is matched with the dialogue input information, and convert the audio unit set into an audio output information that is matched with the dialogue input information; An information output terminal, configured to output the text output information and the audio output information.
25. A computing device, characterized in that, Comprising: A memory, storing an executable program; A processor, configured to run the program, wherein when the program runs, it executes the method according to any one of claims 1 to 23.
26. An electronic device, characterized in that, Comprising: A memory, storing an executable program; A processor, connected to the memory through a bus, configured to run the program, wherein when the program runs, it executes the method according to any one of claims 1 to 23.
27. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 23.
28. A computer program product, characterized in that, Comprising a computer program, which implements the method according to any one of claims 1 to 23 when executed by a processor.