Method and device for training multimodal speech language large model, equipment and medium
Patent Information
- Application Number
- CN202510772945.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-06-10
AI Technical Summary
[0013]根据本公开的一个或多个实施例,可以提升语音数据生成质量,实现更准确的人机语音问答交互。
Smart Images

Figure CN120472888B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the field of speech data processing and data generation technology, specifically to a training method for a multimodal speech and language large model, a speech data generation method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] With the development of computer technology, artificial intelligence-based generative models can be applied to various forms of natural language processing tasks, including the processing of natural language text and natural language speech. In particular, they can generate response content based on user queries to achieve interaction with users.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for training large multimodal speech and language models.
[0006] According to one aspect of this disclosure, a training method for a multimodal speech-language large model is provided, comprising: inputting first query speech data into the multimodal speech-language large model to obtain first response speech data generated by the multimodal speech-language large model; determining the query text corresponding to the first query speech data and the response text corresponding to the first response speech data; determining a first score based on the query text and the response text; determining a second score based on the speech features of the first query speech data and the speech features of the first response speech data, wherein the speech features include at least one of speech clarity, speech rate features, timbre features, intonation features, and emotion features; and adjusting the parameters of the multimodal speech-language large model based on the first score and the second score.
[0007] According to one aspect of this disclosure, a method for generating voice data is provided, comprising: acquiring inquiry voice data from a user; acquiring response voice data generated by the multimodal speech and language large model by inputting the inquiry voice data into a multimodal speech and language large model trained using the above-described training method; and returning the response voice data to the user.
[0008] According to one aspect of this disclosure, a training apparatus for a multimodal speech-language large model is provided, comprising: a first acquisition unit configured to acquire first response speech data generated by the multimodal speech-language large model by inputting first query speech data into the multimodal speech-language large model; a first determination unit configured to determine query text corresponding to the first query speech data and response text corresponding to the first response speech data; a second determination unit configured to determine a first score based on the query text and the response text; a third determination unit configured to determine a second score based on speech features of the first query speech data and speech features of the first response speech data, wherein the speech features include at least one of speech clarity, speech rate features, timbre features, intonation features, and emotion features; and an adjustment unit configured to adjust parameters of the multimodal speech-language large model based on the first score and the second score.
[0009] According to one aspect of this disclosure, a voice data generation apparatus is provided, comprising: a multimodal speech and language large model trained using the aforementioned multimodal speech and language large model training device; a second acquisition unit configured to acquire inquiry voice data from a user; a third acquisition unit configured to acquire response voice data generated by the multimodal speech and language large model by inputting the inquiry voice data into the multimodal speech and language large model; and a return unit configured to return the response voice data to the user.
[0010] According to one aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform at least one of the above-described training method for a multimodal speech-language large model and speech data generation method.
[0011] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform at least one of the above-described training method for a multimodal speech-language large model and speech data generation method.
[0012] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, is capable of implementing at least one of the above-described training method for a large multimodal speech-language model and speech data generation method.
[0013] According to one or more embodiments of this disclosure, the quality of voice data generation can be improved, enabling more accurate human-computer voice question-and-answer interaction.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0016] Figure 1 A schematic diagram of an exemplary system in which various methods described herein may be implemented, according to exemplary embodiments of the present disclosure;
[0017] Figure 2 A flowchart illustrating a method for training a large multimodal speech-language model according to exemplary embodiments of the present disclosure is shown;
[0018] Figure 3 A schematic diagram of the structure of a large-scale multimodal speech-language model according to an exemplary embodiment of the present disclosure is shown;
[0019] Figure 4 A schematic diagram of the structure of a question-answering evaluation model according to an exemplary embodiment of the present disclosure is shown;
[0020] Figure 5 A schematic diagram illustrating the training process of a large multimodal speech-language model according to an exemplary embodiment of the present disclosure is shown.
[0021] Figure 6 A flowchart of a speech data generation method according to an exemplary embodiment of the present disclosure is shown;
[0022] Figure 7 A schematic diagram of a speech data generation process according to an exemplary embodiment of the present disclosure is shown;
[0023] Figure 8 A structural block diagram of a training apparatus for a large multimodal speech-language model according to an exemplary embodiment of the present disclosure is shown;
[0024] Figure 9 A structural block diagram of a speech data generation apparatus according to an exemplary embodiment of the present disclosure is shown;
[0025] Figure 10 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0027] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0028] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0029] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0030] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0031] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable training methods for large multimodal speech-language models.
[0032] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0033] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0034] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to send query voice data. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0035] Client devices 101, 102, 103, 104, 105, and / or 106 may include various categories of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various categories and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices can run a variety of different applications, such as various Internet-related applications, communication applications (e.g., email applications), short message service (SMS) applications, and can use various communication protocols.
[0036] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0037] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0038] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0039] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0040] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0041] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different categories. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0042] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be different categories of databases, such as key-value stores, object stores, or regular stores supported by a file system.
[0043] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0044] In related technologies, when using AI-based data generation models for human-computer dialogue, these models are typically large language models capable of processing and generating natural language text. In this case, when voice dialogue between the user and the computer is required, a cascaded system of speech recognition module—large language model—speech synthesis module can be used to generate response voice data for broadcasting to the user. Specifically, speech recognition technology is used to convert the user's input query speech data into query text, the large language model generates response text based on the query text, and then speech synthesis technology generates the corresponding response voice data. While this approach utilizes the existing capabilities of the large language model, cascading errors can affect the overall voice question-and-answer performance, resulting in insufficient accuracy of the voice responses. Furthermore, this approach fails to leverage AI networks to generate diverse voice responses in terms of speech rate, timbre, and emotion, thus failing to fully meet user needs, and the quality of the response voice data still has room for improvement.
[0045] Based on this, this disclosure provides a training method for a multimodal speech and language large-scale model. It sets reward scores from two perspectives—text content and speech features—for the response speech data directly generated by the multimodal speech and language large-scale model. This allows for optimization of the accuracy of the response content and the quality of the response speech (e.g., higher clarity, matching of speaking emotion to the query data, appropriate speaking speed, etc.) based on the text score and speech score respectively. The optimized model then achieves more accurate speech responses.
[0046] Figure 2 A flowchart of a training method 200 for a large multimodal speech-language model according to an exemplary embodiment of the present disclosure is shown. Figure 2 As shown, method 200 includes:
[0047] Step S201: By inputting the first query voice data into the multimodal speech and language big model, the first response voice data generated by the multimodal speech and language big model is obtained;
[0048] Step S202: Determine the query text corresponding to the first query voice data and the reply text corresponding to the first reply voice data;
[0049] Step S203: Determine the first score based on the query text and the response text;
[0050] Step S204: Determine a second score based on the speech features of the first inquiry speech data and the first response speech data, wherein the speech features include at least one of speech clarity, speech rate features, timbre features, intonation features, and emotion features; and
[0051] Step S205: Based on the first score and the second score, adjust the parameters of the multimodal speech-language large model.
[0052] By applying the above method 200, reward scores from two perspectives—text content and speech features—can be set for the response speech data directly generated by the model during the training process of a multimodal speech and language large model. This allows for more precise optimization of the accuracy of the response content and the quality of the response speech based on the text content score and speech feature score. For example, the response speech can be made clearer, the speech rate more appropriate, or the emotional expression of the response speech can be better matched with the emotional expression of the inquiry speech. In this way, the optimized model can be used to achieve more accurate and high-quality speech responses.
[0053] In some examples, the multimodal speech-language large-scale model can be a pre-trained speech-language large-scale model using a large-scale corpus. This model can intelligently understand and generate content from multiple modalities (e.g., speech and text information). In other examples, the multimodal speech-language large-scale model employs a cross-modal encoder-decoder architecture and attention mechanism to achieve semantic alignment and joint modeling of multimodal information. By applying the above method 200 to optimize the pre-trained large-scale model, the diversity of speech features in the speech data can be further optimized based on fully utilizing the semantic understanding and speech generation capabilities of the pre-trained large-scale model. This results in a more efficient multimodal speech-language large-scale model that can express rich and accurate speech features, thereby achieving more accurate human-computer speech question answering.
[0054] In some examples, the determination of the query text corresponding to the first query voice data and the response text corresponding to the first response voice data in step S202 can be achieved using Automatic Speech Recognition (ASR) technology. That is, text recognition is performed on the first query voice data and the first response voice data, and the query text and response text are determined based on the text recognition results. This allows for a more accurate assessment of the accuracy of the response content based on the text content, thereby improving the quality of the model's response.
[0055] In some examples, step S203 inputs the query text and response text into the scoring model to obtain a first score output by the model. This first score can be used to indicate whether the facts described in the response text are accurate, and whether the content of the response text matches the question's intent. In some examples, the scoring model can be trained under supervision using question-and-answer text pairs labeled with reference scores, and the output of the scoring model can then be used to evaluate the content quality of the response text. In some examples, the scoring model can be built based on a Large Language Model (LLM) trained using a large-scale corpus. For example, it can be fine-tuned using the labeled data mentioned above on top of a pre-trained LLM model, so that the evaluation model can more comprehensively and accurately evaluate the quality of the response content.
[0056] According to some embodiments, determining the second score based on the speech features of the first query speech data and the first response speech data in step S204 includes: in response to determining that the first query speech data includes descriptive information regarding the speech features of the first response speech data, determining the second score based on the descriptive information, the speech features of the first query speech data, and the speech features of the first response speech data. Therefore, it is possible to obtain explicit instructions contained in the query data during the scoring process (e.g., clear requirements for response timbre, tone, and speed), so that the optimized model can output response data that better meets the requirements of the question.
[0057] According to some embodiments, determining the second score based on the speech features of the first inquiry speech data and the first response speech data in step S204 includes: determining the speaker's identity features based on the speech features of the first inquiry speech data; and determining the second score based on the identity features, the response text, and the speech features of the first response speech data. This allows for the capture of speech features from the inquiry data during the scoring process, avoiding contradictions between the response data and the inquiry data. For example, when the inquiry data is a female voice, the response content should not contain any terms of address specific to men, thereby improving the accuracy of the response.
[0058] According to some embodiments, determining the second score based on the speech features of the first query speech data and the first response speech data in step S204 includes: inputting the first query speech data and the first response speech data into a question-and-answer evaluation model to determine the second score output by the question-and-answer evaluation model, wherein the question-and-answer evaluation model is trained using first sample query speech data, first sample response speech data, and reference scores. By training the question-and-answer evaluation model using labeled data, and then using the question-and-answer evaluation model to output reward scores based on the speech features of the speech data during the training process of a multimodal speech and language large model, the efficiency and accuracy of scoring can be improved.
[0059] In some examples, the question-answering evaluation model can be trained using multiple question-answering speech data pairs labeled with reference scores. To improve the accuracy of the model's evaluation of the speech features of the response speech data, the training data can include question-answering speech data pairs with the same text content but different speech features. For example, the training data can simultaneously include: speech data 1 corresponding to content A and timbre B, and speech data 2 corresponding to content A and timbre C. As another example, the training data can also simultaneously include: speech data 3 corresponding to content A and emotion D, and speech data 4 corresponding to content A and emotion E. By configuring question-answering speech data pairs with the same text content but different speech features in the training data, the influence of differences in text content on speech data features can be eliminated. This allows the evaluation model to more accurately learn different speech features (such as speech clarity, speech rate, timbre, intonation, and emotion), thus more accurately evaluating the quality of speech data with different speech features.
[0060] In some examples, steps S203 and S204 can also determine the first and second scores based on other methods, such as using predefined scoring rules. In this example, the predefined scoring rules can include multiple dimensions, such as content conciseness, speech clarity, and whether the speech data expresses positive emotions. By predefining scoring rules and using them to guide the optimized training of the multimodal speech and language model, the output of the multimodal speech and language model can better meet the needs of real-world application scenarios, achieving higher-quality human-computer speech question answering.
[0061] According to some embodiments, adjusting the parameters of the multimodal speech and language large model based on the first score and the second score in step S205 includes: determining reward information based on the first score and the second score; and adjusting the parameters of the multimodal speech and language large model based on the reinforcement learning strategy corresponding to the reward information. Therefore, reinforcement learning training can be performed based on the score information to obtain a multimodal speech and language large model with better performance.
[0062] In some examples, step S205 can involve model tuning using RLHF (Reinforcement Learning with Human Feedback). This means adjusting model parameters based on various types of reinforcement learning strategies, such as proximal policy optimization, group policy relative optimization, and direct preference optimization. By applying scoring reward information and policy optimization algorithms to tune the parameters of a large multimodal speech and language model, the output of the model can be made as close as possible to the expected result of the scoring reward information, resulting in higher-quality response speech data and improving the accuracy of human-computer speech interaction.
[0063] According to some embodiments, method 200 further includes: inputting the second query speech data into the multimodal speech-language large model to obtain second response speech data generated by the multimodal speech-language large model; obtaining reference response speech data corresponding to the second query speech data; and adjusting the parameters of the multimodal speech-language large model based on the second response speech data and the reference response speech data. Thus, by combining model optimization methods that utilize labeled data for supervised training with reward scoring as a basis for training, the quality of the model's output response speech data can be further improved.
[0064] In some examples, the first and second query speech data can be the same. That is, after inputting the query speech data into the multimodal speech and language model, a quality score is given to the model-generated response speech data. Simultaneously, a loss value can be calculated based on the model-generated response speech data and the corresponding reference response speech data for the query speech data. Alternatively, the quality score can be influenced by the difference between the model output and the reference result, combining the reference annotation information and quality score information of the sample data for model optimization training. In some examples, the first and second query speech data can be different, meaning different training methods are applied to the multimodal speech and language model. In one example, the training process of the multimodal speech and language model includes the following two stages: In the first stage, supervised fine-tuning (SFT) can be performed on the pre-trained speech and language model. This involves inputting sample query speech data labeled with reference response speech data into the pre-trained model, calculating the loss value based on the model output and the annotation content, and then adjusting the model parameters based on the loss value. This corresponds to the technique described above of training using the second query speech data and the corresponding reference response speech data for the second query speech data. The training in this first stage could involve using autoregressive learning with a maximum likelihood objective to improve training efficiency and effectiveness. In the second stage, unlabeled first query speech data could be input into a multimodal speech and language model. Content quality and audio quality scores would then be assigned to the model's output. Reinforcement learning training would be performed based on these two reward scores to more accurately optimize the accuracy of the responses and the model's ability to represent diverse speech features, thereby improving model performance.
[0065] According to some embodiments, the multimodal speech-language large model generates the second response speech data in the following manner: generating predictive thinking information based on the second inquiry speech data, wherein the predictive thinking information includes descriptive information of speech features for the second response speech data; and generating the second response speech data based on the predictive thinking information, wherein method 200 further includes: obtaining reference thinking information corresponding to the second inquiry speech data; and adjusting the parameters of the multimodal speech-language large model based on the predictive thinking information and the reference thinking information. Thus, a thought process based on a thought chain can be incorporated into the response data generation process, that is, the large model decomposes the generation task, outputs based on a logical chain of thinking before generation, and uses thinking information to indicate the speech features to be output, so that the generated response speech data is more accurate.
[0066] In some examples, the reasoning mechanism of the multimodal speech and language large-scale model is built based on Chain-of-Thought (CoT) technology. This means the model simulates the human cognitive pattern of "step-by-step thinking and deduction," outputting thought information after receiving input query speech data, and then generating response speech data. By having the model first consider the speech features of the response speech data to be generated, the quality of the generated speech data can be improved. In this example, the sample data annotation during model training includes reference thought information. This allows for precise optimization of the model's thinking process based on the reference thought information and the predicted thought information output by the model, further improving the accuracy of the model's reasoning and the quality of the generated data.
[0067] According to some embodiments, the reference thinking information is text information. Generating predicted thinking information based on the second query voice data includes: encoding the second query voice data into a query semantic vector in a semantic vector space; and generating a thinking semantic vector in the semantic vector space based on the query semantic vector. Furthermore, generating the second response voice data based on the predicted thinking information includes: generating a response semantic vector in the semantic vector space based on the thinking semantic vector; and decoding the response semantic vector into the second response voice data. Additionally, adjusting the parameters of the multimodal speech-language large model based on the predicted thinking information and the reference thinking information includes: decoding the thinking semantic vector into predicted thinking text; and adjusting the parameters of the multimodal speech-language large model based on the predicted thinking text and the reference thinking information. Therefore, text information can be used to more easily and accurately annotate the model's thinking information, thereby improving the accuracy of the model's thinking. In this context, the question-thinking-response chain of voice question answering using the model includes both voice modal data and text modal data. By converting the voice modal data and text modal data into the same semantic vector space, unified intelligent content understanding of multiple modalities is achieved, avoiding the cascading loss caused by modality conversion, optimizing model performance, and improving the accuracy of data generation.
[0068] In some examples, the operation of encoding speech modal data and text modal data into semantic vectors within the speech vector space is implemented using tokenizer technology. Specifically, the tokenizer can split the original speech modal data and text modal data into discrete text tokens and audio tokens, and then map them into a numerical sequence that the computer can process, i.e., encode them into semantic vectors. In this example, the mapping vocabulary of text tokens and the mapping vocabulary of audio tokens are concatenated to form the overall vocabulary space of the multimodal speech and language model. This achieves the mapping of multimodal information to the same semantic vector space, enabling the model to understand and process multimodal information and improving the quality of data generation.
[0069] According to some embodiments, the reference response voice data is marked with voice breakpoints, and the second response voice data consists of a first voice segment and a second voice segment. Adjusting the parameters of the multimodal speech-language model based on the second response voice data and the reference response voice data includes: splitting the reference response voice data into a first reference segment and a second reference segment based on the voice breakpoints; adjusting the parameters of the multimodal speech-language model based on the first voice segment and the first reference segment; and adjusting the parameters of the multimodal speech-language model based on the second voice segment and the second reference segment. In some examples, the multimodal speech-language model is configured to output in stages during data generation. Once a portion of the response voice has been generated, it is output to the user first, while the remaining response voice continues to be generated simultaneously. In this case, the multimodal speech-language model can more quickly deliver voice responses to the user during human-computer voice question-and-answer sessions, reducing user waiting time. By using training data labeled with speech breakpoints to tune and optimize the model, the model can learn when to split the response speech data. This allows it to determine when to output a portion of the response speech first during the data generation process. By outputting the response speech data in the form of speech segment sequences, it is not necessary to wait for all the response data to be generated before the speech is played, thus improving the smoothness of the model's response during human-computer voice interaction and optimizing the user experience.
[0070] Figure 3 A schematic diagram of the structure of a large-scale multimodal speech-language model according to exemplary embodiments of the present disclosure is shown. Figure 3As shown, the multimodal speech and language large-scale model includes a speech encoder 301, a speech decoder 302, and a thinking and generation network 303. In this example, the speech encoder 301 encodes the query speech data into query semantic vectors, enabling the thinking and generation network 303 to think and respond based on the query semantic vectors, and then sequentially outputs thinking semantic vectors and response semantic vectors. In one example, the thinking and generation network 303 first determines the thinking semantic vector based on the query semantic vector, uses the thinking content to indicate the content and speech features to be generated, and then performs further thinking and generation based on the query semantic vector and the thinking semantic vector to obtain the final response semantic vector. In one example, the thinking semantic vector can be decoded into thinking text and output to display explicit thinking information to the user, who can then provide feedback to the model as needed after viewing the thinking information. The speech decoder 302 decodes the response semantic vector into response speech data and outputs it. In this example, the query semantic vector, the thinking semantic vector, and the response semantic vector are in the same semantic vector space. By mapping multimodal information to this semantic vector space, the thinking and generation network 303 can perform unified intelligent understanding and processing of multimodal information to improve the accuracy of data generation.
[0071] Figure 4 A schematic diagram of the structure of a question-answering evaluation model according to an exemplary embodiment of this disclosure is shown. Figure 4 As shown, the question-answering evaluation model includes a speech encoder 401 and an evaluation network 402. The speech encoder 401 encodes the question and response speech data to be evaluated into question semantic vectors and response semantic vectors, enabling the evaluation network 402 to obtain a second score based on these vectors. In some examples, the speech encoder 401 in the question-answering evaluation model can be the same unit as the speech encoder 301 in the thinking and generation network 303 described earlier; that is, the thinking and generation network 303 and the evaluation network 402 in the question-answering evaluation model can perform intelligent understanding and information processing based on the same semantic vector space.
[0072] Figure 5 A schematic diagram illustrating the training process of a large multimodal speech-language model according to an exemplary embodiment of the present disclosure is shown. Figure 5As shown, after inputting the query speech data into the multimodal speech and language model 501 and obtaining the response speech data, the speech recognition system 502 determines the corresponding query text and response text. Then, the first evaluation model 503 outputs a first score based on the query and response texts, indicating whether the response content is accurate and comprehensive. The second evaluation model 504 outputs a second score based on the query and response speech data, indicating whether the response speech data is clear or whether its speech rate, timbre, and emotion meet the requirements of the query data. By training and optimizing the multimodal speech and language model 503 based on the first and second scores, the accuracy of the output response content and the richness of speech expression can be precisely optimized, thereby improving the accuracy of human-computer speech question answering.
[0073] In one example, the multimodal speech-language large model 503 was fine-tuned using labeled sample data prior to the aforementioned training phase. The labeled sample data can simultaneously annotate reference speech responses and reference thinking information, allowing for targeted optimization of the model's thinking process and data generation process to improve training efficiency and optimize model performance.
[0074] In one example, the multimodal speech and language large model 503 outputs response speech data in the form of speech segment sequences. In this case, after the multimodal speech and language large model 503 has output all the response speech segments, it can be concatenated into complete response speech data, which can then be converted into response text and scored. When the labeled sample data contains speech breakpoints from the reference speech response data, the breakpoint information can be used to indicate whether the timing of the splitting of the response speech segments output by the model is accurate. By optimizing based on the labeled information, the model can learn the accurate timing of speech response splitting, thereby improving the fluency of the speech response.
[0075] According to one aspect of this disclosure, a method for generating speech data is provided. Figure 6 A flowchart of a speech data generation method 600 according to an exemplary embodiment of the present disclosure is shown. Figure 6 As shown, method 600 includes:
[0076] Step S601: Obtain voice data of user inquiry;
[0077] Step S602: By inputting the query voice data into the multimodal speech and language large model trained using method 200, the response voice data generated by the multimodal speech and language large model is obtained; and
[0078] Step S603: Return the reply voice data to the user.
[0079] By applying the aforementioned multimodal speech and language model for human-computer voice question answering, we can return richer and more accurate voice response data to users, including voice features such as timbre, emotion, and speech rate. This makes the response content more in line with user needs and improves the user experience.
[0080] Figure 7 A schematic diagram of a speech data generation process according to an exemplary embodiment of the present disclosure is shown. Figure 7 As shown, the multimodal speech and language model 700, after receiving inquiry voice 01, can first think and output stage-specific thinking information A, and then output response voice a based on thinking information A. While response voice a is playing, the model can continue to think and output the next stage of thinking information B, and then output response voice b based on thinking information B. When the model determines that it has output all the response content for inquiry voice 01, the end of response voice b can include a symbol indicating that the answer has been completed, to indicate the end of the question-and-answer round for inquiry voice 01. After the previous round ends, the user can continue to input inquiry voice 02. The multimodal speech and language model 700 can think and generate data based on the new inquiry voice 02 to output thinking information C and response voice c. By having the model output response voice data in the form of a sequence of voice segments during the data generation process, serialized voice output can be achieved during voice interaction, without waiting for all response data to be generated before voice playback, reducing user waiting time, improving the fluency of voice response, and thus improving user experience.
[0081] According to one aspect of this disclosure, a training apparatus for a large multimodal speech-language model is provided. Figure 8 A structural block diagram of a training apparatus 800 for a large multimodal speech-language model according to an exemplary embodiment of the present disclosure is shown. Figure 8 As shown, the device 800 includes:
[0082] The first acquisition unit 801 is configured to acquire the first response voice data generated by the multimodal speech and language model by inputting the first query voice data into the multimodal speech and language model.
[0083] The first determining unit 802 is configured to determine the query text corresponding to the first query voice data and the reply text corresponding to the first reply voice data.
[0084] The second determining unit 803 is configured to determine a first score based on the query text and the response text;
[0085] The third determining unit 804 is configured to determine a second score based on the speech features of the first inquiry speech data and the speech features of the first response speech data, wherein the speech features include at least one of speech clarity, speech rate features, timbre features, intonation features, and emotion features; and
[0086] The adjustment unit 805 is configured to adjust the parameters of the multimodal speech-language large model based on the first score and the second score.
[0087] According to some embodiments, the third determining unit 804 is configured to: in response to determining that the first inquiry voice data includes descriptive information about the voice features of the first response voice data, determine the second score based on the descriptive information, the voice features of the first inquiry voice data, and the voice features of the first response voice data.
[0088] According to some embodiments, the third determining unit 804 includes: a first determining subunit configured to determine the speaker's identity features based on the voice features of the first inquiry voice data; and a second determining subunit configured to determine the second score based on the identity features, the reply text, and the voice features of the first reply voice data.
[0089] According to some embodiments, the third determining unit 804 is configured to: determine the second score output by the question-and-answer evaluation model by inputting the first inquiry voice data and the first response voice data into the question-and-answer evaluation model, wherein the question-and-answer evaluation model is trained using the first sample inquiry voice data, the first sample response voice data and the reference score.
[0090] According to some embodiments, the adjustment unit 805 includes: a third determining subunit configured to determine reward information based on the first score and the second score; and a first adjustment subunit configured to adjust the parameters of the multimodal speech-language large model based on the reinforcement learning strategy corresponding to the reward information.
[0091] According to some embodiments, the first acquisition unit 801 is further configured to acquire second response voice data generated by the multimodal speech-language large model by inputting the second inquiry voice data into the multimodal speech-language large model. The device 800 further includes: a fourth acquisition unit configured to acquire reference response voice data corresponding to the second inquiry voice data, and wherein the adjustment unit 805 is further configured to adjust the parameters of the multimodal speech-language large model based on the second response voice data and the reference response voice data.
[0092] According to some embodiments, the multimodal speech-language large model is configured to: generate predictive thinking information based on the second inquiry speech data, wherein the predictive thinking information includes descriptive information of speech features for the second response speech data; and generate the second response speech data based on the predictive thinking information, wherein the device 800 further includes: a fifth acquisition unit configured to acquire reference thinking information corresponding to the second inquiry speech data, and the adjustment unit 805 configured to adjust the parameters of the multimodal speech-language large model based on the predictive thinking information and the reference thinking information.
[0093] According to some embodiments, the reference thinking information is text information, and the multimodal speech-language large model is configured to: encode the second query speech data into a query semantic vector in a semantic vector space; generate a thinking semantic vector in the semantic vector space based on the query semantic vector; generate a response semantic vector in the semantic vector space based on the thinking semantic vector; and decode the response semantic vector into the second response speech data, wherein the adjustment unit 805 includes: a decoding subunit configured to decode the thinking semantic vector into predicted thinking text; and a second adjustment subunit configured to adjust the parameters of the multimodal speech-language large model based on the predicted thinking text and the reference thinking information.
[0094] According to some embodiments, the reference response speech data is marked with speech breakpoints, the second response speech data is composed of a first speech segment and a second speech segment, and the adjustment unit 805 includes: a splitting subunit configured to split the reference response speech data into a first reference segment and a second reference segment based on the speech breakpoints; and a third adjustment subunit configured to adjust the parameters of the multimodal speech-language large model based on the first speech segment and the first reference segment; and to adjust the parameters of the multimodal speech-language large model based on the second speech segment and the second reference segment.
[0095] According to one aspect of this disclosure, a voice data generation apparatus is provided. Figure 9 A structural block diagram of a speech data generation apparatus 900 according to an exemplary embodiment of the present disclosure is shown. Figure 9 As shown, the device 900 includes:
[0096] The multimodal speech and language large model 901 is obtained by training using the training device 800 of the multimodal speech and language large model as described above;
[0097] The second acquisition unit 902 is configured to acquire inquiry voice data from the user;
[0098] The third acquisition unit 903 is configured to acquire response voice data generated by the multimodal speech-language model by inputting the query voice data into the multimodal speech-language model; and
[0099] Return unit 904 is configured to return the response voice data to the user.
[0100] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0101] According to one aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform at least one of the above-described methods for training a large multimodal speech-language model and for generating speech data.
[0102] According to one aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform at least one of the above-described training method for a large multimodal speech-language model and speech data generation method.
[0103] According to one aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein, when the computer program is executed by a processor, it implements at least one of the above-described training method for a multimodal speech-language large model and the speech data generation method.
[0104] refer to Figure 10 The present invention describes a structural block diagram of an electronic device 1000 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0105] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0106] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, output unit 1007, storage unit 1008, and communication unit 1009. Input unit 1006 can be any type of device capable of inputting information to device 1000. Input unit 1006 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 1007 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1008 may include, but is not limited to, a hard disk and an optical disk. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0107] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as methods for training a large multimodal speech-language model or methods for generating speech data. For example, in some embodiments, the methods for training a large multimodal speech-language model or methods for generating speech data can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the methods for training a large multimodal speech-language model or methods for generating speech data described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured by any other suitable means (e.g., by means of firmware) to perform a training method for a large multimodal speech-language model or a speech data generation method.
[0108] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0112] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0113] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0114] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0115] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A training method for a large-scale multimodal speech-language model, comprising: By inputting the first query voice data into the multimodal speech and language big model, the first response voice data generated by the multimodal speech and language big model is obtained; Determine the query text corresponding to the first query voice data and the reply text corresponding to the first reply voice data; A first score is determined based on the query text and the response text; A second score is determined based on the speech features of the first inquiry speech data and the speech features of the first response speech data, wherein the speech features include at least one of speech clarity, speech rate features, timbre features, tone features, and emotion features; Based on the first score and the second score, the parameters of the multimodal speech-language large model are adjusted; By inputting the second query voice data into the multimodal speech-language model, the second response voice data generated by the multimodal speech-language model is obtained in the following manner: Predictive thinking information is generated based on the second inquiry voice data, wherein the predictive thinking information includes descriptive information about the voice features of the second response voice data; and The second response voice data is generated based on the predicted thinking information; Obtain the reference response voice data corresponding to the second inquiry voice data; Obtain the reference thinking information corresponding to the second inquiry voice data; Based on the second response speech data and the reference response speech data, the parameters of the multimodal speech-language large model are adjusted; and Based on the predicted thinking information and the reference thinking information, the parameters of the multimodal speech and language large model are adjusted.
2. The method as described in claim 1, wherein, The process of determining the second score based on the voice features of the first inquiry voice data and the first response voice data includes: In response to determining that the first inquiry voice data includes descriptive information about the voice features of the first response voice data, a second score is determined based on the descriptive information, the voice features of the first inquiry voice data, and the voice features of the first response voice data.
3. The method of claim 1, wherein, The process of determining the second score based on the voice features of the first inquiry voice data and the first response voice data includes: The speaker's identity is determined based on the speech features of the first inquiry speech data; and The second score is determined based on the identity features, the reply text, and the voice features of the first reply voice data.
4. The method of claim 1, wherein, The process of determining the second score based on the voice features of the first inquiry voice data and the first response voice data includes: The second score output by the question-and-answer evaluation model is determined by inputting the first query voice data and the first response voice data into the question-and-answer evaluation model, wherein the question-and-answer evaluation model is trained using the first sample query voice data, the first sample response voice data and the reference score.
5. The method of claim 1, wherein, The adjustment of the parameters of the multimodal speech-language large model based on the first score and the second score includes: Based on the first score and the second score, reward information is determined; and Based on the reinforcement learning strategy corresponding to the reward information, the parameters of the multimodal speech and language large model are adjusted.
6. The method of claim 1, wherein, The reference thinking information is text information, and the generation of predictive thinking information based on the second inquiry voice data includes: The second query voice data is encoded into a query semantic vector in the semantic vector space; and Based on the query semantic vector, a thinking semantic vector is generated in the semantic vector space. Furthermore, the step of generating the second response voice data based on the predicted thinking information includes: Based on the aforementioned semantic vector of thought, a response semantic vector is generated in the semantic vector space; and The response semantic vector is decoded into the second response voice data. Furthermore, adjusting the parameters of the multimodal speech-language large model based on the predicted thinking information and the reference thinking information includes: Decode the thought semantic vector into predicted thought text; and Based on the predicted thinking text and the reference thinking information, the parameters of the multimodal speech-language large model are adjusted.
7. The method according to any one of claims 1-6, wherein, The reference response voice data is marked with voice breakpoints, and the second response voice data consists of a first voice segment and a second voice segment. Furthermore, adjusting the parameters of the multimodal speech-language large model based on the second response speech data and the reference response speech data includes: Based on the aforementioned voice breakpoint, the reference response voice data is split into a first reference segment and a second reference segment; Based on the first speech segment and the first reference segment, adjust the parameters of the multimodal speech-language large model; and Based on the second speech segment and the second reference segment, the parameters of the multimodal speech-language large model are adjusted.
8. A method for generating speech data, comprising: Obtain voice data of user inquiries; By inputting the query speech data into a multimodal speech-language large model trained using the method described in any one of claims 1-7, the response speech data generated by the multimodal speech-language large model is obtained; and The response voice data is returned to the user.
9. A training device for a large-scale multimodal speech-language model, comprising: The first acquisition unit is configured to acquire the first response voice data generated by the multimodal speech and language model by inputting the first query voice data into the multimodal speech and language model. The first determining unit is configured to determine the query text corresponding to the first query voice data and the reply text corresponding to the first reply voice data; The second determining unit is configured to determine a first score based on the query text and the response text; The third determining unit is configured to determine a second score based on the speech features of the first inquiry speech data and the speech features of the first response speech data, wherein the speech features include at least one of speech clarity, speech rate features, timbre features, intonation features, and emotion features; and The adjustment unit is configured to adjust the parameters of the multimodal speech-language large model based on the first score and the second score. The first acquisition unit is further configured to acquire second response voice data generated by the multimodal speech-language model by inputting the second query voice data into the multimodal speech-language model in the following manner: Predictive thinking information is generated based on the second inquiry voice data, wherein the predictive thinking information includes descriptive information about the voice features of the second response voice data; and The second response voice data is generated based on the predicted thinking information. The training device also includes: The fourth acquisition unit is configured to acquire reference response voice data corresponding to the second inquiry voice data; and The fifth acquisition unit is configured to acquire reference thinking information corresponding to the second inquiry voice data. The adjustment unit is further configured to: Based on the second response speech data and the reference response speech data, the parameters of the multimodal speech-language large model are adjusted; and Based on the predicted thinking information and the reference thinking information, the parameters of the multimodal speech and language large model are adjusted.
10. A voice data generation apparatus, comprising: A large multimodal speech and language model trained using the device as described in claim 9; The second acquisition unit is configured to acquire inquiry voice data from the user; The third acquisition unit is configured to acquire response voice data generated by the multimodal speech and language model by inputting the query voice data into the multimodal speech and language model. as well as The return unit is configured to return the response voice data to the user.
11. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
13. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Generative large language model training method and model-based man-machine voice interaction method
CN116127046A
Vehicle interaction method and device, model training method and device, server and storage medium
CN116758913A