Method and device for generating voice reply for input voice

By connecting the speech encoder and multiple speech predictors in series, the problem of delay and insufficient quality in speech signal processing of large language models is solved, and faster speech generation speed and higher reply quality are achieved.

CN120452439APending Publication Date: 2025-08-08ALIPAY (HANGZHOU) INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510422174.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing large language models have problems such as large voice response delay and poor quality in speech signal processing, which makes it difficult to effectively process the timing characteristics and information density of speech signals, resulting in slow voice generation speed and poor quality.

Method used

By connecting a tandem speech encoder, a large language model and multiple speech predictors, multi-word element prediction is used to generate speech replies, improve speech decoding speed and enhance structural representation of speech signals.

Benefits of technology

It significantly improves the output rate of pronunciation word elements, reduces the delay in voice response, and improves the quality of voice reply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452439A_ABST
    Figure CN120452439A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for generating voice reply for input voice, and the method comprises the steps: obtaining first voice input by a user, inputting the first voice into a voice encoder, and obtaining a first feature; obtaining a second feature based on the first feature through a preset large language model; through a voice decoder and N voice predictors which are sequentially connected in series, N + 1 first lexical element features which are sequentially arranged are obtained based on the second feature, the voice decoder outputs a first first lexical element feature in the N + 1 first lexical element features based on the second feature, and the first second lexical element feature is a second lexical element feature in the N + 1 first lexical element features; the ith voice predictor in the N voice predictors outputs the (i + 1) th first lexical element feature based on the ith first lexical element feature in the N + 1 first lexical element features; and generating a voice reply based on the N + 1 first lexical element features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of large model technology, and more particularly, to a method and apparatus for generating a voice response to an input voice. Background Art

[0002] A large language model (LLM) refers to a deep learning model based on natural language processing that is pre-trained on a large-scale text corpus and contains parameters in the hundreds of millions or more. Large language models are often used for tasks such as semantic understanding and common sense reasoning based on text data. However, this text-based model essentially lacks the ability to process voice signals and cannot directly implement voice input, understanding, and generation. A solution for implementing voice response based on a large model can respond to voice commands by connecting a voice encoder, a large language model, and a voice decoder in series, but there are also problems with large voice response delays and poor voice response quality. Summary of the Invention

[0003] The embodiments in this specification aim to provide a method and apparatus for obtaining training data for fine-tuning a large model, which can significantly improve the speed of speech decoding, reduce speech response delay, and improve the quality of speech responses, thereby addressing the shortcomings of the existing technology.

[0004] According to a first aspect, there is provided a method for generating a voice response for input speech, comprising:

[0005] Obtaining a first speech input by a user, and inputting the first speech into a speech encoder to obtain a first feature;

[0006] Obtaining a second feature based on the first feature by presetting a large language model;

[0007] By sequentially connecting a speech decoder and N speech predictors, N+1 first word-unit features arranged in sequence are obtained based on the second feature, wherein the speech decoder outputs the first first word-unit feature among the N+1 first word-unit features based on the second feature, and the i-th speech predictor among the N speech predictors outputs the i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features; and a speech response is generated based on the N+1 first word-unit features.

[0008] In a possible implementation, the speech decoder outputs a first first word-unit feature among the N+1 first word-unit features based on the second feature, and the i-th speech predictor among the N speech predictors outputs an i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features, including:

[0009] The speech decoder outputs a first first word-unit feature among the N+1 first word-unit features and a first processing feature based on the second feature;

[0010] The i-th speech predictor among the N speech predictors outputs a third processed feature and the i+1-th first word unit feature based on the second processed feature output by the speech decoder or speech predictor arranged previously and the i-th first word unit feature among the N+1 first word unit features.

[0011] In one possible implementation, generating a speech response based on the N+1 first word-unit features includes:

[0012] A second word-unit feature is generated according to the first word-unit feature through a classification head, and a speech response is generated based on the second word-unit feature.

[0013] In one possible implementation, generating a speech response based on the second word-unit feature includes:

[0014] The second word-unit feature is input into a speech generator to obtain a speech response.

[0015] In a possible implementation, obtaining the second feature based on the first feature by presetting a large language model includes:

[0016] The first feature is input into a downsampling adapter to obtain a third feature, where the dimension of the third feature is smaller than that of the first feature. The third feature is input into a preset large model to obtain a second feature.

[0017] In one possible implementation, sequentially arranged N+1 first word-unit features are obtained based on the second features by sequentially connecting a speech decoder and N speech predictors, including:

[0018] Inputting the second feature into the voice adapter to obtain a fourth feature;

[0019] By sequentially connecting the speech decoders and N speech predictors, N+1 first word-unit features arranged in sequence are obtained according to the fourth feature.

[0020] In one possible implementation, the method further includes:

[0021] The second feature is input into a text decoder to obtain a first text corresponding to the second speech.

[0022] In one possible implementation, the speech decoder and N speech predictors are pre-trained by the following process:

[0023] Obtaining a second voice input by a user, and inputting the second voice into a speech encoder to obtain a third feature;

[0024] By presetting a large language model, a fourth feature is obtained based on the third feature;

[0025] Obtaining, by the speech decoder and the N speech predictors, N+1 third word-unit features arranged in sequence based on the fourth feature, wherein the speech decoder outputs a first third word-unit feature among the N+1 predicted word-unit features based on the fourth feature, and the i-th speech predictor among the N speech predictors outputs an i+1-th third word-unit feature based on the i-th predicted word-unit feature among the N+1 third word-unit features;

[0026] Obtain word-unit labels corresponding to the N+1 third word-unit features, determine a first loss based on the N+1 third word-unit features and the word-unit labels; and update network parameters of the speech decoder and N speech predictors based on the first loss.

[0027] In one possible implementation, determining a first loss based on the N+1 third word features and the word tag includes: determining N+1 prediction losses based on each third word feature in the N+1 third word features and the corresponding word tag, and determining the first loss based on the sum of the N+1 prediction losses.

[0028] In one possible implementation, determining the first loss based on the sum of the N+1 prediction losses includes: determining the first loss based on the weighted sum of the N+1 prediction losses, wherein the weighted weight of each prediction loss is determined based on the arrangement position of the speech decoder or speech predictor that generates the third word feature corresponding to the prediction loss in the speech decoder and the N speech predictors.

[0029] In one possible implementation, the method further includes:

[0030] According to the first loss, network parameters of the preset large language model are updated.

[0031] According to a second aspect, there is provided an apparatus for obtaining a voice response to a voice input, the apparatus comprising:

[0032] an acquiring unit configured to acquire a first speech input by a user, input the first speech into a speech encoder, and obtain a first feature;

[0033] A first processing unit is configured to obtain a second feature based on the first feature by using a preset large language model;

[0034] a second processing unit configured to obtain N+1 first word-unit features arranged in sequence based on the second feature by sequentially connecting a speech decoder and N speech predictors, wherein the speech decoder outputs a first first word-unit feature among the N+1 first word-unit features based on the second feature, and an i-th speech predictor among the N speech predictors outputs an i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features;

[0035] The generation unit is configured to generate a speech response based on the N+1 first word-unit features.

[0036] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in the first aspect.

[0037] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the first aspect is implemented.

[0038] By utilizing one or more of the methods, devices, computing devices, and storage media in the above aspects, the output rate of speech words can be greatly improved, thereby increasing the speed of speech decoding and generating speech responses, reducing speech response delays, and improving the quality of speech responses. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0040] Figure 1 A schematic diagram showing a solution for implementing speech response based on a large model;

[0041] Figure 2 A schematic diagram illustrating a method for generating a voice response to an input voice according to an embodiment of the present specification;

[0042] Figure 3 A flowchart illustrating a method for generating a voice response to an input voice according to an embodiment of the present specification is shown;

[0043] Figure 4 A schematic diagram illustrating a method for generating a voice response to an input voice according to another embodiment of the present specification;

[0044] Figure 5 A schematic diagram illustrating a method for generating a voice response to an input voice according to another embodiment of the present specification;

[0045] Figure 6 A structural diagram of a device for generating a voice response to an input voice according to an embodiment of the present specification is shown. DETAILED DESCRIPTION

[0046] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments derived by those skilled in the art based on the embodiments in this specification without creative effort shall fall within the scope of protection of this specification.

[0047] A large language model (LLM) is a deep learning model based on natural language processing that is trained on large-scale text corpora and contains parameters in the hundreds of millions or more. Large language models are often used for tasks such as semantic understanding and common sense reasoning based on text data. However, this text-based model inherently lacks the ability to process voice signals and cannot directly implement voice input, understanding, and generation. A solution for implementing voice response based on a large model, by connecting a voice encoder, a large language model, and a voice decoder in series, can achieve basic responses to voice commands. Figure 1 A schematic diagram of a solution for implementing speech response based on a large model is shown. Figure 1 As shown, the speech encoder first extracts speech features from the input speech signal. The large language model then performs semantic understanding and reasoning based on these speech features, generating latent state representations. Finally, the speech decoder uses these latent states to gradually generate speech output using an autoregressive word-unit prediction method. This involves multiple rounds of progressive generation, with each round predicting and outputting only a single speech word-unit.

[0048] However, this solution also has the following problems: First, compared with text signals, speech signals have more complex temporal characteristics and higher information density. The same semantic content often requires several times more speech words to represent. The word-word prediction autoregressive decoding method suitable for generating text responses has difficulty in outputting speech words at a high speed, resulting in a low speed of generating speech responses and a significant problem of speech response lag. Second, speech usually has multi-level structural features (such as phonemes, syllables, semantics, etc.), and a single speech word usually lacks clear semantic representation capabilities. It often requires a coordinated combination of multiple words to express basic phonemes or semantic units. This makes it difficult for the word-word prediction autoregressive decoding method to effectively extract the complex structural representation inherent in the speech signal, resulting in poor quality of speech output.

[0049] In order to solve the above technical problems, the embodiments of this specification provide a method for generating a voice response for an input voice. Figure 2 A schematic diagram showing a method for generating a voice response to an input voice according to an embodiment of the present specification is shown. Figure 2 As shown, first, user input speech can be obtained, speech features can be obtained through a speech encoder, and the speech features can be input into a large language model to obtain speech response features output by the large language model. Next, the speech response features are input into a decoding network comprising a speech decoder and multiple speech predictors (e.g., speech predictor 1, speech predictor 2, speech predictor 3) connected in series. The speech decoder can generate the first word-meta feature based on the speech response features, and each speech predictor can output the next word-meta feature based on the word-meta feature output by the previous speech predictor / speech decoder. Furthermore, the output speech response can be generated based on the word-meta features output by the speech decoder and multiple speech predictors.

[0050] This method has the following advantages: First, it can perform multi-token prediction (Multi Token Prediction) on the output speech by connecting a speech decoder and N speech predictors in series. Compared with the current single-token prediction, it can significantly increase the rate of output speech tokens, thereby increasing the speed of speech decoding and the speed of generating speech responses, and reducing the delay of speech responses. Second, this method can also effectively obtain a structural representation of the speech signal by predicting multiple tokens with a dependent relationship in each prediction round, thereby improving the quality of speech responses generated based on the multiple tokens.

[0051] The detailed process of this method is further explained below. Figure 3 A flow chart of a method for generating a voice response for an input voice according to an embodiment of this specification is given. Figure 3 Said method comprises at least the following steps:

[0052] Step S301: obtaining a first speech input by a user, inputting the first speech into a speech encoder, and obtaining a first feature;

[0053] Step S303: obtaining a second feature based on the first feature by using a preset large language model;

[0054] Step S305: Obtaining N+1 first word-unit features arranged in sequence based on the second feature by sequentially connecting a speech decoder and N speech predictors, wherein the speech decoder outputs a first first word-unit feature among the N+1 first word-unit features based on the second feature, and the i-th speech predictor among the N speech predictors outputs an i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features;

[0055] Step S307: Generate a voice response based on the N+1 first word features.

[0056] First, in step S301, a first speech input by a user is obtained. The first speech is input into a speech encoder to obtain a first feature. The first speech is a speech signal input by the user. In different embodiments, the user inputting the speech may be a different specific user, and the first speech input by the user may be a speech for different specific purposes, which is not limited in this specification. In this step, the obtained first speech can be input into a speech encoder to obtain a speech feature (i.e., the first feature). In a specific embodiment, it can be expressed as: in, Indicates the first sound, z 1:T represents the first feature, and T represents the length of the input speech sequence. The speech encoder can be used to convert the speech signal into a corresponding speech feature representation. In different embodiments, the first feature can be obtained using different specific types of speech encoders, and this specification does not limit this.

[0057] Then, in step S303, a second feature can be obtained based on the first feature by presetting a large language model. As mentioned above, a large language model (LLM) refers to a deep learning model based on natural language processing that is pre-trained on a large text corpus and contains parameters at or above the billion level. In different embodiments, the preset large language model can be a large language model of different specific types or different neural network structures, and this specification does not limit this.

[0058] In different embodiments, the specific method of obtaining the second feature based on the first feature may be different. In one embodiment, the first feature can be input into a downsampling adapter to obtain a third feature, the dimension of the third feature is smaller than the first feature, and the third feature is input into a preset large model to obtain the second feature, such as Figure 4 In a specific embodiment, it can be expressed as: h 1:T =DownAdaptor(z 1:T ), h′ 1:T′ =LLM(h 1:T ), where z 1:T represents the first feature, DownAdaptor() represents the downsampling adapter, LLM() represents the large language model, h 1:T Represents the third feature, h′ 1:T′ represents the second feature, T represents the length of the input speech sequence, and T′ represents the length of the semantic sequence of the output speech response. The downsampling adapter can reduce the dimensionality of the speech features input to the large language model, thereby reducing the amount of token data input to the large language model. The specific neural network structure of the downsampling adapter can vary in different embodiments. In one example, it can consist of a single fully connected layer.

[0059] After obtaining the second feature, in one embodiment, the second feature can also be input into a text decoder to obtain a first text corresponding to the second speech. The first text, that is, the text corresponding to the speech response to the first speech, such as Figure 4 shown.

[0060] Next, in step S305, a speech decoder and N speech predictors connected in series may be used to obtain N+1 first word-unit features arranged in sequence based on the second feature. The speech decoder may output the first first word-unit feature among the N+1 first word-unit features based on the second feature. The i-th speech predictor among the N speech predictors may output the i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features. N may be a natural number greater than or equal to 1, and the specific value of N may vary in different embodiments.

[0061] In different embodiments, the specific method of obtaining N+1 first word-unit features arranged in sequence according to the second feature by sequentially connecting a speech decoder and N speech predictors can be different. In order to further improve the quality of the speech reply feature used to generate word-unit features in subsequent steps, in one embodiment, the second feature can also be input into a speech adapter to obtain a fourth feature; and by sequentially connecting a speech decoder and N speech predictors, N+1 first word-unit features arranged in sequence according to the fourth feature are obtained. In a specific embodiment, the fourth feature obtained by the speech adapter can be expressed as: h′1′ :T′ =TTSAdaptor(h′ 1:T′ ). Among them, TTSAdaptor() represents the voice adapter, h′ 1:T′ Represents the second feature, h′1′ :T′ represents the fourth feature, and T′ represents the semantic sequence length of the output voice response. In different embodiments, the specific type of the voice adapter can be different. In one embodiment, the voice adapter can be, for example, a Transformer neural network pre-trained with voice training data.

[0062] The speech decoder and speech predictor can be used to further extract features from the input speech features and generate predicted word-unit features. In different embodiments, the speech decoder and speech predictor can each have different neural network structures. In one embodiment, the speech decoder can be, for example, a multi-layer Transformer neural network, and the speech predictor can be, for example, a single-layer Transformer neural network.

[0063] In a specific embodiment, the speech decoder may output the first first word-unit feature among the N+1 first word-unit features and the first processing feature based on the second feature; the i-th speech predictor among the N speech predictors may output the third processing feature and the i+1-th first word-unit feature based on the second processing feature output by the previous speech decoder or speech predictor and the i-th first word-unit feature among the N+1 first word-unit features. In the above example of generating the fourth feature based on the second feature, the process can be expressed as:

[0064]

[0065] Among them, SpeechDecoder() represents the speech decoder, MTP1() represents the first speech predictor, and MTP N () represents the Nth speech predictor, h′1′ :T′ represents the fourth feature, T′ represents the semantic sequence length of the output speech response; Indicates the preposition word. No preposition word is required for the first input. represents the processing characteristics of the speech decoder output, Represents the first word feature output by the speech decoder, represents the processed features output by the first speech predictor, represents the first word feature output by the first speech predictor, represents the processed features output by the Nth speech predictor, Represents the first word-unit feature output by the Nth speech predictor.

[0066] Since each speech predictor outputs the processing features and predicted word features of the current speech predictor based on further feature extraction processing of the processing features output by the previous speech predictor (or speech decoder), the first word feature of N+1 stores the temporal dependency and the dependency of the feature processing level, thereby not only improving the speed of speech decoding and the speed of generating speech responses by predicting the output of multiple word features in each round, but also effectively obtaining the structural representation of the speech signal through multiple words with dependencies, thereby improving the quality of the generated speech responses.

[0067] Figure 5 FIG. 1 is a schematic diagram showing a method for generating a voice response to an input voice according to another embodiment of the present specification. Figure 5 As shown, in the decoding network, the speech decoder and multiple speech predictors (e.g., speech predictor 1, speech predictor 2, and speech predictor 3) are connected in series. The speech feature H0 (the fourth feature) output by the speech adapter can be input into the speech decoder. The speech decoder generates processing feature H1 and word-unit feature h1, and transmits processing feature H1 and word-unit feature h1 to the next speech predictor 1. Speech predictor 1 can generate processing feature H2 and word-unit feature h2 based on processing feature H2 and word-unit feature h2... Similarly, speech predictor 1 can send processing feature H2 and word-unit feature h2 to the next speech predictor 2. Speech predictor 2 can generate processing feature H3 and word-unit feature h3 based on processing feature H2 and word-unit feature h2, and send these to the next speech predictor 3. Speech predictor 3 can generate word-unit feature h4 based on processing feature H3 and word-unit feature h3.

[0068] Thereafter, in step S307, a speech response may be generated based on the N+1 first word-meta features. In one embodiment, a classification head may be used to generate second word-meta features based on the first word-meta features, and a speech response may be generated based on the second word-meta features. Since the first word-meta features may include semantic features of the predicted word-meta, the second word-meta features obtained by the classification head may indicate the classification result of the predicted word-meta.

[0069] In one specific embodiment, the second word-unit feature can be input into a speech generator to obtain a speech response. The speech generator can synthesize a speech signal based on the word-unit classification result indicated by the second word-unit feature, and output the speech signal as the speech response. In different specific embodiments, the speech response can be generated using different specific types of speech generators, and this specification does not limit this.

[0070] In the above, the first word feature is generated In the example, generating the second word feature based on the first word feature can be expressed as:

[0071]

[0072] Among them, OutHead0(), OutHead1()...OutHead N () represents the classification head, SpeechVocoder() represents the speech generator, t represents the predicted initial position, Represents the first word feature output by the speech decoder, Indicates based on The generated second word feature (indicating the classification result), represents the first word feature output by the first speech predictor, Indicates based on The generated second word feature, represents the first word feature output by the Nth speech predictor, Indicates based on The generated second word feature, Indicates the output voice response.

[0073] exist Figure 5 In the example shown, the speech decoder can input word-unit feature h1 (the first word-unit feature) into classification head 1 to obtain word-unit feature p1. Similarly, speech predictor 1 can input word-unit feature h2 into classification head 1 to obtain word-unit feature p2. Speech predictor 2 can input word-unit feature h2 into classification head 2 to obtain word-unit feature p2. Speech predictor 3 can input word-unit feature h3 into classification head 3 to obtain word-unit feature p3. Speech predictor 4 can input word-unit feature h4 into classification head 4 to obtain word-unit feature p4. Word-unit features p1, p2, p3, and p4 can then be input into the speech generator to obtain the output speech response. In a specific example, the speech response can also be sent to the user.

[0074] In actual production scenarios, voice responses can be generated through multiple rounds of word prediction. In one embodiment, in each round of multiple rounds of word prediction, N+1 predicted word features can be generated by a speech decoder and N speech predictors, and the N+1 predicted word features are input into the speech generator to obtain a voice response segment and return it to the user. The next round of word prediction can start from the position of the next word predicted in the previous round, further generate N+1 predicted word features, and input the newly generated N+1 predicted word features into the speech generator to obtain a new voice response segment and return it to the user. In this way, with N+1 words as the word length of a word prediction, voice response segments can be continuously generated and returned to the user until the entire voice response is output.

[0075] In addition, in one embodiment, the speech decoder and N speech predictors can be pre-trained through the following process: obtaining a second speech input by a user, inputting the second speech into a speech encoder to obtain a third feature; obtaining a fourth feature based on the third feature by presetting a large language model; obtaining N+1 third word-unit features arranged in sequence based on the fourth feature by the speech decoder and the N speech predictors, wherein the speech decoder outputs the first third word-unit feature among the N+1 predicted word-unit features based on the fourth feature, and the i-th speech predictor among the N speech predictors outputs the i+1-th third word-unit feature based on the i-th predicted word-unit feature among the N+1 third word-unit features; obtaining word-unit labels corresponding to the N+1 third word-unit features, and determining a first loss based on the N+1 third word-unit features and the word-unit labels; and updating the network parameters of the speech decoder and the N speech predictors based on the first loss.

[0076] In different embodiments, the specific method for determining the first loss may vary. In one embodiment, N+1 prediction losses may be determined based on each of the N+1 third word-unit features and the corresponding word-unit label, and the first loss may be determined based on the sum of the N+1 prediction losses. In one embodiment, the first loss is determined based on the weighted sum of the N+1 prediction losses, where the weight of each prediction loss is determined based on the position of the speech decoder or speech predictor that generates the third word-unit feature corresponding to the prediction loss in the speech decoder and the N speech predictors. In different specific embodiments, the specific method for determining the prediction loss based on the third word-unit feature and the corresponding word-unit label may vary. In one specific embodiment, the prediction loss may be determined using a cross-entropy loss function. In one example, in a speech decoder and N speech predictors connected in series, the weighted prediction loss corresponding to the word-unit feature extracted by the speech decoder or speech predictor with the first position is higher than the weighted prediction loss corresponding to the word-unit feature extracted by the speech decoder or speech predictor with the last position.

[0077] In a specific embodiment, the first loss is determined based on the sum of N+1 predicted losses and can be expressed as:

[0078]

[0079] Where L represents the first loss, N represents the number of speech predictors, k represents the sequence number of the speech predictor, and L sd represents the prediction loss determined based on the third word-unit feature output by the speech decoder and the corresponding word-unit label, represents the prediction loss determined based on the third word-unit feature output by the k-th speech predictor and the corresponding word-unit label, λ represents the attenuation coefficient, and λ∈(0,1), λ k represents λ to the power of k.

[0080] In one embodiment, the network parameters of the preset large language model may also be updated based on the first loss. The specific methods for updating the network parameters of the preset large language model may vary in different embodiments. In one embodiment, for example, all parameters of the first large model may be updated using a backpropagation algorithm based on the first loss. In another embodiment, some parameters of the first large model may be updated based on the first loss, such as one or more parameters of a newly added layer, newly added low-rank parameters, or parameters of some existing layers.

[0081] In the above-mentioned embodiment of further extracting features from the speech response output by the large language model through the speech adapter, and in the embodiment of reducing the dimension of the speech features input to the large language model through the downsampling adapter, the network parameters of the speech adapter and / or the downsampling adapter can also be updated based on the first loss.

[0082] According to yet another embodiment, a device for obtaining a voice response to a voice input is also provided. Figure 6 A structural diagram of a device for obtaining a voice response to a voice input according to an embodiment of this specification is shown. Figure 6 As shown, the apparatus 600 includes:

[0083] An acquiring unit 602 is configured to acquire a first speech input by a user, input the first speech into a speech encoder, and obtain a first feature;

[0084] A first processing unit 604 is configured to obtain a second feature based on the first feature by using a preset large language model;

[0085] The second processing unit 606 is configured to obtain N+1 first word-unit features arranged in sequence based on the second feature by sequentially connecting a speech decoder and N speech predictors, wherein the speech decoder outputs a first first word-unit feature among the N+1 first word-unit features based on the second feature, and an i-th speech predictor among the N speech predictors outputs an i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features;

[0086] The generating unit 608 is configured to generate a speech response based on the N+1 first word-unit features.

[0087] Another aspect of the embodiments of this specification provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute any one of the above methods.

[0088] On the other hand, the embodiments of this specification provide a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, any one of the above methods is implemented.

[0089] It should be understood that the descriptions such as “first” and “second” in this article are only used to distinguish similar concepts for the sake of simplicity of description and do not have any other limiting effect.

[0090] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0091] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0092] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the future development of computer technology, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0093] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flow charts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any particular order.

[0094] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more of the present specifications, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0095] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0096] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0098] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0099] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0100] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0101] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0102] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0103] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.

[0104] The foregoing description is merely an example of one or more embodiments of this specification and is not intended to limit the one or more embodiments of this specification. Those skilled in the art will appreciate that various modifications and variations of one or more embodiments of this specification are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification are intended to be included within the scope of the claims.

Claims

1. A method for generating a voice response to a voice input, comprising: Obtaining a first speech input by a user, and inputting the first speech into a speech encoder to obtain a first feature; Obtaining a second feature based on the first feature by presetting a large language model; Obtaining N+1 first word-unit features arranged in sequence based on the second feature by sequentially connecting a speech decoder and N speech predictors, wherein the speech decoder outputs a first first word-unit feature among the N+1 first word-unit features based on the second feature, and the i-th speech predictor among the N speech predictors outputs an i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features; Generate a speech response based on the N+1 first word-unit features.

2. The method according to claim 1, wherein The speech decoder outputs a first first word-unit feature among the N+1 first word-unit features based on the second feature, and the i-th speech predictor among the N speech predictors outputs an i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features, including: The speech decoder outputs a first first word-unit feature among the N+1 first word-unit features and a first processing feature based on the second feature; The i-th speech predictor among the N speech predictors outputs a third processed feature and the i+1-th first word unit feature based on the second processed feature output by the speech decoder or speech predictor arranged previously and the i-th first word unit feature among the N+1 first word unit features.

3. The method according to claim 1, wherein Generating a speech response based on the N+1 first word-unit features, including: A second word-unit feature is generated according to the first word-unit feature through a classification head, and a speech response is generated based on the second word-unit feature.

4. The method according to claim 3, wherein: Generate a voice response based on the second word feature, including: The second word-unit feature is input into a speech generator to obtain a speech response.

5. The method according to claim 1, wherein By presetting a large language model and based on the first feature, a second feature is obtained, including: The first feature is input into a downsampling adapter to obtain a third feature, where the dimension of the third feature is smaller than that of the first feature. The third feature is input into a preset large model to obtain a second feature.

6. The method according to claim 1, wherein By sequentially connecting the speech decoder and N speech predictors, sequentially arranged N+1 first word features are obtained based on the second feature, including: Inputting the second feature into the voice adapter to obtain a fourth feature; By sequentially connecting the speech decoders and N speech predictors, N+1 first word features arranged in sequence are obtained according to the fourth feature.

7. The method according to claim 1, further comprising: The second feature is input into a text decoder to obtain a first text corresponding to the second speech.

8. The method according to claim 1, wherein The speech decoder and N speech predictors are pre-trained by the following process: Obtaining a second voice input by a user, and inputting the second voice into a speech encoder to obtain a third feature; By presetting a large language model, a fourth feature is obtained based on the third feature; Obtaining, by the speech decoder and the N speech predictors, N+1 third word-unit features arranged in sequence based on the fourth feature, wherein the speech decoder outputs a first third word-unit feature among the N+1 predicted word-unit features based on the fourth feature, and the i-th speech predictor among the N speech predictors outputs an i+1-th third word-unit feature based on the i-th predicted word-unit feature among the N+1 third word-unit features; Obtain word-unit labels corresponding to the N+1 third word-unit features, determine a first loss based on the N+1 third word-unit features and the word-unit labels; and update network parameters of the speech decoder and N speech predictors based on the first loss.

9. The method according to claim 8, wherein Determining a first loss based on the N+1 third word features and the word tag includes: determining N+1 prediction losses based on each third word feature in the N+1 third word features and the corresponding word tag, and determining the first loss based on the sum of the N+1 prediction losses.

10. The method according to claim 9, wherein: Determining a first loss based on the sum of the N+1 prediction losses includes: determining the first loss based on a weighted sum of the N+1 prediction losses, wherein the weighted weight of each prediction loss is determined based on the arrangement position of the speech decoder or speech predictor that generates the third word-unit feature corresponding to the prediction loss in the speech decoder and the N speech predictors.

11. The method according to claim 8, further comprising: According to the first loss, network parameters of the preset large language model are updated.

12. A device for obtaining a voice response to a voice input, the device comprising: an acquiring unit configured to acquire a first speech input by a user, input the first speech into a speech encoder, and obtain a first feature; A first processing unit is configured to obtain a second feature based on the first feature by using a preset large language model; a second processing unit configured to obtain N+1 first word-unit features arranged in sequence based on the second feature by sequentially connecting a speech decoder and N speech predictors, wherein the speech decoder outputs a first first word-unit feature among the N+1 first word-unit features based on the second feature, and an i-th speech predictor among the N speech predictors outputs an i+1-th first word-unit feature based on the i-th first word-unit feature among the N+1 first word-unit features; The generation unit is configured to generate a speech response based on the N+1 first word-unit features.

13. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 11.

14. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 11 is implemented.