Voice large model reasoning method and device for long voice
By compressing and combining speech signals and dynamic training, the high computational cost and scarce data of speech large models in long speech processing are solved, and efficient long speech comprehension is achieved.
Patent Information
- Application Number
- CN202510356151.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
Existing speech models face the problems of scarce training data and high computational costs when processing long speech, resulting in limited exploration in long speech processing.
The speech training signal is encoded through the information extraction module, and the similarity and text content of adjacent frames are compressed and merged. Combined with the dynamic compression training scheme, the calculation cost of the model processing long speech input is reduced.
While maintaining performance, it significantly reduces the computational cost and time cost, improves long speech comprehension capabilities, and achieves efficient long speech processing.
Smart Images

Figure CN120260547A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language understanding and speech understanding, and particularly relates to a speech large model inference method, device, electronic device, computer-readable storage medium, and computer program product for long speech. Background Art
[0002] Speech interaction is a key focus in current artificial intelligence research. In recent years, thanks to the rapid development of large language models (LLMs), significant progress has also been made in speech large models (LSLMs). By extending the speech processing capabilities of LLMs, LSLMs can directly understand speech signals, perform analysis and inference, and excel in various tasks such as speech recognition, speech translation, and speech understanding. This ability to process and understand diverse speech signals has become the core focus of LSLMs research. Traditional methods typically adopt a cascaded processing framework, i.e., first transcribing speech into text and then processing it with LLMs. However, this method has the problem of error propagation and will lose valuable paralinguistic information, such as intonation information. To overcome these limitations, recent research has shifted to an end-to-end processing paradigm, enabling LSLMs to directly process speech signals and perform inference. These methods can be roughly divided into two categories: one aligns the output space of a pre-trained speech encoder with the embedding representation of LLMs to process speech inputs while retaining some capabilities of LLMs; the other discretizes speech into speech units, enabling LSLMs to process speech units in the same way as text units. However, due to the limitation of long speech training data, current LSLMs can usually only process short speech segments, typically with a duration of no more than 30 seconds.
[0003] The processing of long speech by speech large models remains an under-explored area. It mainly faces two major challenges. First, different from the rich and diverse short speech datasets, there is currently a lack of publicly available long speech alignment and instruction training data, and the cost of generating long speech data is relatively high. Second, long speech sequences pose huge computational requirements for language large models. For content with the same meaning, the sequence of speech representations is usually more than four times longer than its corresponding text sequence, resulting in higher computational costs. Therefore, speech large models face significant challenges in modeling long speech sequences, which come both from the scarcity of training data and the increase in computational costs. These challenges limit the exploration of speech large models in long speech processing.
[0004] Therefore, there is an urgent need to propose a method to expand the capabilities of language large models in the case of scarce long speech data, enabling them to efficiently process long speech inputs. Summary of the Invention
[0005] The objective of the present invention is to solve the problems faced by speech large models in processing long speeches, and a speech compression and training solution is proposed to enhance the ability of speech large models to process long speech inputs, which can greatly reduce the inference cost when the model processes long speech inputs while maintaining performance.
[0006] In view of the deficiencies of the prior art, such as Figure 4 shown, the present invention proposes a large model inference method for long speeches, which includes:
[0007] Initial step: Obtain a speech training signal with labeled training labels, encode the speech training signal through an information extraction module to obtain the original speech representation of the speech training signal, and compress and combine the original speech representation according to the text content and inter-frame similarity of the original speech representation to obtain a compressed speech representation;
[0008] Training step: Input the compressed speech representation into a large language model, perform an inference task to obtain an inference result corresponding to the speech training signal, and construct a loss function based on the inference result and the training label to train the information extraction module;
[0009] Inference step: Input a long speech signal into the trained information extraction module to obtain the compressed speech representation of the long speech signal, and input it into the large language model to obtain an inference result corresponding to the long speech signal.
[0010] The large model inference method for long speeches, wherein the initial step includes:
[0011] Similarity calculation step: Evaluate the similarity e j between adjacent speech frames h j+1 in the current original speech representation through the cosine similarity between adjacent speech frames h j,j+1 :
[0012]
[0013] Text content calculation step: Select the adjacent speech frames with the highest similarity in the current original speech representation as a speech segment, and for the i-th speech frame h j in the original speech representation in the speech segment, accumulate non-empty tags through the following formula to obtain its text content d j :
[0014]
[0015] Combination step: Combine all frames within the speech segment by using weighted average;
[0016] In the loop step, it is judged whether the length of the original speech representation after merging is less than or equal to the preset length. If so, the current original speech representation is saved as the compressed speech representation. Otherwise, the original speech representation is used as the current original speech representation to execute the similarity calculation step again.
[0017] The large model inference method for long speech as described above, wherein the loss function is:
[0018]
[0019] where L is a set of lengths of compressed speech representations, and IF(h, L) represents applying a compression operation to the original speech representation h to obtain a compressed speech representation with length L.
[0020] As Figure 5 shown, the present invention also proposes a large model inference device for long speech, which includes:
[0021] An initial module that obtains a speech training signal with a marked training label, encodes the speech training signal through an information extraction module to obtain the original speech representation of the speech training signal, and performs compression and combination on the original speech representation according to the text content and inter-frame similarity of the original speech representation to obtain a compressed speech representation;
[0022] A training module that inputs the compressed speech representation into a large language model, executes an inference task to obtain an inference result corresponding to the speech training signal, and constructs a loss function according to the inference result and the training label to train the information extraction module;
[0023] An inference module that inputs a long speech signal into the trained information extraction module to obtain a compressed speech representation of the long speech signal, and inputs it into the large language model to obtain an inference result corresponding to the long speech signal.
[0024] The large model inference device for long speech as described above, wherein the initial module includes:
[0025] A similarity calculation module that evaluates the similarity e of adjacent frames through the cosine similarity between adjacent speech frames h j and h j+1 in the current original speech representation: j,j+1 :
[0026]
[0027] A text content calculation module that selects the adjacent speech frame with the highest similarity in the current original speech representation as a speech segment, and accumulates non-empty markers through the following formula for the i-th speech frame h j in the original speech representation in the speech segment to obtain its text content d j :
[0028]
[0029] The merging module merges all the frames within the speech segment in a weighted average manner;
[0030] The loop module determines whether the length of the original speech representation after merging is less than or equal to a preset length. If so, it saves the current original speech representation as the compressed speech representation. Otherwise, it uses the original speech representation as the current original speech representation and executes the similarity calculation module again.
[0031] For the large model inference device for long speech, the loss function is as follows:
[0032]
[0033] where L is a set of lengths of compressed speech representations, and IF(h,L) represents applying a compression operation to the original speech representation h to obtain a compressed speech representation with length L.
[0034] The present invention also proposes a client for any one of the large model inference devices for long speech.
[0035] The present invention also proposes an electronic device, which includes one of the large model inference devices for long speech. The electronic device is either connected to an information display device, and the information display device is used to display the inference result with display parameters, attributes set by the user or through an artificial intelligence model.
[0036] The present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the large model inference method for long speech.
[0037] The present invention also proposes a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the large model inference method for long speech.
[0038] As can be seen from the above solutions, the advantages of the present invention are as follows:
[0039] For the short speech QA task solution, this method greatly exceeds other speech compression solutions at different compression ratios. The scores are evaluated by a large language model, with a full score of 5 points. For the long speech QA solution, this method also achieves the best generation performance while reducing 70% of the computational amount and 60% of the computational time. The experimental results of this method and other methods are listed in Figure 1 and Figure 3 .
[0040] In summary, the invention enhances the long speech understanding ability at a minimal cost, greatly reducing the inference cost and inference time while ensuring high generation quality. Description of the Drawings
[0041] Figure 1 It is a comparison chart of the response quality of the short speech QA task at different compression ratios;
[0042] Figure 2 It is a schematic diagram of the model process of the present invention;
[0043] Figure 3 It is a graph of the model response quality in the long speech task;
[0044] Figure 4 It is a flowchart of the method of the present invention;
[0045] Figure 5 It is a module diagram of the device of the present invention;
[0046] Figure 6 It is a schematic diagram of the structure of the first electronic device of the present invention;
[0047] Figure 7 It is a schematic diagram of the application environment structure of the first electronic device of the present invention;
[0048] Figure 8 It is a schematic diagram of the structure of the second electronic device of the present invention.
[0049] Reference Signs:
[0050] A - The first electronic device;
[0051] B - Large model inference device for long speech;
[0052] C - Data acquisition device;
[0053] D - Information display device;
[0054] 1000 - The second electronic device;
[0055] Ⅰ - Computing unit;
[0056] Ⅱ - ROM;
[0057] Ⅲ - RAM;
[0058] Ⅳ - Bus;
[0059] Ⅴ - Interface;
[0060] Ⅵ - Input unit;
[0061] Ⅶ - Output unit;
[0062] Ⅷ - Storage medium;
[0063] IX - Communication Unit. Detailed Implementation Manner
[0064] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0065] Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0066] The processor described in the present invention is the control center of an electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0067] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0068] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smartphones, tablets, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.
[0069] The memory is used to store the software program for implementing the solution of the present invention and is controlled by a processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0070] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not limit it. The actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or combine some components, or have different component arrangements.
[0071] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0072] It should also be understood that the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0073] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items (pieces)" or its similar expressions refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0074] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0075] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0076] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0077] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0078] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.
[0079] In order to enable the speech large model to process long language inputs, the present invention efficiently compresses the continuous speech representations generated by the speech encoder. This enables the sequence of speech representations to be reduced before being input into the large language model, thereby reducing the efficiency of the large language model in processing long speech inputs.
[0080] The inventors noticed that the large language model would not adapt to the compressed speech representations after long speech inputs were compressed, and the sequence length of the compressed speech representations was roughly similar to that of short speech inputs. For this reason, the inventors proposed a dynamic compression training scheme. Instead of using long speech training data, the model uses short speech data during training. By dynamically adjusting the compression ratio, the model not only adapts to the compressed representations but also maintains a sense of the length of longer speech representation sequences.
[0081] The present invention enables the speech large model to have the ability to understand and reason about long speech with a relatively small training cost. Through the compression scheme, the length of the long speech sequence is reduced, thereby reducing the cost of the speech large model in processing long speech inputs, and maintaining performance with the cooperation of the dynamic compression scheme. To achieve the above technical effects, the present invention proposes the following key technical points:
[0082] Key point 1: By evaluating the text content in the speech representation and comparing the similarity between adjacent frames, the speech large model can greatly reduce the speech representation sequence; while maintaining performance in short speech tasks, it reduces the inference cost by half in terms of the computational amount and inference time adopted.
[0083] Key point 2: On the basis of Key point 1, a dynamic compression training scheme is added to the speech large model to ensure that the speech large model adapts to speech representations with different compression ratios, thereby enhancing the long speech processing ability of the speech large model. On the basis of Key point 1, the performance of long speech understanding is ensured, and at the same time, compared with the speech large model that does not compress speech representations, the computational cost is reduced by 70% and the computational time is reduced by 60%, ensuring that the model can complete the response to a speech input of about 20 minutes in length within 1.5 seconds.
[0084] To make the above features and effects of the present invention more clearly understandable, specific embodiments are given below and are described in detail in conjunction with the accompanying drawings of the specification. The present specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0085] In order to enhance the long speech understanding ability of the speech large model, the present invention proposes a unified information extraction module and a dynamic compression training scheme. The information extraction module reduces the speech representation sequence through a compression strategy based on semantics and frame similarity. The dynamic compression training scheme enables the large language model to adapt to speech representations with different compression ratios, achieving a balance between the inference efficiency and the generation performance of the model.
[0086] like Figure 2 As shown, according to the processing capability of the speech encoder, given a speech signal of a specific Hz, for example, 16 kHz, s=(s1, s2, ..., s N ), N represents the Nth frame of the speech signal, and a speech frame refers to a moment of speech. The speech encoder will perform convolution and downsampling operations on the speech signal to generate a speech representation h = (h1, h2, ..., h J ), N is a multiple of J, h j It is a speech representation obtained by convolution and downsampling of multiple speech frames. The subsequent information extraction module will merge the speech representations based on similarity and text density (content). We first introduce the indicators for measuring text content and speech frame similarity.
[0087] For text content, we use the connection time classification CTC decoder to measure the text content. The CTC decoder predicts h j The text content represented by h j If no one speaks during this period, that is, there is no voice content, then h j Mark a j Empty, otherwise h j Mark a j Represents a speech word. For a speech frame h j , whose textual content can be obtained by accumulating non-empty tokens:
[0088]
[0089] where d j Represents h j The text content, p ctc Decode for CTC, a j For frame similarity, we use adjacent frames h j and h j+1 The cosine similarity between them is used to evaluate the similarity of adjacent frames:
[0090]
[0091] After introducing the completion metric, we introduce our information extraction module. Given a speech representation h of length T(m), we first determine the length of the speech representation after this round of compression:
[0092]
[0093] Where L represents the character length of the final desired compressed speech representation. The number of speech representations that need to be reduced in this round is calculated as follows:
[0094] r(m) = T(m) - T(m + 1)
[0095] Subsequently, we first use the similarity measurement index between adjacent frames to select the adjacent speech frames that are most similar to r(m). The recognized continuous speech frames are recognized as a segment. For each segment, we use the CTC decoder to recognize the text content of the speech frames in each segment, and use the weighted average method to merge all the frames within a segment into one frame.
[0096]
[0097] h j h ∈ Seg represents the speech frames belonging to a segment. This process will generate a speech representation of length T(m + 1). If T(m + 1) does not reach the length we preset, we will start the next fusion process. Otherwise, the current representation is the final compressed representation.
[0098] Subsequently, we introduce a dynamic compression training method. There are two considerations for introducing this method. First, it enables the large language model to adapt to compressed representations with different compression ratios. Second, the current <s, x, y> triple training data mainly contains short speech and short video segments, where x represents the question and y represents the answer. By sampling the length L of the compressed representation, our method can maintain the perception of the speech sequence corresponding to its window length without being overly biased towards overly compressed sequences. The dynamic compression training method for speech is as follows:
[0099]
[0100] where L dct is the loss function, L is the set of sampled final speech sequence lengths, IF(h, L) represents applying a compression operation to the speech representation to obtain a compressed representation of length L, and p represents the function of predicting y based on x, IF(h, L). After the training process, the model can effectively transfer the short speech ability and short video to the long speech task.
[0101] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To reduce repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0102] As Figure 5 shown, the present invention also proposes a large model inference device for long speech, which includes:
[0103] Initial module, which obtains a voice training signal with marked training labels, encodes the voice training signal through an information extraction module to obtain the original voice representation of the voice training signal, and compresses and combines the original voice representation based on the text content and inter-frame similarity of the original voice representation to obtain a compressed voice representation;
[0104] Training module, which inputs the compressed voice representation into a large language model, performs an inference task to obtain an inference result corresponding to the voice training signal, and constructs a loss function based on the inference result and the training label to train the information extraction module;
[0105] Inference module, which inputs a long voice signal into the trained information extraction module to obtain a compressed voice representation of the long voice signal, and inputs it into the large language model to obtain an inference result corresponding to the long voice signal.
[0106] The large model inference device for long voice, wherein the initial module includes:
[0107] Similarity calculation module, which evaluates the similarity e of adjacent frames through the cosine similarity between adjacent voice frames h j and h j+1 in the current original voice representation: j,j+1 :
[0108]
[0109] Text content calculation module, which selects the adjacent voice frames with the highest similarity in the current original voice representation as a voice segment, and accumulates non-empty marks for the i-th voice frame h j in the original voice representation in the voice segment to obtain its text content d j :
[0110]
[0111] Combination module, which combines all frames within the voice segment in a weighted average manner;
[0112] Loop module, which determines whether the length of the combined original voice representation is less than or equal to a preset length. If so, it saves the current original voice representation as the compressed voice representation. Otherwise, it uses the original voice representation as the current original voice representation and executes the similarity calculation module again.
[0113] The large model inference device for long voice, wherein the loss function is:
[0114]
[0115] Where L is a set of lengths of compressed speech representations, and IF(h, L) represents applying a compression operation to the original speech representation h to obtain a compressed speech representation of length L.
[0116] The present invention also proposes a client for any one of the large model inference devices for long speech.
[0117] As Figure 6 shown, in another embodiment, the present invention also proposes a first electronic device A, which includes the large model inference device for long speech described above.
[0118] As Figure 7 shown, the first electronic device A can also be connected to a data acquisition device C and an information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire long speech signals, and the information display device D is used to display the inference results obtained by the analysis of the present invention.
[0119] Among them, the information display device D can process and organize the data output by the first electronic device A based on an information display mechanism to improve the readability of the data output by the first electronic device A. This information display mechanism can be preset manually. For example, visualizing the data output by the first electronic device A, which can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. Present the key information specified by the user to the user, so that the user can understand this information more timely without having to access a secondary page or scroll the page, saving the user's operations. Or this information display mechanism can be an artificial intelligence AI display model, which can learn the key information of the user according to the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.
[0120] The present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the large model inference method for long speech provided by the above various methods.
[0121] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the large model inference method for long speech. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).
[0122] Figure 7 FIG. shows a schematic block diagram of a second electronic device 1000 that can be used to implement embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0123] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory II (ROM) or computer programs loaded from a storage medium VIII into a random access memory (RAM) III. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other through a bus IV. An input / output (I / O) interface V is also connected to the bus IV.
[0124] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a magnetic disk, an optical disc, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0125] The computing unit I can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit I executes the various methods and processes described above, such as method steps S1 - S3. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM II and / or the communication unit IX. When the computer program is loaded into the RAM III and executed by the computing unit I, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit I can be configured to execute the method in any other appropriate way (e.g., by means of firmware).
[0126] Although the embodiments of the present invention have been disclosed as above, it is not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrations shown and described herein.
Claims
1. A large model inference method for long speech, characterized in that, Including: Initial step: Obtain a voice training signal with labeled training labels, encode the voice training signal through an information extraction module to obtain the original voice representation of the voice training signal, and perform compression synthesis on the original voice representation according to the text content and inter-frame similarity of the original voice representation to obtain a compressed voice representation; Training step: Input the compressed voice representation into a large language model, perform an inference task to obtain an inference result corresponding to the voice training signal, and construct a loss function based on the inference result and the training label to train the information extraction module; Inference step: Input a long voice signal into the trained information extraction module to obtain a compressed voice representation of the long voice signal, and input it into the large language model to obtain an inference result corresponding to the long voice signal.
2. The large model inference method for long speech according to claim 1, wherein The initial step includes: Steps for calculating similarity, by using the cosine similarity between adjacent speech frames h j and h j+1 in the current original speech representation to evaluate the similarity e of adjacent frames j,j+1 : Text content calculation steps, select the adjacent speech frame with the highest similarity in the current original speech representation as the speech segment, and for the i-th speech frame h in the original speech representation of this speech segment j , obtain its text content d by accumulating non-empty tags through the following formula j : Merging step: Merge all frames within the voice segment in a weighted average manner; Loop step: Determine whether the length of the merged original voice representation is less than or equal to a preset length. If so, save the current original voice representation as the compressed voice representation. Otherwise, use the original voice representation as the current original voice representation and perform the similarity calculation step again.
3. The large model inference method for long speech according to claim 1, wherein, The loss function is: where L is a set of lengths of compressed voice representations, and IF(h, L) represents applying a compression operation to the original voice representation h to obtain a compressed voice representation of length L.
4. A large model inference device for long speech, characterized in that, Including: Initial module: Obtain a voice training signal with labeled training labels, encode the voice training signal through an information extraction module to obtain the original voice representation of the voice training signal, and perform compression synthesis on the original voice representation according to the text content and inter-frame similarity of the original voice representation to obtain a compressed voice representation; Training module: Input the compressed voice representation into a large language model, perform an inference task to obtain an inference result corresponding to the voice training signal, and construct a loss function based on the inference result and the training label to train the information extraction module; Inference module: Input a long voice signal into the trained information extraction module to obtain a compressed voice representation of the long voice signal, and input it into the large language model to obtain an inference result corresponding to the long voice signal.
5. The large model inference device for long speech according to claim 4, characterized in that The initial module includes: Similarity calculation module, which evaluates the similarity e of adjacent frames through the cosine similarity between adjacent speech frames h j and h j+1 in the current original speech representation j,j+1 : The text content calculation module selects the adjacent speech frame with the highest similarity in the current original speech representation as the speech segment, and accumulates the non-empty tags through the following formula to obtain its text content d j for the i-th speech frame h in the original speech representation of the speech segment j as follows: Merging module: Merge all frames within the voice segment in a weighted average manner; Loop module: Determine whether the length of the merged original voice representation is less than or equal to a preset length. If so, save the current original voice representation as the compressed voice representation. Otherwise, use the original voice representation as the current original voice representation and perform the similarity calculation module again.
6. The large model inference device for long speech according to claim 4, wherein The loss function is: where L is a set of lengths of compressed voice representations, and IF(h, L) represents applying a compression operation to the original voice representation h to obtain a compressed voice representation of length L.
7. A client for use in any one of the large model inference devices for long voices described in claims 4-6.
8. An electronic device, characterized in that, Including a large model inference device for long voices described in claims 4-6, the electronic device is either connected to an information display device, and the information display device is used to display the inference result with display parameters, attributes set by the user, or through an artificial intelligence model.
9. A computer-readable storage medium storing a computer program which, when executed by a processor, implements the steps of the large model inference method for long speech according to any one of claims 1-3.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the large model inference method for long speech according to any one of claims 1-3.