Multi-language speech recognition model optimization method and device, equipment and storage medium
By introducing low-rank decomposition matrix, fine-tuning and Tucker decomposition into the multilingual speech recognition model, the parameter structure of the model is optimized, and the performance bottleneck of the multilingual speech recognition model is solved when processing multilingual input, achieving more efficient and accurate recognition effects.
Patent Information
- Application Number
- CN202510308049.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
The existing multilingual speech recognition model faces performance bottlenecks when processing multilingual input, which is manifested as slow processing speed and high memory consumption, limiting its application on real-time applications and devices with limited resources.
By introducing a low-rank decomposition matrix and fine-tuning it, incremental weight tensors are generated and then Tucker decomposed, the parameter structure of the multilingual speech recognition model is optimized.
It significantly improves the speed and accuracy of the multilingual speech recognition model, reduces the word error rate, and provides an effective solution for practical applications.
Smart Images

Figure CN120164469A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi - language speech recognition model optimization, and particularly to a method, device, equipment and storage medium for optimizing a multi - language speech recognition model. Background Art
[0002] Existing speech recognition models, especially OpenAI's multi - language speech recognition models, face significant performance bottlenecks when processing multi - language inputs. Although these models perform excellently in terms of accuracy and can recognize the pronunciations and intonations of multiple languages, in practical applications, the problems of slow processing speed and high memory consumption severely restrict their real - time application capabilities. For example, in scenarios that require instant feedback, such as meeting recordings, online translation, or customer service, this delay may lead to a decline in the user experience and even affect the timeliness of decision - making.
[0003] Currently, multi - language speech recognition models usually require a large amount of computing resources to parse the features of different languages, resulting in a significant extension of their response time in a multi - language environment. In addition, the high memory consumption makes it difficult to run on devices with limited resources, which restricts their applications in mobile devices and edge computing environments. To solve these problems, an innovative technical solution is urgently needed that can significantly improve the efficiency of multi - language speech recognition while ensuring the accuracy of the model. Summary of the Invention
[0004] Based on this, in view of the XX technical problems of the prior art, a method, device, equipment and storage medium for optimizing a multi - language speech recognition model are proposed.
[0005] In the first aspect, a method for optimizing a multi - language speech recognition model is provided, and the method includes:
[0006] Obtain a multi - language speech recognition model to be adjusted, wherein a low - rank decomposition matrix is introduced into the parameter structure of the multi - language speech recognition model to be adjusted;
[0007] Fine - tune each language model in the multi - language speech recognition model, and generate an incremental weight tensor of the fine - tuned multi - language speech recognition model based on the fine - tuned multi - language speech recognition model;
[0008] Based on performing Tucker decomposition on the incremental weight tensor of the fine - tuned multi - language speech recognition model, obtain an optimized multi - language speech recognition model.
[0009] In the second aspect, a device for optimizing a multi - language speech recognition model is provided, and the device includes:
[0010] An acquisition module for acquiring a multilingual speech recognition model to be adjusted, wherein a low-rank decomposition matrix is introduced into the parameter structure of the multilingual speech recognition model to be adjusted;
[0011] A fine-tuning module for fine-tuning each language model in the multilingual speech recognition model and generating an incremental weight tensor of the fine-tuned multilingual speech recognition model based on the fine-tuned multilingual speech recognition model;
[0012] An optimization module for obtaining an optimized multilingual speech recognition model based on performing Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model.
[0013] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned multilingual speech recognition model optimization method are implemented.
[0014] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned multilingual speech recognition model optimization method are implemented.
[0015] The multilingual speech recognition model optimization method proposed by the present invention obtains a multilingual speech recognition model to be adjusted, wherein a low-rank decomposition matrix is introduced into the parameter structure of the multilingual speech recognition model to be adjusted, then fine-tunes each language model in the multilingual speech recognition model, and generates an incremental weight tensor of the fine-tuned multilingual speech recognition model based on the fine-tuned multilingual speech recognition model. Finally, an optimized multilingual speech recognition model is obtained based on performing Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model. The present invention can significantly improve the speed and accuracy of the multilingual speech recognition model, reduce the word error rate, and provide an effective solution for practical applications. Experimental results show that the model using the method of the present invention performs better than traditional models in a multilingual environment. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Among them:
[0018] Figure 1It is an application environment diagram of the multi - language speech recognition model optimization method in an embodiment;
[0019] Figure 2 It is a flowchart of the multi - language speech recognition model optimization method in an embodiment;
[0020] Figure 3 It is a structural block diagram of the multi - language speech recognition model optimization device in an embodiment;
[0021] Figure 4 It is a structural block diagram of a computer device in an embodiment;
[0022] Figure 5 It is a structural block diagram of a computer device in another embodiment. Detailed implementation manners
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above - mentioned drawings are intended to cover non - exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above - mentioned drawings are used to distinguish different objects and not to describe a specific order.
[0024] Reference to "an embodiment" in this context means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0026] The multi - language speech recognition model optimization method provided by the embodiments of the present invention can be applied, for example, in Figure 1In the application environment, the client 110 communicates with the server 120 through the network. The server 120 can optimize the multi-language speech recognition model proposed by the present invention through the client 110. First, obtain the multi-language speech recognition model to be adjusted. Among them, a low-rank decomposition matrix is introduced into the parameter structure of the multi-language speech recognition model to be adjusted. Then, fine-tune each language model in the multi-language speech recognition model, and generate an incremental weight tensor of the fine-tuned multi-language speech recognition model based on the fine-tuned multi-language speech recognition model. Finally, perform Tucker decomposition on the incremental weight tensor of the fine-tuned multi-language speech recognition model to obtain the optimized multi-language speech recognition model. The present invention can significantly improve the speed and accuracy of the multi-language speech recognition model, reduce the word error rate, and provide an effective solution for practical applications. Experimental results show that the model using the method of the present invention performs better than the traditional model in a multi-language environment. Among them, the client 110 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0027] Please refer to Figure 2 as shown in Figure 2 a flowchart of a multi-language speech recognition model optimization method provided by an embodiment of the present invention, including the following steps:
[0028] Step S101: Obtain the multi-language speech recognition model to be adjusted, where a low-rank decomposition matrix is introduced into the parameter structure of the multi-language speech recognition model to be adjusted;
[0029] Among them, the multi-language speech recognition model can be the Whisper model.
[0030] In one embodiment, the multi-language speech recognition model is initialized, the pre-trained Whisper model is loaded, and corresponding fine-tuning parameters are set for each target language. Then, the LoRA technology is used to introduce a trainable low-rank decomposition matrix, and each language is fine-tuned separately. By optimizing the parameters of the low-rank matrix, the performance of the model on a specific language is improved. When using the low-rank adaptation (LoRA) technology, the low-rank matrix reduces the number of parameters, thereby reducing the computational complexity while retaining the expressive power of the model. During the fine-tuning process, by gradually adjusting the parameters of the low-rank matrix, the specialized adaptation to each language is realized, ensuring the flexibility and accuracy of the model in a multi-language environment.
[0031] Specifically, introducing low-rank decomposition matrices into the parameter structure of the multilingual speech recognition model to be adjusted can be achieved by designing two trainable low-rank matrices A and B, which will be embedded into the model structure to capture the unique acoustic features of different languages. The introduction of low-rank matrices aims to effectively reduce the number of parameters required by the model, thereby reducing computational complexity and memory requirements. For example, for each language, the low-rank decomposition matrix can learn its unique pronunciation, prosody, and intonation features, enabling the model to make more accurate predictions when processing different languages. This strategy not only improves the flexibility of the model but also ensures its adaptability in a multilingual environment.
[0032] Step S102: Fine-tune each language model in the multilingual speech recognition model, and based on the fine-tuned multilingual speech recognition model, generate an incremental weight tensor of the fine-tuned multilingual speech recognition model;
[0033] In this embodiment, during the fine-tuning process, a loss function is set and the error between the model output and the actual label is calculated, and the parameters are updated based on the error. Finally, the performance of the fine-tuned model is evaluated, and the fine-tuning strategy is adjusted if necessary based on metrics such as the word error rate (WER).
[0034] Specifically, each language model is fine-tuned separately. During this process, the fine-tuning of each language will be adjusted according to its specific speech and language attributes to ensure that the Whisper model can efficiently process input data in multiple languages. For example, for different languages such as Chinese, English, and Spanish, the fine-tuning process will be customized according to their respective acoustic characteristics and related datasets. By optimizing the loss function, each fine-tuned language model can learn the pronunciation features of a specific language, thereby improving its performance in practical applications. In addition, this process also allows adjusting the learning rate and the choice of optimizer during training to further enhance the learning effect of the model. Set W i as the product of the low-rank matrix A i and B i . Mathematically, this process can be represented by the following formula: h i = W0x i + ΔW i x i = W0x i + B i A i x i . In this formula, i represents the index of the language, and i ∈ {1, 2,..., N - 1}. The goal of fine-tuning is to achieve the best mapping of the input signal x i by combining the output W i of the base model W0 and the low-rank adaptation module.
[0035] In the generation stage of the fine-tuned model output representation, the model combines the results of the base model and the low-rank adaptation module to form the best mapping of the input signal. During this process, the fine-tuned model will generate corresponding output representations according to the specific output requirements of each language. By combining the results of the base model and the low-rank adaptation module, it is ensured that the model can not only maintain high efficiency but also improve the overall accuracy when processing multilingual inputs. After fine-tuning, evaluate the performance of the model on different languages to ensure that it can effectively handle the speech and language attributes of each language. This step is to verify whether the fine-tuning is successful and further adjust the model parameters or fine-tuning strategy according to the evaluation results.
[0036] Step S103: Perform Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model to obtain an optimized multilingual speech recognition model.
[0037] In this embodiment, the weights and biases of each language model in the fine-tuned multilingual speech recognition model are integrated to obtain an incremental weight tensor. Then, perform Tucker decomposition on the incremental weight tensor to extract the core tensor and the corresponding projection matrices. This process helps to reduce the number of model parameters while retaining multilingual features. Next, integrate the obtained core tensor and projection matrices to construct a unified Tucker integration model. This model can not only effectively capture the shared features of different languages but also significantly reduce the storage requirements and computational complexity, ensuring the efficiency of the model. Tucker decomposition into a core tensor and the corresponding projection matrices effectively reduces the storage requirements and maintains the performance of the model. This process not only optimizes the memory usage but also enhances the model's ability to process different languages.
[0038] Specifically, it includes the following steps:
[0039] Step 1: Load the parameters of multiple language models in the fine-tuned multilingual speech recognition model, including weights and biases. These language models are fine-tuned by LoRA to adapt to specific tasks and datasets. Then, select the tensor decomposition method to be performed.
[0040] Step 2: If directly performing Tucker decomposition on the increment is selected, then: (1) directly load the weights of the models fine-tuned for N languages; (2) integrate the parameters of these models into an incremental weight tensor Here, N represents the number of language models, I represents the input dimension, and 0 represents the output dimension. The construction of this incremental weight tensor is to uniformly manage and optimize the parameters of different language models, enabling efficient subsequent decomposition and aggregation. (3) Select whether to use LoRA. If the answer is no, then do not use LoRA in the next tucker compression step: for the generated incremental weight tensor Perform incremental Tucker decomposition. The core of this step is to decompose the three-dimensional tensor into a core tensor and three corresponding projection matrices U (1) 、U (2) 、U (3) . Specifically, through incremental Tucker decomposition, we can represent the incremental weight tensor as: Here, is the core tensor is the projection matrix corresponding to the number of languages, and are the projection matrices corresponding to the input and output dimensions respectively. Through this decomposition, features in different dimensions can be effectively extracted, thereby reducing the complexity of the model.
[0041] If the choice is yes, then LoRA will be used in the next Tucker compression step: perform incremental Tucker non-compression decomposition on the generated incremental weight tensor . The core of this step is to decompose the three-dimensional tensor into a core tensor and three corresponding projection matrices P1, P2, and P3. Specifically, through incremental Tucker non-compression decomposition, we can represent the incremental weight tensor as: ΔW≈C×1P1×2P2×3P3. Here, is the core tensor, and P1, P2, and P3 are the projection matrices corresponding to the input and output dimensions respectively. Among them, the size of the core tensor is equal to the size of the incremental weight tensor . Apply the LoRA method to the core tensor and the three corresponding projection matrices P1, P2, and P3 respectively to reduce their ranks. This process will generate a low-rank version of the core tensor C′ and low-rank versions of the projection matrices P′1, P′2, P′3. By adopting LoRA, we can further extract important features in the model while reducing redundant parameters, thereby optimizing the performance and efficiency of the model. In this way, incremental Tucker decomposition not only retains the key information of the original tensor but also improves the flexibility and adaptability of the model in practical applications through rank reduction.
[0042] If the LoRA-Tucker decomposition method is selected, then: (1) Load the corresponding LoRA matrices A and B for each of the models fine-tuned for N languages; (2) Aggregate the LoRA matrices of all models into two three-dimensional LoRA weight tensors, namely and Among them, R represents the rank of LoRA, I represents the input dimension, and O represents the output dimension. In this way, the parameters after LoRA fine-tuning can be effectively managed, and a unified format can be provided for subsequent processing; (3) After generating the LoRA weight tensors A and B, perform Tucker decomposition on these two three-dimensional tensors to obtain their respective core tensors and corresponding projection matrices. The specific formulas are as follows: In these two formulas, and are the core tensors corresponding to the LoRA weight tensors respectively, and
[0043]
[0044]
[0045] Step 3: Regardless of the Tucker decomposition method used, integrate the Tucker decomposition parameters of all language models to form a comprehensive Tucker integrated model. This integrated model contains the Tucker decomposition parameters of all N language models, enabling the model to utilize knowledge from different languages simultaneously and improve the performance of multilingual tasks.
[0046] In one embodiment, the multilingual speech recognition model to be adjusted introduces a trainable low-rank decomposition matrix using LoRA technology.
[0047] In one embodiment, the incremental weight tensor is the integration of the weights and biases of each language model in the fine-tuned multilingual speech recognition model.
[0048] In one embodiment, after the step of obtaining the optimized multilingual speech recognition model by performing Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model, it includes:
[0049] Extract the core tensor and the corresponding projection matrix based on the Tucker decomposition of the incremental weight tensor of the fine-tuned multilingual speech recognition model;
[0050] Integrate the core tensor and the projection matrix to obtain the Tucker integrated model.
[0051] In one embodiment, the multilingual speech recognition model optimization method includes:
[0052] Step S201: Obtain a list of compression ratio factors;
[0052] Step S202: Traverse different combinations of compression factors in the compression ratio factor list, and calculate the word error rate for each combination of compression factors.
[0053] Step S203: Select the target combination of compression factors based on the word error rate, and apply the target combination of compression factors to the multilingual speech recognition model.
[0054] In this embodiment, an Adaptive Compression Path Module (ACPM) is designed and implemented to optimize the compression strategy of the model. The required total compression ratio and parameter k are received, and a compression ratio factor list is obtained using this information. Then, different combinations of compression factors are traversed, and the Word Error Rate (WER) is calculated for each combination to evaluate its performance. After evaluation, the best-performing compression combination is selected and applied to the model to ensure optimized model performance at a specific compression ratio. Finally, ACPM can flexibly adjust the compression path, maintaining the efficient operation of the model in different tasks while saving computational resources.
[0055] In the Adaptive Compression Path Module, ACPM monitors the performance metrics of each layer of the model, such as the word error rate and the number of model parameters, and dynamically selects the parameters that have the greatest impact on performance. This method ensures the flexibility of the model in different tasks and scenarios, adapts to different compression requirements, and thus achieves the best performance and storage balance.
[0056] Specifically, it includes the following steps:
[0057] Step 1: Receive the required total compression ratio and parameter k, that is, the maximum number of compression factors considered in each dimension.
[0058] Step 2: Use the total compression ratio for factorization to generate a factor list containing all possible compression factors. This factor list will be used in subsequent steps to facilitate the determination of the optimal compression configuration in different dimensions.
[0059] Step 3: Initialize the best compression ratio to one-fold compression, and the compression ratios for the three dimensions are 1, 1, 1, that is, no compression.
[0060] Step 4: The Adaptive Compression Path Module traverses the factor list and generates a new compression configuration for each factor combination. Specifically, the module generates a new list of compression ratios based on the current best compression configuration and the new factor combination.
[0061] Step 5: Calculate the corresponding word error rate for each new compression combination. This calculation process will evaluate the performance of the new configuration, aiming to find the configuration with the best performance.
[0062] Step 6: After evaluating all possible compression combinations, the adaptive compression path module selects the best-performing compression configuration from them. This configuration corresponds to the lowest word error rate (WER) and is recorded as the current best configuration. Finally, the module returns this best compression path for subsequent model optimization and compression decisions.
[0063] The present invention first obtains a multilingual speech recognition model to be adjusted, wherein a low-rank decomposition matrix is introduced into the parameter structure of the multilingual speech recognition model to be adjusted, and then each language model in the multilingual speech recognition model is fine-tuned, and based on the fine-tuned multilingual speech recognition model, an incremental weight tensor of the fine-tuned multilingual speech recognition model is generated. Finally, based on performing Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model, an optimized multilingual speech recognition model is obtained. The present invention can significantly improve the speed and accuracy of the multilingual speech recognition model, reduce the word error rate, and provide an effective solution for practical applications. Experimental results show that the model using the method of the present invention performs better than traditional models in a multilingual environment.
[0064] Please refer to Figure 3 As shown, in one embodiment, a multilingual speech recognition model optimization device is provided, and the device includes:
[0065] An acquisition module 10, configured to acquire a multilingual speech recognition model to be adjusted, wherein a low-rank decomposition matrix is introduced into the parameter structure of the multilingual speech recognition model to be adjusted;
[0066] A fine-tuning module 20, configured to fine-tune each language model in the multilingual speech recognition model, and generate an incremental weight tensor of the fine-tuned multilingual speech recognition model based on the fine-tuned multilingual speech recognition model;
[0067] An optimization module 30, configured to perform Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model to obtain an optimized multilingual speech recognition model.
[0068] In one embodiment, a computer device is provided, and the computer device may be a server, and its internal structure diagram may be as Figure 4As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multi-language speech recognition model optimization method.
[0069] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a multi-language speech recognition model optimization method.
[0070] In one embodiment, a computer device is proposed, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0071] Obtain a multi-language speech recognition model to be adjusted. Among them, a low-rank decomposition matrix is introduced into the parameter structure of the multi-language speech recognition model to be adjusted;
[0072] Fine-tune each language model in the multi-language speech recognition model, and based on the fine-tuned multi-language speech recognition model, generate an incremental weight tensor of the fine-tuned multi-language speech recognition model;
[0073] Based on performing Tucker decomposition on the incremental weight tensor of the fine-tuned multi-language speech recognition model, obtain an optimized multi-language speech recognition model.
[0074] The present invention can significantly improve the speed and accuracy of the multi-language speech recognition model, reduce the word error rate, and provide an effective solution for practical applications. Experimental results show that the model using the method of the present invention performs better than traditional models in a multi-language environment.
[0075] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0076] Obtain a multi-language speech recognition model to be adjusted, wherein a low-rank decomposition matrix is introduced into the parameter structure of the multi-language speech recognition model to be adjusted;
[0077] Fine-tune each language model in the multi-language speech recognition model, and based on the fine-tuned multi-language speech recognition model, generate an incremental weight tensor of the fine-tuned multi-language speech recognition model;
[0078] Based on performing Tucker decomposition on the incremental weight tensor of the fine-tuned multi-language speech recognition model, obtain an optimized multi-language speech recognition model.
[0079] The present invention can significantly improve the speed and accuracy of the multi-language speech recognition model, reduce the word error rate, and provide an effective solution for practical applications. Experimental results show that the model using the method of the present invention performs better than traditional models in a multi-language environment.
[0080] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0081] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0082] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0083] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A multilingual speech recognition model optimization method, characterized in that: The multilingual speech recognition model optimization method comprises: Acquire a multilingual speech recognition model to be adjusted, wherein a low-rank decomposition matrix is introduced into a parameter structure of the multilingual speech recognition model to be adjusted; Fine-tune each language model in the multilingual speech recognition model, and generate an incremental weight tensor of the fine-tuned multilingual speech recognition model based on the fine-tuned multilingual speech recognition model; Based on Tucker decomposition of the incremental weight tensor of the fine-tuned multilingual speech recognition model, an optimized multilingual speech recognition model is obtained.
2. The multilingual speech recognition model optimization method according to claim 1, characterized in that: The multilingual speech recognition model to be adjusted adopts LoRA technology to introduce a trainable low-rank decomposition matrix.
3. The multilingual speech recognition model optimization method according to claim 1, characterized in that: The incremental weight tensor is an integration of the weights and biases of each language model in the multilingual speech recognition model after fine-tuning.
4. The multilingual speech recognition model optimization method according to claim 1, characterized in that: The step of performing Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model to obtain an optimized multilingual speech recognition model includes: Extracting a core tensor and a corresponding projection matrix based on Tucker decomposition of the incremental weight tensor of the fine-tuned multilingual speech recognition model; The core tensor and projection matrix are integrated to obtain the Tucker integrated model.
5. The multilingual speech recognition model optimization method according to claim 1, characterized in that: The multilingual speech recognition model optimization method comprises: Get the compression factor list; Traversing different compression factor combinations in the compression factor list, and calculating the word error rate for each compression factor combination; A target compression factor combination is selected based on word error rate and applied to a multilingual speech recognition model.
6. A multilingual speech recognition model optimization device, characterized in that: The multilingual speech recognition model optimization device comprises: An acquisition module, used for acquiring a multilingual speech recognition model to be adjusted, wherein a low-rank decomposition matrix is introduced into a parameter structure of the multilingual speech recognition model to be adjusted; A fine-tuning module, configured to fine-tune each language model in the multilingual speech recognition model, and generate an incremental weight tensor of the fine-tuned multilingual speech recognition model based on the fine-tuned multilingual speech recognition model; The optimization module is used to perform Tucker decomposition on the incremental weight tensor of the fine-tuned multilingual speech recognition model to obtain an optimized multilingual speech recognition model.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the multilingual speech recognition model optimization method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multilingual speech recognition model optimization method according to any one of claims 1 to 5 are implemented.