Incremental Tucker decomposition method and device for multi-language automatic speech recognition
The core tensor and language factor matrix are updated through the incremental Tucker decomposition method, which solves the memory and computing overhead of multilingual automatic speech recognition system when expanding new languages, and realizes efficient language expansion and resource utilization.
Patent Information
- Application Number
- CN202510294328.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-10
AI Technical Summary
When expanding new languages, multilingual automatic speech recognition systems face increased memory and computing overhead, and the prior art is difficult to dynamically expand languages without recalculating the entire model decomposition.
The incremental Tucker decomposition method is used to update the initial core tensor and language factor matrix, calculate the incremental weight and project it into the compressed space, and generate a multilingual automatic speech recognition model containing a new language.
It reduces memory consumption, reduces computing overhead, improves the scalability and incremental learning ability of the system, and can effectively support the dynamic language expansion of large-scale multilingual speech recognition systems.
Smart Images

Figure CN120126458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and particularly to an incremental Tucker decomposition method and device for multilingual automatic speech recognition. Background Art
[0002] The goal of a multilingual automatic speech recognition (ASR) system is to transcribe or translate speech data in multiple languages through a unified model, and it has been widely applied in many fields such as global communication, media production, and cross - language education. With the continuous progress of speech recognition technology, although multilingual ASR systems have achieved remarkable results, the challenges they face still include how to efficiently expand new languages on the basis of existing models while maintaining the scalability, performance, and resource utilization efficiency of the system. Especially in resource - constrained environments, how to achieve the rapid expansion of large - scale language models and maintain good speech recognition accuracy is a key issue.
[0003] Therefore, a Tucker - based LoRA integration method is proposed to enhance the multilingual automatic speech recognition (ASR) model by integrating low - rank adaptation (LoRA) and Tucker decomposition. LoRA enables efficient cross - language fine - tuning by optimizing low - rank matrices, but its memory requirements grow linearly with the number of tasks. Tucker decomposition solves this problem by compressing the aggregated LoRA parameters into a compact representation, thereby reducing memory and computational overhead. Although Tucker decomposition improves efficiency, it has poor flexibility for dynamic expansion and requires recalculation when adding new tasks.
[0004] To overcome these deficiencies, the present application proposes an incremental Tucker decomposition method and device for multilingual automatic speech recognition. Summary of the Invention
[0005] The purpose of the present application is to provide an incremental Tucker decomposition method and device for multilingual automatic speech recognition, aiming to solve the above problems.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] In a first aspect, the present application provides an incremental Tucker decomposition method for multilingual automatic speech recognition, and the steps include:
[0008] Performing Tucker decomposition on the automatic speech recognition model parameters of at least one known language to obtain an initial core tensor and language factor matrices;
[0009] Based on obtaining an added new language, calculating incremental weights according to the automatic speech recognition model parameters of the new language;
[0010] Using an incremental update method, project the incremental weights into a compressed space to obtain incremental core vectors; wherein, the compressed space is defined by the initial core tensor and the language factor matrix;
[0011] Generate a multi - language automatic speech recognition model including the new language through the incremental core vectors; use the multi - language automatic speech recognition model to transcribe or translate speech data including the new language.
[0012] In a second aspect, the present application provides an incremental Tucker decomposition device for multi - language automatic speech recognition, specifically including:
[0013] Model construction module: Perform Tucker decomposition on the parameters of an automatic speech recognition model of at least one known language to obtain an initial core tensor and a language factor matrix;
[0014] Language integration module: Based on obtaining an added new language, calculate incremental weights according to the parameters of the automatic speech recognition model of the new language; use an incremental update method to project the incremental weights into a compressed space to obtain incremental core vectors; wherein, the compressed space is defined by the initial core tensor and the language factor matrix;
[0015] Model reconstruction module: Generate a multi - language automatic speech recognition model including the new language through the incremental core vectors;
[0016] Speech recognition module: Use the multi - language automatic speech recognition model to transcribe or translate speech data including the new language;
[0017] Storage module: Used to store the initial core tensor, the language factor matrix, and the parameters required for calculating incremental weights; wherein, the core tensor and the language factor matrix are obtained after performing Tucker decomposition on the parameters of an automatic speech recognition model of at least one known language, and are respectively used to capture shared structural information in the model parameters and encode language - specific features and shared features.
[0018] In a third aspect, the present application provides a computer device, the computer device includes a processor and a memory coupled to the processor, wherein, the memory stores program instructions for implementing an incremental Tucker decomposition method for multi - language automatic speech recognition; the processor is used to execute the program instructions stored in the memory to implement an incremental Tucker decomposition for multi - language automatic speech recognition.
[0019] In a fourth aspect, the present application provides a storage medium, storing program instructions executable by a processor, the program instructions being used to execute an incremental Tucker decomposition method for multi - language automatic speech recognition.
[0020] The present application provides an incremental Tucker decomposition method and apparatus for multilingual automatic speech recognition, which has the following beneficial effects:
[0021] (1) By introducing incremental Tucker decomposition, the model is incrementally updated without recalculating the entire model decomposition, reducing memory consumption; especially when adding a new language, instead of recalculating the entire model, only the core tensor and the language factor matrix need to be updated. As the number of languages increases, the memory consumption is significantly controlled, improving the scalability of the multilingual automatic speech recognition model.
[0022] (2) When adding a new language, the new language is integrated by incrementally updating the core tensor and the language factor matrix, avoiding repeated calculations every time a new language is extended, thereby significantly reducing the computational overhead and improving efficiency and flexibility.
[0023] (3) This application combines LoRA and Tucker decomposition. Each new language is integrated into the model through low-rank matrices and incremental Tucker updates, reducing the memory and computational requirements for incremental expansion, and also avoiding the problems of repeated training and global updates in traditional methods. Description of the Drawings
[0024] Figure 1 It is a schematic flowchart of an incremental Tucker decomposition method for multilingual automatic speech recognition according to Embodiment 1 of the present application;
[0025] Figure 2 It is a schematic structural diagram of an incremental Tucker decomposition apparatus for multilingual automatic speech recognition according to Embodiment 2 of the present application;
[0026] Figure 3 It is a schematic structural diagram of a computer device according to Embodiment 3 of the present application;
[0027] Figure 4 It is a schematic structural diagram of a storage medium according to Embodiment 4 of the present application. Detailed Embodiments
[0028] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0029] The following analyzes the solutions in the prior art in combination with related technologies.
[0030] In the prior art, incremental learning has been proposed to solve the problem of expanding multilingual models. For example, MoEx-tend proposed a modular method that allows adding new tasks without forgetting existing tasks. Although these methods can effectively support incremental updates, they still fail to solve the problems of parameter and memory growth in multilingual ASR.
[0031] LoRA is the most widely used low-rank adaptation technique at present. It can update the parameters of the model through low-rank matrices without retraining the entire model. The LoRA method is particularly suitable for incremental learning of multilingual ASR models because it allows training a low-rank adaptation module for each language, thus avoiding the need to train the entire model for each new language. However, the memory consumption of LoRA grows linearly with the number of languages, which limits its application in large-scale multilingual scenarios. Tucker decomposition is an efficient tensor compression technique that reduces the number of parameters and maintains the expressive power of the model by decomposing a multi-dimensional tensor into a core tensor and multiple factor matrices. Tucker decomposition can effectively compress the model, but it requires recomputing the entire decomposition when integrating new languages, which incurs high computational overhead in a dynamic multilingual environment.
[0032] The following significant drawbacks exist in the existing technologies: Although LoRA can reduce the amount of parameter updates through low-rank matrices and improve training efficiency, when the number of languages increases, the memory requirement grows linearly. This means that for each additional language module, the memory consumption of the model increases, which becomes non-scalable in large-scale multilingual ASR systems. Secondly, Tucker decomposition is an effective tensor compression technique, but a significant drawback is that when new languages need to be added, the entire decomposition of the model must be recomputed, which not only increases the computational overhead but also leads to inefficiency, especially in scenarios where languages need to be added dynamically. Existing multilingual ASR systems often need to retrain the entire model or adjust all parameters when supporting new languages, resulting in huge resource consumption when scaling to new languages. This approach not only reduces the efficiency of the system but also makes the operation complex and computationally expensive when dynamically expanding languages. At the same time, many existing methods provide the ability of incremental learning, allowing new languages to be added to a trained model, but these methods often cannot effectively handle the memory consumption problem after the addition of new language modules. Especially in low-resource environments, how to maintain efficient incremental learning without retraining the entire model is an important challenge.
[0033] Therefore, this application proposes an incremental Tucker decomposition method and device for multilingual automatic speech recognition, effectively solving the memory consumption and computational overhead problems of traditional methods, while improving the scalability and incremental learning ability of the system, providing a more efficient and flexible solution for large-scale, multilingual automatic speech recognition.
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0035] Embodiment 1
[0036] Please refer to Figure 1 , which is a schematic flowchart of an incremental Tucker decomposition method for multi-language automatic speech recognition in Embodiment 1 of the present application; the steps include:
[0037] S1: Perform Tucker decomposition on the automatic speech recognition model parameters of at least one known language to obtain an initial core tensor and language factor matrices.
[0038] In this embodiment, first, for the automatic speech recognition models of at least one known language, low-rank adaptation methods are used to generate low-rank matrix update parameters for each language. The low-rank adaptation method fine-tunes the automatic speech recognition model by introducing low-rank matrices B and A to adjust the weights ΔW of the pre-trained model; the formula is expressed as:
[0039] ΔW = BA,
[0040] where B ∈ R d×r , A ∈ R r×k , and r is the rank of the low-rank matrix;
[0041] For language i, the fine-tuned parameters are expressed as:
[0042] h i = W 0 x i + ΔW i x i = W 0 x i + B i A i x i ,
[0043] where x i is the input of language i; ΔW i = B i A i ; h i is the output of language i after calculation in the model layer; W 0 is the original weight of the model layer.
[0044] Specifically, let the original weight matrix of the automatic speech recognition model of the known language be W 0 Rd×k , for language i, a low-rank adaptation method is used to generate two low-rank matrices B∈R for each language d×r and A∈R r×k , where the rank r = 8, which can be adjusted according to the task.
[0045] The weight update of language i is calculated by ΔW=BA. Assuming d=768, k=768, and r=8, the total number of parameters of ΔW is 2*768*8=12288, which significantly reduces the memory requirement compared to the full parameter fine-tuning of 768*768=589824.
[0046] Combining the low-rank adaptation method LoRA with Tucker decomposition not only utilizes the low-rank adaptability of LoRA, but also reduces memory consumption with the compression capability of Tucker decomposition. Specifically, in the Tucker decomposition process, the automatic speech recognition model parameters of at least one known language are subjected to Tucker decomposition to obtain an initial core tensor and a language factor matrix. The initial core tensor is used to capture the shared structural information in the automatic speech recognition model parameters; the language factor matrix is used to encode language-specific features and shared features, including but not limited to input dimensions, output dimensions, and language-specific feature matrices.
[0047] Tucker decomposition decomposes the multilingual weight tensor ΔW into:
[0048] ΔW=C× 1 U (1) × 2 U (2) × 3 U (3) ,
[0049] Among them, C is the initial core tensor, U (1) , U (2) , U (3) They are language factor matrix, input factor matrix, and output factor matrix respectively.
[0050] S2: Based on the acquisition of the added new language, the incremental weight is calculated according to the automatic speech recognition model parameters of the new language.
[0051] In this embodiment, whenever a new language is added, the new language module is effectively integrated into the existing model through incremental Tucker decomposition without recalculating the entire Tucker decomposition. Specifically, by updating the core tensor and language factor matrix, the compressed representation of the existing language is maintained, and the new language is effectively integrated, which significantly improves resource utilization efficiency while maintaining good recognition accuracy.
[0052] For the newly acquired language, the incremental weight ΔW inc The expression is:
[0053] ΔW inc = U (2) × 2 C inc × 3 U (3) ,
[0054] wherein, × 2 and × 3 are tensor-matrix multiplications along the input mode and the output mode respectively; C inc is the incremental core vector.
[0055] S3: Using an incremental update method, project the incremental weights into the compression space to obtain an incremental core vector; wherein, the compression space is defined by the initial core tensor and the language factor matrix.
[0056] In this embodiment, project the incremental weights of the new language into the compression space to calculate the incremental core vector; the formula is expressed as:
[0057] C inc ≈ U (2)+ × 2 ΔW inc × 3 U (3)+ ,
[0058] wherein, c inc is the incremental core vector; ΔW inc is the incremental weight; U (2)+ and U (3)+ are the pseudo-inverses of the language factor matrix.
[0059] After obtaining the incremental core vector, it further includes: adjusting the language factor matrix for expanding to the language space including the new language to ensure the unification and compression of the parameters of the new language and the existing languages.
[0060] The updated language factor matrix is expressed as:
[0061]
[0062] wherein, I n2 is the identity matrix corresponding to the new language; is the updated language factor matrix.
[0063] S4: Generate a multi-language automatic speech recognition model including the new language through the incremental core vector; transcribe or translate the speech data including the new language using the multi-language automatic speech recognition model.
[0064] In this embodiment, the weight tensor is reconstructed through the inverse process of Tucker decomposition by means of the incremental core vector and the updated language factor matrix; the reconstructed weight tensor contains information of all known languages and new languages, and a multi-language automatic speech recognition model including the new language is generated;
[0065] The reconstructed weight tensor is expressed as:
[0066]
[0067] where C inc is the incremental core vector, and is the updated language factor matrix.
[0068] Based on this, this application is verified on the Mozilla Common Voice dataset, and the experimental results are as follows: Compared with the traditional LoRA method, this application reduces the parameter requirements by up to 80%, and can effectively integrate new language modules without the need to recalculate the complete decomposition, significantly reducing the memory burden. Through experiments on 52 languages, the incremental Tucker framework can be effectively extended to more languages without significantly increasing the number of parameters or reducing the performance. For example, when expanding from 32 languages to 52 languages, the WER of the incremental Tucker model only increases slightly, and at the same time, the number of parameters only increases slightly, while the LoRA incremental model needs to significantly increase the number of parameters to maintain a similar WER performance. For low-resource languages, this application shows higher recognition accuracy than traditional methods and can effectively reduce the error rate. Therefore, the method proposed in this application can not only improve the parameter efficiency of the multi-language ASR system, but also be comparable to traditional methods in terms of performance, and can flexibly adapt to the integration of new language modules.
[0069] In summary, Embodiment 1 of this application combines the low-rank adaptability of LoRA and the compression ability of Tucker decomposition, significantly reducing the memory requirements of the multi-language ASR model, and at the same time avoiding the problem of linear growth of memory requirements as the number of languages increases. By incrementally updating the core tensor and the language factor matrix, new languages are efficiently integrated into the existing Tucker decomposition, avoiding the computational overhead of retraining the entire model, so that when adding new languages, it will not cause a sharp increase in computational and storage requirements, thus achieving more efficient language expansion.
[0070] Embodiment 2
[0071] Please refer to Figure 2 , which is a schematic structural diagram of an incremental Tucker decomposition device for multi-language automatic speech recognition according to Embodiment 2 of this application; the specific content includes:
[0072] Model construction module: Perform Tucker decomposition on the parameters of an automatic speech recognition model for at least one known language to obtain an initial core tensor and language factor matrices;
[0073] Language integration module: Based on the newly added language obtained, calculate incremental weights according to the parameters of the automatic speech recognition model for the new language; Use an incremental update method to project the incremental weights into a compressed space to obtain incremental core vectors; wherein, the compressed space is defined by the initial core tensor and the language factor matrices;
[0074] Model reconstruction module: Generate a multi-language automatic speech recognition model including the new language through the incremental core vectors;
[0075] Speech recognition module: Use the multi-language automatic speech recognition model to transcribe or translate speech data including the new language;
[0076] Storage module: Used to store the initial core tensor, the language factor matrices, and the parameters required for calculating incremental weights; wherein, the core tensor and the language factor matrices are obtained after performing Tucker decomposition on the parameters of an automatic speech recognition model for at least one known language, and are respectively used to capture the shared structure information in the model parameters and encode language-specific features and shared features.
[0077] Embodiment 3
[0078] Please refer to Figure 3 , which is a schematic structural diagram of a computer device according to Embodiment 3 of this application. The computer device 50 includes a processor 51 and a memory 52 coupled to the processor 51.
[0079] The memory 52 stores program instructions for implementing the above-mentioned incremental Tucker decomposition method for multi-language automatic speech recognition.
[0080] The processor 51 is used to execute the program instructions stored in the memory 52 to implement an incremental Tucker decomposition for multi-language automatic speech recognition.
[0081] Among them, the processor 51 can also be called a CPU (Central Processing Unit, central processing unit).
[0082] The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0083] Example 4
[0084] Please refer to Figure 4 , which is a schematic structural diagram of the storage medium according to Example 4 of the present application. The storage medium of the embodiment of the present application stores a program file 61 capable of implementing all the above methods. Among them, the program file 61 can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes, or devices such as computers, servers, mobile phones, and tablets.
[0085] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, device, article or method. Without further limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, device, article or method including that element.
[0086] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
[0087] Although the embodiments of the present application have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.
[0088] Of course, the present invention can also have other various embodiments. Based on this embodiment, other embodiments obtained by those of ordinary skill in the art without any creative work belong to the scope protected by the present invention.
Claims
1. An incremental Tucker decomposition method for multilingual automatic speech recognition, characterized in that: include: Perform Tucker decomposition on the parameters of the automatic speech recognition model of at least one known language to obtain an initial core tensor and a language factor matrix; Based on acquiring the added new language, calculating the incremental weight according to the automatic speech recognition model parameters of the new language; Using an incremental update method, projecting the incremental weight into a compressed space to obtain an incremental core vector; wherein the compressed space is defined by the initial core tensor and the language factor matrix; A multilingual automatic speech recognition model including a new language is generated by using the incremental core vector; and the multilingual automatic speech recognition model is used to transcribe or translate speech data including the new language.
2. The incremental Tucker decomposition method for multilingual automatic speech recognition according to claim 1, characterized in that: The step of performing Tucker decomposition on the automatic speech recognition model parameters of at least one known language to obtain an initial core tensor and a language factor matrix specifically includes: In the Tucker decomposition process, the initial core tensor is used to capture the shared structural information in the parameters of the automatic speech recognition model; the language factor matrix is used to encode language-specific features and shared features, including but not limited to input dimensions, output dimensions and language-specific feature matrices; Tucker decomposition decomposes the multilingual weight tensor ΔW into: ΔW=C×1U (1) ×2U (2) ×3U (3) , Among them, C is the initial core tensor, U (1) , U (2) , U (3) They are language factor matrix, input factor matrix, and output factor matrix respectively.
3. The incremental Tucker decomposition method for multilingual automatic speech recognition according to claim 2, characterized in that: The step of obtaining the added new language and calculating the incremental weight according to the automatic speech recognition model parameters of the new language specifically includes: The incremental weight ΔW inc The expression is: ΔW inc =U (2) ×2C inc ×3U (3) , Where ×2 and ×3 are tensor-matrix multiplications along the input mode and output mode, respectively; C inc is the incremental core vector.
4. The incremental Tucker decomposition method for multilingual automatic speech recognition according to claim 3, characterized in that: The step of adopting the incremental updating method to project the incremental weight into the compression space to obtain the incremental core vector specifically includes: The incremental weight of the new language is projected into the compressed space to calculate the incremental core vector; the formula is expressed as: C inc ≈U (2)+ ×2ΔW inc ×3U (3)+ , Among them, C inc is the incremental core vector; ΔW inc is the incremental weight; U (2)+ and U (3)+ is the pseudo-inverse of the language factor matrix.
5. The incremental Tucker decomposition method for multilingual automatic speech recognition according to claim 4, characterized in that: After obtaining the incremental core vector, the method further includes: Adjust the language factor matrix to expand the language space to include the new language; the updated language factor matrix is expressed as: Among them, I n2 is the identity matrix corresponding to the new language; is the updated language factor matrix.
6. The incremental Tucker decomposition method for multilingual automatic speech recognition according to claim 5, characterized in that: The step of generating a multilingual automatic speech recognition model including a new language through the incremental core vector specifically includes: Reconstructing the weight tensor through the inverse process of Tucker decomposition using the incremental core vector and the updated language factor matrix; the reconstructed weight tensor contains information of all known languages and the new language, and generates a multilingual automatic speech recognition model including the new language; The reconstructed weight tensor is expressed as: Among them, C inc is the incremental core vector, The updated language factor matrix.
7. The incremental Tucker decomposition method for multilingual automatic speech recognition according to claim 6, characterized in that: The method further comprises: For an automatic speech recognition model of at least one known language, a low-rank adaptation method is used to generate a low-rank matrix update parameter for each language; The low-rank adaptation method fine-tunes the automatic speech recognition model by introducing low-rank matrices B and A to adjust the weight ΔW of the pre-trained model; the formula is expressed as: ΔW=BA, Among them, B∈R d×r , A∈R r×k , r is the rank of the low-rank matrix; For language i, the fine-tuned parameters are expressed as: h i =W0x i +ΔW i x i =W0x i +B i A i x i , Among them, x i is the input of language i; ΔW i =B i A i ;h i is the output of language i after calculation at the model layer; W0 is the original weight of the model layer.
8. An incremental Tucker decomposition device for multilingual automatic speech recognition, characterized in that: include: Model building module: Tucker decomposition of the automatic speech recognition model parameters of at least one known language is performed to obtain the initial core tensor and language factor matrix; Language integration module: based on acquiring the added new language, calculating the incremental weight according to the automatic speech recognition model parameters of the new language; Using an incremental update method, projecting the incremental weight into a compressed space to obtain an incremental core vector; wherein the compressed space is defined by the initial core tensor and the language factor matrix; Model reconstruction module: generating a multilingual automatic speech recognition model including a new language through the incremental core vector; Speech recognition module: using the multilingual automatic speech recognition model to transcribe or translate speech data containing a new language; Storage module: used to store the initial core tensor, the language factor matrix and the parameters required for calculating the incremental weight; wherein the core tensor and the language factor matrix are obtained after Tucker decomposition of the automatic speech recognition model parameters of at least one known language, and are used to capture the shared structural information in the model parameters and encode language-specific features and shared features respectively.
9. A computer device, characterized in that: The computer device includes a processor and a memory coupled to the processor, wherein the memory stores program instructions for implementing an incremental Tucker decomposition method for multi-language automatic speech recognition as described in any one of claims 1-7; the processor is used to execute the program instructions stored in the memory to implement an incremental Tucker decomposition method for multi-language automatic speech recognition.
10. A storage medium, characterized in that: The method stores program instructions executable by a processor, wherein the program instructions are used to execute the incremental Tucker decomposition method for multi-language automatic speech recognition as described in any one of claims 1 to 7.