A method, apparatus, medium, and device for training large models based on speech tasks
Patent Information
- Application Number
- CN202511984838.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-12-26
AI Technical Summary
[0005]本发明提供了一种基于语音任务的大模型训练方法、装置、介质及设备,以解决现有技术中无法准确高效地训练出各种语音任务模型的问题
Smart Images

Figure CN121768368B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech models, and more particularly to a method, apparatus, medium, and device for training large-scale models based on speech tasks. Background Technology
[0002] Most AutoML systems can only complete parameter combination selection within a preset discrete search space. They cannot dynamically integrate external domain knowledge, nor do they have high-level semantic reasoning capabilities. They are also difficult to perform targeted optimizations based on the specific context of speech tasks. At the same time, traditional AutoML lacks an adaptive learning mechanism based on historical experimental data. Each optimization process is independent of the others and cannot form a continuously evolving knowledge loop, resulting in high iteration costs and greatly limited optimization space.
[0003] With the rapid development of Large Language Models (LLMs), they have demonstrated powerful capabilities in natural language understanding and code generation. However, their application in the systematic collaboration of automated training for speech task models is still in its early stages. Currently, most large models passively respond to single instructions, lacking a complete closed-loop structure of "experiment planning—code generation—training execution—result feedback—solution regeneration." Furthermore, they are not deeply integrated with mainstream model training frameworks such as PyTorch and TensorFlow, making it impossible to adapt to the full-process requirements of speech task model training and hindering the conversion of the knowledge advantages of large models into actual improvements in model performance.
[0004] Furthermore, speech tasks have unique characteristics that distinguish them from image and text tasks. Taking speech enhancement as an example, the model not only needs to recover the signal spectrum, but also needs to take into account the requirements of speech intelligibility, naturalness, and low latency. Traditional supervised training methods (such as SEGAN, DCCRN, etc.) rely heavily on manual adjustment of model parameters and complex combinations of loss functions (such as STOI and PESQ loss). This process is not only cumbersome, but also extremely sensitive to noise distribution, making it difficult to achieve good generalization under different devices, sampling rates, and sound field conditions. These shortcomings make it impossible for existing technologies to accurately and efficiently train various speech task models. Summary of the Invention
[0005] This invention provides a method, apparatus, medium, and device for training large-scale models based on speech tasks, in order to solve the problem that existing technologies cannot accurately and efficiently train various speech task models.
[0006] Firstly, this application provides a method for training large models based on speech tasks, including: Obtain task description data, baseline model parameters, and experimental objectives for the speech task; wherein, the baseline model is a preset initial model trained based on a speech network architecture; The task description data and the experimental target are used as joint retrieval query information. Based on the retrieval enhancement generation technology, technical knowledge data related to the speech task are retrieved from the preset knowledge base. Based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data, structured prompt words are constructed; the structured prompt words are input into a preset large model so that the large model generates speech task model training code and experimental planning scheme that meet the requirements of the experimental objectives; The training code is executed based on the experimental planning scheme to iteratively train the benchmark model and obtain the first speech task model; the first speech task model is tested based on the preset test script and the performance index requirements in the experimental objectives to obtain test result data. The test results data or error information that occurs during iterative training are fed back to the large model and the historical experimental record data are updated until the preset iteration termination condition is met, thus obtaining the optimized speech task processing model.
[0007] This application provides a clear input foundation for optimizing automated speech enhancement models by acquiring task description data, baseline model parameters, and experimental objectives for speech tasks. Utilizing Retrieval-Enhanced Generation (RAG) technology to retrieve technical knowledge data related to speech tasks from a pre-defined knowledge base allows for the dynamic integration of cross-domain optimization strategies and methods, providing richer and more targeted knowledge support for model training. By combining task descriptions, experimental objectives, technical knowledge, and historical experimental records to construct structured prompts, and inputting them into a large model to generate training code and experimental plans, this application achieves efficient transformation from knowledge acquisition to code generation, avoiding the tedious process of traditional manual code writing. Simultaneously, leveraging the logical reasoning capabilities of the large model, it can explore better model structures and training strategies. Iterative training of the baseline model based on the experimental plans using the training code gradually optimizes model performance. The model is tested using pre-defined test scripts and performance metrics, and the test results or error information during training are fed back to the large model to further update historical experimental data, forming a closed-loop feedback mechanism. This process allows the system to learn and adjust in each iteration, avoiding repeated errors and continuously optimizing the model structure and training strategy until the preset iteration termination condition is met, ultimately resulting in an optimized speech task processing model. This application effectively solves the problem that existing technologies cannot accurately and efficiently train various speech task models.
[0008] Furthermore, the acquisition of the task description data, baseline model parameters, and experimental objectives for the speech task specifically includes: The task description data includes the type constraints of the speech task processing model, data format requirements, computing power limitations, real-time requirements, and dataset distribution description. The baseline model parameters include the number of network layers, the feature input dimension, and the network architecture type; the network architecture type is a convolutional neural network, a long short-term memory network, or a combination of a convolutional neural network and a long short-term memory network; the feature input dimension is the speech feature dimension based on short-time Fourier transform; The experimental objectives include performance constraints, lightweight requirements for speech task processing models, low-latency operation standards, and performance improvement requirements in pre-defined low signal-to-noise ratio scenarios.
[0009] This application provides a precise input foundation for the automated training of speech enhancement models through detailed task description data and benchmark model parameters. The task description data covers the type constraints of the speech task processing model, data format requirements, computational limitations, real-time requirements, and dataset distribution, enabling the system to optimize the model specifically according to the needs of specific application scenarios. The benchmark model parameters clearly define the number of network layers, feature input dimensions, and network architecture type, providing a clear definition of the initial model structure. The network architecture type includes Convolutional Neural Networks (CNN), Long Short-Term Memory Networks (LSTM), or combinations thereof, and the feature input dimension is based on Short-Time Fourier Transform (STFT). These details provide a stable starting point for model training. The experimental objectives further clarify performance constraints, model lightweight requirements, low-latency operation standards, and performance improvement needs in low signal-to-noise ratio scenarios. These objectives allow the system to focus on key performance indicators during optimization, ensuring that the generated model not only has theoretical optimization potential but also meets the performance requirements of complex scenarios such as low latency and low signal-to-noise ratio in practical applications.
[0010] Furthermore, the step of using the task description data and the experimental objective as joint retrieval query information, and retrieving technical knowledge data related to the speech task from a preset knowledge base based on retrieval enhancement generation technology, specifically involves: The task description data and experimental objectives are vector-encoded based on a preset embedding model to generate a joint query vector; wherein, the preset embedding model is the qwen3-embedding-4b model. The joint query vector is input into a preset vector database to match and obtain several related technical documents whose vector similarity meets a preset threshold; wherein, the preset vector database is the Milvus database; The aforementioned related technical documents are input into a preset semantic understanding model, so that the semantic understanding model extracts technical knowledge data, including network structure optimization, data augmentation schemes, and hyperparameter configuration ranges, from the related technical documents according to preset extraction rules. The preset semantic understanding model is the qwen3-plus model.
[0011] This application achieves knowledge-driven, efficient model optimization by combining task description data with experimental objectives and utilizing Retrieval Augmentation (RAG) technology to retrieve technical knowledge data related to speech tasks from a pre-defined knowledge base. Specifically, firstly, the task description and experimental objectives are vector-encoded using a pre-defined qwen3-embedding-4b model to generate a joint query vector. This process transforms complex task requirements into a quantifiable vector form, facilitating efficient subsequent retrieval. Subsequently, the joint query vector is input into the Milvus vector database, and highly relevant technical literature is obtained through vector similarity matching. This step ensures the accuracy and relevance of the retrieval results. Finally, the matched literature is input into the qwen3-plus semantic understanding model, and technical knowledge data, including network structure optimization, data augmentation schemes, and hyperparameter configuration ranges, is extracted according to pre-defined rules. This retrieval process based on embedding encoding and semantic understanding not only rapidly acquires cross-domain optimization knowledge but also provides rich strategies and methods for model training, thereby significantly improving the performance and generalization ability of speech enhancement models. It solves the problems of limited knowledge acquisition and singular optimization directions in traditional methods, demonstrating significant innovation and practicality.
[0012] Furthermore, based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data, structured prompt words are constructed; these structured prompt words are then input into a preset large model, enabling the large model to generate speech task model training code and experimental planning schemes that meet the experimental objective requirements. Specifically: Based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data, structured prompt words are constructed that include objective dimensions, reference dimensions, and historical experimental record dimensions. Among them, the target dimension is filled with task description data and experimental objectives, the reference dimension is filled with technical knowledge data, and the historical experiment record dimension is filled with code modifications to completed experiments and corresponding experimental results. The preset generation constraint instruction is embedded into the structured prompt word to obtain a structured prompt word carrying the generation constraint instruction; The structured cue words carrying generation constraints are input into a pre-set large model, so that the large model generates speech task model training code and experimental planning scheme that meet the experimental objectives based on the content of the structured cue words carrying generation constraints.
[0013] This application achieves a high degree of automation and precise optimization in speech enhancement model training by constructing structured prompt words that include target dimensions, reference dimensions, and historical experimental record dimensions, and inputting them into a pre-set large model to generate training code and experimental planning schemes for speech task models. Specifically, firstly, task description data and experimental objectives are used as the target dimension, technical knowledge data as the reference dimension, and code modifications and corresponding experimental results from completed experiments as the historical experimental record dimension, all integrated into structured prompt words. This process organically combines task requirements, external knowledge, and historical experience, providing comprehensive and accurate input for model generation. Subsequently, pre-set generation constraint instructions are embedded into the structured prompt words, further clarifying the specific requirements of the generation task and ensuring the relevance and effectiveness of the output results. Finally, the structured prompt words carrying generation constraints are input into the large model, enabling it to generate training code and experimental planning schemes that meet the experimental target requirements based on the prompt word content. This generation method based on structured prompt words not only fully utilizes historical experimental data and external knowledge, avoiding repeated errors, but also dynamically adjusts the optimization direction, thereby significantly improving the training efficiency and performance of speech enhancement models. It solves the problems of reliance on human experience and low optimization efficiency in traditional methods, demonstrating significant innovation and practicality.
[0014] Furthermore, the step of executing the training code based on the experimental planning scheme to iteratively train the benchmark model to obtain the first speech task model; and then testing the first speech task model based on the preset test script and the performance index requirements in the experimental objectives to obtain test result data, specifically as follows: The training code is executed based on either a modular restricted update mode or a global adaptive update mode. Combined with a preset dataset, the baseline model is iteratively trained to obtain the first speech task model. The modular restricted update mode only modifies the optimizer, loss function, and learning rate scheduler. The global adaptive update mode can reconstruct the overall training framework. The preset dataset is a dataset obtained by mixing preset clean speech with preset environmental noise at a random signal-to-noise ratio, superimposing the room impact response, and applying random gain scaling. Based on the preset test script and combined with the speech performance index requirements in the experimental objectives, the first speech task model was tested, and the corresponding performance index values were statistically analyzed and compiled into test result data. The voice performance indicators include a voice quality index not lower than a preset quality threshold and a voice intelligibility index not lower than a preset intelligibility threshold.
[0015] This application executes training code using either a modular restricted update mode or a global adaptive update mode. It iteratively trains a baseline model using a pre-set dataset to obtain a first speech task model. The model is then tested based on a pre-set test script and the performance metrics required in the experimental objectives, yielding test results. This process achieves efficient optimization and performance verification of the speech enhancement model. Specifically, the modular restricted update mode only modifies the optimizer, loss function, and learning rate scheduler, ensuring the stability of the main program; while the global adaptive update mode can reconstruct the overall training framework, further enhancing the flexibility and potential of model optimization. The pre-set dataset simulates complex speech environments in real-world scenarios by mixing clean speech with environmental noise at a random signal-to-noise ratio, superimposing room impact responses, and applying random gain scaling, providing diverse data support for model training. The first speech task model is tested based on the pre-set test script, and performance metrics such as speech quality and speech intelligibility are statistically analyzed and compiled into test results data, ensuring that the model can meet the requirements of speech quality and speech intelligibility not being lower than the pre-set quality threshold and speech intelligibility not being lower than the pre-set intelligibility threshold in practical applications. This method of combining training and testing not only improves the model's robustness and generalization ability in complex environments, but also ensures the model's practical application effectiveness through rigorous evaluation of performance metrics. It solves the problem of the disconnect between model optimization and practical application needs in traditional methods, and has significant innovation and practicality.
[0016] Furthermore, the process of obtaining the optimized speech task processing model until the preset iteration termination condition is met specifically involves: Real-time determination of whether the current iteration state meets the preset iteration termination condition; the iteration termination condition includes the number of experimental rounds reaching a preset round threshold, the model performance index reaching the experimental target requirement, or the number of consecutive error reports reaching a preset threshold; If the iteration termination condition is met, stop the model iteration training and determine the speech task model obtained in the last round of training as the optimized speech task processing model.
[0017] This application achieves efficient optimization and stable output of the speech task processing model by real-time determination of whether the current iteration state meets the preset iteration termination conditions. Specifically, the iteration termination conditions include the number of experimental rounds reaching a preset threshold, the model performance indicators meeting the experimental target requirements, or the number of consecutive error reports reaching a preset threshold. This mechanism ensures that the model training process can stop in time when the optimization target is reached or unsustainable errors occur, avoiding unnecessary waste of computational resources and invalid iterations. When the number of experimental rounds reaches the preset threshold, the system can ensure that the model has undergone sufficient training and optimization; when the model performance indicators meet the experimental target requirements, the system can lock the optimal model in time to ensure that the model performance meets the actual application requirements; and when the number of consecutive error reports reaches the preset threshold, the system can avoid falling into an invalid optimization loop, thereby protecting the stability and reliability of the system. Finally, the model obtained from the last round of training when the termination conditions are met is determined as the optimized speech task processing model, ensuring dual optimization of the model in terms of performance and stability. This solves the problems of lack of an effective termination mechanism and low optimization efficiency in the traditional model training process, and has significant innovation and practicality.
[0018] Secondly, this application provides a large-scale model training device based on speech tasks. The large-scale model training device based on speech tasks includes: The acquisition module is used to acquire task description data, baseline model parameters, and experimental objectives for the speech task; wherein the baseline model is a preset initial model trained based on a speech network architecture. The retrieval module is used to use the task description data and the experimental target as joint retrieval query information, and retrieve technical knowledge data related to the speech task from a preset knowledge base based on retrieval enhancement generation technology. The module is used to construct structured prompt words based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data; and input the structured prompt words into a preset large model so that the large model generates speech task model training code and experimental planning scheme that meet the requirements of the experimental objectives. The training module is used to execute the training code based on the experimental planning scheme to iteratively train the benchmark model to obtain the first speech task model; and to test the first speech task model based on the preset test script and the performance index requirements in the experimental objectives to obtain test result data. The generation module is used to feed back the test result data or error information that occurs during iterative training to the large model and update the historical experimental record data until the preset iteration termination condition is met, so as to obtain the optimized speech task processing model.
[0019] This application utilizes a systematic device to achieve efficient optimization and automated iteration of a speech task processing model. The acquisition module accurately collects task descriptions, baseline model parameters, and experimental objectives, providing a clear basis for subsequent optimization. The retrieval module employs retrieval enhancement generation technology to acquire relevant technical knowledge from a knowledge base, broadening optimization approaches. The construction module integrates task requirements, knowledge data, and historical experience into structured prompts, inputting them into the large model to generate targeted training code and experimental plans, ensuring the scientific validity and effectiveness of the optimization direction. The training module iteratively trains the baseline model according to the planned scheme and rigorously evaluates model performance through preset test scripts, ensuring model quality. The generation module feeds back test results or error information to the large model, dynamically updating historical records until the iteration termination condition is met, ultimately outputting the optimized speech task processing model. This device, through the synergistic effect of its modules, forms a closed-loop optimization process, not only improving model optimization efficiency but also significantly enhancing the model's performance and generalization ability in complex speech tasks.
[0020] Thirdly, this application provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the large model training method based on a speech task as described above. Its beneficial effects are the same as those of the large model training method based on a speech task provided in the first aspect of this application.
[0021] Fourthly, this application provides a terminal device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement any of the large model training methods based on speech tasks as described in the first aspect. Attached Figure Description
[0022] Figure 1 : A schematic diagram of an embodiment of the large model training method based on speech tasks provided in this application; Figure 2 : A schematic diagram illustrating an embodiment of the large-scale model training process based on speech tasks provided in this application; Figure 3 : A schematic diagram of an embodiment of the speech task-based large model training device provided in this application. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Example 1 Please refer to Figure 1 In order to solve the problem that existing technologies cannot accurately and efficiently train various speech task models, this invention provides a large model training method based on speech tasks, including steps S01-S05.
[0025] S01: Obtain the task description data, baseline model parameters, and experimental objectives for the speech task; wherein the baseline model is a preset initial model trained based on a speech network architecture.
[0026] In a preferred embodiment of this invention, the acquisition of the task description data, baseline model parameters, and experimental objectives for the speech task specifically involves: Collect and organize task description data corresponding to speech tasks, including but not limited to speech enhancement, speech denoising, and speech separation. This data must be clearly described by the user in text form, comprehensively covering the core constraints and requirements of speech task processing. Specifically, it includes constraints on the type of speech task processing model, such as explicitly requiring a certain layer of the model to use a convolutional neural network, long short-term memory network, deep neural network, or a combination of both, while limiting the maximum number of network layers; it also needs to specify data format requirements, such as specifying that the model must be compatible with speech data formats such as WAV, PCM, and MP3, and different sampling rates such as 16kHz and 48kHz; it also needs to indicate computing power constraints (such as the number of available GPUs, the size of video memory, and other hardware constraints), real-time requirements (such as the hard indicator that the end-to-end processing latency must be controlled within 50ms), and dataset distribution descriptions (such as the average duration of speech in the training dataset, the mixing method of noise and clean speech, the signal-to-noise ratio range, and the gain adjustment rules, etc.). In addition, the training optimization goals should be clearly defined by the user. They can be flexibly set to directions such as model miniaturization, reducing the number of parameters, reducing the amount of computation, or improving the speech intelligibility index of the test set without changing the amount of computation.
[0027] While acquiring the task description data, the parameters of the baseline model need to be retrieved simultaneously. The baseline model is an initial model that has been pre-trained based on a basic speech network architecture. Its parameters should include three core elements: the number of network layers, the feature input dimension, and the network architecture type. The number of network layers can be set to a basic architecture of 3-5 layers according to the initial requirements of the task. The feature input dimension is based on the speech feature dimension of short-time Fourier transform by default to adapt to the time-frequency domain characteristics of speech signals. The network architecture type can be a convolutional neural network, a long short-term memory network, or a hybrid architecture combining the two. This initial model has been pre-trained based on a basic speech dataset and has basic speech signal processing capabilities, and can be directly used as the base model for subsequent iterative optimization.
[0028] Furthermore, experimental objectives need to be clearly defined based on actual application scenarios. These objectives must balance performance and operational requirements, specifically including performance constraints, such as requiring the model to achieve a speech quality index (PESQ) of no less than 2.0, a speech intelligibility index (STOI) of no less than 0.85, and a significant improvement in the signal-to-noise ratio (SI-SNR) in low signal-to-noise ratio (SNR<5dB) environments. Simultaneously, lightweight requirements for the speech task processing model need to be clearly defined, such as controlling the number of model parameters within a preset threshold to adapt to mobile deployment. Low-latency operation standards also need to be set to ensure that the end-to-end processing latency of the model in real-time scenarios does not exceed 50ms. Additionally, performance improvement requirements in low signal-to-noise ratio scenarios should be included to provide clear direction for subsequent model iteration and optimization.
[0029] S02: Using the task description data and the experimental target as joint retrieval query information, and based on retrieval enhancement generation technology, retrieve technical knowledge data related to the speech task from the preset knowledge base.
[0030] In a preferred embodiment of this invention, the step of using the task description data and the experimental objective as joint retrieval query information, and retrieving technical knowledge data related to the speech task from a preset knowledge base based on retrieval enhancement generation technology, specifically involves: After acquiring the task description data, baseline model parameters, and experimental objectives for the speech task, the system will proceed to the technical knowledge retrieval stage. This stage uses enhanced generation techniques to filter matching technical knowledge data from a pre-defined knowledge base. The specific execution flow is as follows: First, the system integrates the task description data with the experimental objectives to form a joint retrieval query. The task description data covers the type constraints of the speech task processing model, data format requirements, computational limitations, real-time requirements, and dataset distribution. The experimental objectives include performance constraints, model lightweighting requirements, low-latency operation standards, and performance improvement requirements in low signal-to-noise ratio scenarios. Then, the system calls a pre-defined embedding model to vectorize the joint retrieval query, generating a unified joint query vector. The embedding model used is the qwen3-embedding-4b model, which can transform textual query information into semantically meaningful vector data, laying the foundation for subsequent accurate retrieval.
[0031] Next, the system inputs the generated joint query vector into a pre-defined vector database, the Milvus database, which pre-stores embedding vectors corresponding to a massive amount of technical literature, code examples, and methodological materials in the fields of speech processing and machine learning. Based on a vector similarity algorithm, the database matches several related technical documents whose similarity to the joint query vector meets a preset threshold. In this process, the top 100 related technical documents with the highest similarity are selected as the search results by default, thus ensuring the relevance and coverage of the search content.
[0032] In the vector database matching stage, this process defaults to selecting the top 100 related technical documents with the highest similarity as the search results. When issuing knowledge extraction instructions to the semantic understanding model, the specific prompts are a combination of user input and fixed instructions. The fixed instructions are: "Key information extraction: Organize the following into modules: 'Research background and problems' (including shortcomings in existing research), 'Research methods' (including data sources, experimental design / model architecture), 'Core results' (annotating key data and statistical significance, such as P-value and accuracy), and 'Conclusions and prospects' (including research limitations and future directions). Each module should be precisely extracted with short sentences to avoid redundancy." This ensures that the extracted technical knowledge data has the characteristics of being structured and highly practical.
[0033] Finally, the system inputs all the aforementioned related technical literature into a pre-defined semantic understanding model, namely the qwen3-plus model, and simultaneously issues a pre-defined knowledge extraction instruction. This instruction requires the model to organize the relevant content of "Research Background and Problems," "Research Methods," "Core Results," and "Conclusions and Prospects" from the literature in modules, and to extract the key technical points that can be used for speech task model optimization. After receiving the instruction and literature data, the model automatically completes semantic understanding and information compression, and finally outputs technical knowledge data including network structure optimization schemes, data augmentation strategies, hyperparameter configuration ranges, feature engineering techniques, and loss function design schemes, providing key technical support for subsequent structured prompt word construction and training code generation.
[0034] S03: Based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data, construct structured prompt words; input the structured prompt words into a preset large model so that the large model generates speech task model training code and experimental planning scheme that meet the experimental objective requirements.
[0035] In a preferred embodiment of this invention, the step of constructing structured prompt words based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data; and inputting the structured prompt words into a preset large model to enable the large model to generate speech task model training code and experimental planning scheme that meet the requirements of the experimental objectives, specifically: After completing the retrieval of technical knowledge data related to the speech task, the system will enter the stage of generating structured prompt words, training code, and experimental planning schemes. The specific execution process is as follows: First, the structured prompt words are constructed. The system integrates task description data, experimental objectives, technical knowledge data, and preset historical experimental record data. At the code level, a dictionary structure with three keys—objective, reference, and past experimental records—is built, forming a three-dimensional structured prompt word framework. The "objective" key in the dictionary is a character type, containing the user-inputted task description data, and supplemented with additional prompt words in a fixed format. The "reference" key is a list type, with each element containing JSON-formatted search technical literature. The "past experimental records" key is also a list type, with each element containing JSON-formatted modifications to completed experimental code and corresponding experimental results. Correspondingly, in the prompt word framework, the objective dimension is fully populated with the task description data and experimental objectives. The model type constraints, data format requirements, computational limitations, real-time requirements, and dataset distribution descriptions in the task description data are correlated with the experimental objectives. The performance constraints, lightweight model requirements, low-latency operation standards, and performance improvement needs in low signal-to-noise ratio scenarios in the standard collectively clarify the core direction of model training. The reference dimension is filled with technical knowledge data obtained through retrieval and refinement, covering technical points that can be directly used for model optimization, such as network structure optimization suggestions, data augmentation schemes, hyperparameter configuration ranges, and feature engineering techniques. The historical experiment record dimension will record the code modifications and corresponding experimental results of completed experiments, including previously tried feature engineering methods, model structure variants, hyperparameter combinations, and changes in model test indicators for each scheme, providing a reference for avoiding duplication in the generation of subsequent schemes.
[0036] After the basic framework is built, the system will pre-set the generation constraint instructions and embed them into the structured prompts. The generation constraint instructions include basic requirements such as not repeating historical experimental schemes, generating modular Python code that can be embedded, and outputting multiple sets of parallel experimental schemes, as well as specific requirements for adapting to speech tasks. For scenarios involving only feature engineering optimization, the fixed prompt is: "Please write a Python function named `enhance_engineer`. This function should receive a feature with dimensions `batch_size`, `seq_len`, and `feature_dim`. This feature should be similar to Struts-Focused Features (SFT). This function should return a new array with dimensions `batch_size`, `seq_len`, and `feature_dim`, which represents the feature engineering result of the input array. Please try a feature engineering method, such as addition, subtraction, multiplication, or division. You can use NumPy or Math libraries. Do not output any non-code content; directly output the executable function's code. No additional information is needed, including specifying that it's Python code or providing comments. I will directly call this function. I will experiment with each method one by one and send you the methods I've tried. If you see a method you've already tried, please do not try it again. If I don't send you any tried methods, it means..." This is the first attempt; you can try simple feature engineering methods first. Note that this is a real-time audio processing system. For scenarios requiring reconstruction of the entire training code, the fixed prompt is adjusted to: "The previous training code and experimental results are attached below. Please rewrite the training code, referring to the provided documentation. Solutions already tried do not need to be retried. Combine untried solutions to create a new codebase. Do not change the model input / output, or the function names and storage location of the model file. You can add or modify intermediate steps such as data augmentation, feature engineering, model structure, loss function, optimizer, and hyperparameters. Carefully consider whether the code conforms to the rules before outputting; otherwise, errors will occur, wasting the solution." Simultaneously, the speech-specific requirements will be explicitly written into the prompt, specifically: "Based on the changes in PESQ / STOI metrics from the last test, automatically determine whether the problem lies in the 'feature extraction stage' or the 'model structure stage'; if the model fails under low SNR conditions, the large model will automatically recommend the 'bandwidth attention network' or 'multi-scale feature fusion' solution," thus forming a complete structured prompt carrying generation constraints.
[0037] Subsequently, the system inputs structured prompts carrying generation constraints into a pre-set large model. The selected large model can be chosen from models with code generation and logic planning capabilities, such as qwen3-plus, qwen3-max, and qwen3-coder-plus, depending on actual needs. The code generated by the model is automatically written into a pre-defined document and then called in the training code, without the need for manual intervention in the code migration process. Based on the three-dimensional core information and generation constraint instructions within the prompts, the large model first deeply understands the core requirements of the speech task and the existing technical reference directions, then combines historical experimental records to avoid duplicate solutions, and finally generates speech task model training code and experimental planning schemes that meet the experimental objectives. The training code will be presented in a modular form, including feature engineering functions, model structure variant functions, data augmentation strategy functions, and hyperparameter configuration information. All code is pure Python code that can be directly embedded into existing training frameworks, without comments or redundant non-code content. The experimental planning scheme will clearly define the number of parallel experiments, the core optimization direction of each experiment, and the expected performance indicators for verification. For example, parallel experiments with lightweight model structures will be planned for low latency requirements, and multi-parameter variant experiments of frequency band attention networks will be planned for low signal-to-noise ratio scenarios to ensure that the experiments cover the core dimensions of model optimization.
[0038] S04: Execute the training code based on the experimental planning scheme to iteratively train the benchmark model and obtain the first speech task model; test the first speech task model based on the preset test script and the performance index requirements in the experimental objectives to obtain test result data.
[0039] In a preferred embodiment of this invention, the step of executing the training code based on the experimental planning scheme to iteratively train the benchmark model yields a first speech task model; and the first speech task model is tested based on the preset test script and the performance indicators in the experimental objectives to obtain test result data, specifically as follows: After generating the training code and experimental plan for the speech task model, this application will enter the model iterative training and performance testing phase. The specific execution process is as follows: This application will call the generated training code to perform iterative training on the benchmark model according to the training mode specified in the experimental plan. The training mode is divided into two types: modular restricted update and global adaptive update. If the modular restricted update mode is selected, only the optimizer, loss function and learning rate scheduler of the benchmark model will be adjusted, while the rest of the core training framework will remain unchanged, so as to ensure the stability of the training process. If the global adaptive update mode is selected, the overall training framework can be reconstructed, including the optimization and adjustment of core links such as data loading process, network structure hierarchy and model evaluation logic.
[0040] During training, the system uses a pre-set dedicated dataset to iterate the model's parameters. This dataset is built based on open-source, commercially available speech resources. Specifically, it mixes clean speech data, TTS-synthesized speech data, and various environmental noise data with random signal-to-noise ratios (normally distributed), and superimposes room impact response to simulate spatial reverberation. At the same time, it applies random gain scaling to simulate the signal acquisition differences of different devices. To improve the model's robustness to complex noise scenarios, it also dynamically performs SpecAugment spectral masking and time stretching operations at 0.8-1.2 times speed on the dataset. The dataset is divided into 80% training set, 10% validation set, and 10% test set to ensure no data cross-contamination.
[0041] Throughout the training process, the system will monitor the computing power resource usage, model training loss value, and validation set accuracy data in real time. When the training rounds reach the preset value of the experimental plan or the validation set indicators tend to stabilize, the training will be stopped and the model at this time will be determined as the first speech task model.
[0042] After training is complete, the system automatically calls the pre-set test script and conducts a comprehensive test on the first speech task model, based on the performance indicators specified in the experimental objectives. The model files and parameter storage locations from the training phase are pre-set, and the test code directly reads the model information from the agreed locations and executes the test script. During testing, the system reads the processed speech data output by the first speech task model and compares it with the corresponding clean reference speech data, focusing on calculating three core indicators: first, a speech quality indicator based on the PESQ algorithm, measuring the subjective listening quality of the speech; second, a speech intelligibility indicator based on the STOI algorithm, assessing the recognizability of the speech content; and third, a signal-to-noise ratio indicator based on the SI-SNR algorithm, determining the actual effect of noise suppression. After testing, the system calculates the specific values of each indicator and analyzes the trends of the indicators compared to the baseline model. This data is then organized and archived in a pre-set JSON format. The test results are also archived synchronously and will be incorporated into the prompt words and fed back to the large model at the start of the next iteration, forming a complete test-feedback loop. This ultimately generates complete test result data, providing a quantitative reference for subsequent model iteration and optimization.
[0043] S05: Feed back the test result data or error information that occurs during iterative training to the large model and update the historical experimental record data until the preset iteration termination condition is met, and obtain the optimized speech task processing model.
[0044] In a preferred embodiment of this invention, the step of feeding back the test result data or error information that occurs during iterative training to the large model and updating the historical experimental record data until a preset iteration termination condition is met, thereby obtaining an optimized speech task processing model, specifically involves: If runtime errors occur during model iterative training, such as memory overflow, gradient explosion, incorrect code path, or abnormal function call, the system will automatically capture complete error logs, extracting the specific stage where the error occurred, the error type, and the core error information. Subsequently, the system will integrate the training code generated by the large model in that round with the error information, input it into the preset large model, and issue a specific prompt instruction. This instruction states, "I embedded the code generated by the large model into my code and executed it, but an error occurred. Please write a prompt message to guide the large model on what not to do. I will also input this prompt message into the large model." This requires the large model to generate targeted error-prevention prompts, clearly defining the error types and operations to be avoided in subsequent code generation. Simultaneously, the system will update the historical experimental record data with the error information, corresponding code, and error-prevention prompts, completing the update of the error record for that round.
[0045] If no errors occur during training, the system will structure the test results data according to the performance metrics defined in the experimental objectives. It will focus on categorizing the specific values of the speech quality index (PESQ), speech intelligibility index (STOI), and signal-to-noise ratio (SI-SNR), as well as the improvement or trend of each index compared to the baseline model. After processing, the system will uniformly feed back the training code, experimental configuration, and test results data for that round to the pre-set large model, providing a quantitative reference for subsequent adjustments to the large model's training scheme. Simultaneously, the system will completely record the above information into historical experimental data, forming a traceable experimental iteration chain.
[0046] This application verifies in real time whether the current iteration status meets the preset iteration termination conditions after each round of feedback. The iteration termination conditions include three core scenarios: First, the number of experimental rounds reaches a preset round threshold, meaning the maximum number of experiments preset by the system has been exhausted; second, the model performance indicators meet the experimental target requirements, such as the speech quality indicator reaching a preset threshold in low signal-to-noise ratio scenarios, the speech intelligibility indicator meeting practical application standards, and the low latency constraint also being met; third, the number of consecutive error reports reaches a preset threshold, meaning the system experiences unrecoverable operational errors for multiple consecutive rounds, determining that the current direction has no optimization value. When any termination condition is met, the system will stop model iteration training and select the speech task model with the best performance indicators obtained in the last training round as the final optimized speech task processing model.
[0047] If the termination conditions are not met, the system will return to the structured prompt word construction stage, combine the updated historical experimental record data, regenerate training code and experimental planning schemes that meet the new requirements, and start the next round of model iteration training.
[0048] Furthermore, the method of this application has broad applicability and can be adapted to various practical application scenarios of speech tasks. The following are specific implementation examples: Scenario 1: Real-time mobile call noise reduction (e.g., VoIP applications): In this scenario, the task description data clearly states that the model needs to be adapted to 16kHz single-channel PCM format speech data, using 30ms frame length processing, and must meet the constraint of limited network bandwidth; the experimental objectives are set as follows: end-to-end latency not exceeding 50ms, CPU utilization on mid-tier mobile processors (SoCs) not exceeding 10%, and speech quality index (PESQ) not lower than 2.0 under a -5dB signal-to-noise ratio environment.
[0049] After acquiring the task description data, baseline model parameters, and experimental objectives, the system uses retrieval enhancement generation techniques to obtain relevant technical knowledge data on mobile lightweight model design and low-latency feature engineering. It then constructs structured prompt words and generates targeted training code. During training, a modular, constrained update mode is adopted, focusing on optimizing the model's computational efficiency and latency performance. The dataset is generated by mixing clean speech with various environmental noises (such as traffic noise and human voice interference) at a random signal-to-noise ratio using a 16kHz sampling rate.
[0050] After training, the first speech task model met the experimental objectives after testing, the iteration terminated, and it was determined to be the optimized speech task processing model. During deployment, the model was quantized to int8 precision and deployed at the edge using the PyTorch Mobile or ONNX runtime framework. In the mobile acquisition stage, speech activity detection (VAD) and frame packetization were performed first. The model ran in a frame-by-frame inference manner, and the output denoised PCM frames were encoded by Opus and sent as RTP packets to ensure the smoothness of real-time calls and the effect of noise reduction.
[0051] Scenario 2: Voice enhancement for hearing aids / wearable devices: The task description data for this scenario includes a multi-channel microphone array signal with an input of 48kHz sampling rate, which needs to be adapted to a dedicated DSP / FPGA hardware deployment environment. The experimental goal is to achieve hard real-time requirements with a latency of less than 10ms, a device power consumption of no more than 50mW, and at the same time ensure speech intelligibility and naturalness.
[0052] Based on the technical knowledge related to task description and experimental target retrieval, the system focuses on acquiring key technologies such as multi-channel signal processing, model fixed-point implementation, and hardware acceleration path design, generating training code that includes model pruning, int16 quantization, and dedicated computing kernel development. During training, a global adaptive update mode is adopted, and the model structure is reconstructed to adapt to hardware resources. The dataset is superimposed with room impact response to simulate a near-field sound environment.
[0053] The optimized voice task processing model is deployed on a dedicated DSP / FPGA, employing a processing approach that combines directional beamforming with a learning front-end. A low-latency overlap-add pipeline is implemented in the hardware firmware, and input / output buffer addresses and interrupt hooks are explicitly defined through firmware-level APIs to ensure efficient communication with the headset firmware. The final low-latency enhanced audio output is directly fed into the headset digital-to-analog converter (DAC) to meet the stringent operating requirements of wearable devices.
[0054] Scenario 3: Meeting Recording Enhancement and ASR Front-End Processing: The task description for this scenario explicitly states that the input data is a multi-channel conference recording with a 16kHz sampling rate, supporting long-term recording requirements. The processed data needs to be adapted to an automatic speech recognition (ASR) system. The experimental objectives are to reduce the ASR word error rate (WER) by no less than 5% and improve the speech quality index (PESQ) by no less than 0.3.
[0055] The system retrieves relevant technical knowledge on multi-channel dereverberation, speech denoising and ASR joint optimization, and generates training code that includes multi-channel fusion and denoising processing. During training, server-side batch inference is used for adaptation optimization. The dataset includes mixed data of conference room ambient noise, multi-person dialogue speech and different reverberation conditions.
[0056] The optimized model is deployed on the server side. It first performs multi-channel fusion déreverberation processing on the meeting recordings, then performs noise reduction and enhancement. The processed audio is then synchronously input into the ASR system. The system uses the enhanced audio quality and ASR word error rate as joint evaluation indicators, and finally outputs enhanced WAV format audio and ASR transcribed text for subsequent search and playback, which significantly improves the usability of meeting recordings and ASR recognition accuracy.
[0057] like Figure 2 As shown, Figure 2This is a flowchart illustrating the large-scale model training process for speech tasks in this embodiment: The process begins with the user specifying the task type, computing power, metrics, and other requirements. First, component 1 (RAG matching query) matches highly relevant documents / code segments in the knowledge base. Then, the original full text and project code are obtained from the matching results and input into the large model to extract relevant summaries for the corresponding task. Subsequently, component 2 (creating prompt words) concatenates user requirements, task information, summaries, experimental logs, etc., into structured prompt words in a fixed format and limits the output format. This is then input into the large model to generate training code. Next, automatic experiments are started. The code is adjusted and run through two methods: "inserting code into restricted modules" or "the large model independently modifies the code." After training is completed, testing is performed. The large model extracts the test results or error information and sends these contents back to the prompt word construction stage to form an iterative loop until a stop signal is received or the preset number of loops is reached, at which point the process ends.
[0058] In summary, this application provides a clear input foundation for optimizing automated speech enhancement models by acquiring task description data, baseline model parameters, and experimental objectives for speech tasks. Utilizing Retrieval Enhancement Generation (RAG) technology to retrieve technical knowledge data related to speech tasks from a pre-defined knowledge base allows for the dynamic integration of cross-domain optimization strategies and methods, thus providing richer and more targeted knowledge support for model training. By constructing structured prompts based on task descriptions, experimental objectives, technical knowledge, and historical experimental records, and inputting them into a large model to generate training code and experimental plans, this application achieves an efficient transformation from knowledge acquisition to code generation, avoiding the tedious process of traditional manual code writing. Furthermore, leveraging the logical reasoning capabilities of the large model enables the exploration of better model structures and training strategies. Iterative training of the baseline model based on the experimental plans allows for gradual optimization of model performance. Testing the model using pre-defined test scripts and performance metrics, and feeding back test results or error information from the training process to the large model, further updates historical experimental data and forms a closed-loop feedback mechanism. This process allows the system to learn and adjust in each iteration, avoiding repeated errors and continuously optimizing the model structure and training strategy until the preset iteration termination condition is met, ultimately resulting in an optimized speech task processing model. This application effectively solves the problem that existing technologies cannot accurately and efficiently train various speech task models.
[0059] Example 2 Please refer to Figure 3 This is a large model training device based on speech tasks provided in the embodiments of this application.
[0060] In this embodiment, the large model training device based on speech tasks includes an acquisition module 10, a retrieval module 20, a construction module 30, a training module 40, and a generation module 50.
[0061] The acquisition module 10 is used to acquire task description data, benchmark model parameters, and experimental objectives for the speech task; wherein the benchmark model is a preset initial model trained based on a speech network architecture. The retrieval module 20 is used to use the task description data and the experimental target as joint retrieval query information, and retrieve technical knowledge data related to the speech task from a preset knowledge base based on retrieval enhancement generation technology. The construction module 30 is used to construct structured prompt words based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data; and input the structured prompt words into a preset large model so that the large model generates speech task model training code and experimental planning scheme that meet the requirements of the experimental objectives. Training module 40 is used to execute the training code based on the experimental planning scheme to iteratively train the benchmark model to obtain the first speech task model; and to test the first speech task model based on the preset test script and the performance index requirements in the experimental objectives to obtain test result data. The generation module 50 is used to feed back the test result data or error information that occurs during the iterative training process to the large model and update the historical experimental record data until the preset iteration termination condition is met, so as to obtain the optimized speech task processing model.
[0062] For ease of description and brevity, the embodiments of the present invention include all the implementation methods described in the above-described embodiments of the large model training method based on speech tasks, and will not be repeated here.
[0063] Example 3: This application provides a computer-readable storage medium, which includes a stored computer program, wherein the computer program controls the device where the computer-readable storage medium is located to execute the aforementioned large model training method based on speech tasks when it is executed. The large-scale model training method based on speech tasks, when implemented as a software functional unit and used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0064] Example 4 This embodiment provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements any of the large model training methods based on speech tasks as described in Embodiment 1.
[0065] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for training large models based on speech tasks, characterized in that, include: Obtain task description data, baseline model parameters, and experimental objectives for the speech task; wherein, the baseline model is a preset initial model trained based on a speech network architecture; The task description data and the experimental target are used as joint retrieval query information. Based on the retrieval enhancement generation technology, technical knowledge data related to the speech task are retrieved from the preset knowledge base. Based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data, structured prompt words are constructed; the structured prompt words are input into a preset large model so that the large model generates speech task model training code and experimental planning scheme that meet the requirements of the experimental objectives; The training code is executed based on the experimental planning scheme to iteratively train the benchmark model and obtain the first speech task model; the first speech task model is tested based on the preset test script and the performance index requirements in the experimental objectives to obtain test result data. The test results data or error information that occurs during iterative training are fed back to the large model and the historical experimental record data are updated until the preset iteration termination condition is met, thus obtaining the optimized speech task processing model.
2. The method for training large models based on speech tasks according to claim 1, characterized in that, The acquisition of the task description data, baseline model parameters, and experimental objectives for the speech task specifically includes: The task description data includes the type constraints of the speech task processing model, data format requirements, computing power limitations, real-time requirements, and dataset distribution description. The baseline model parameters include the number of network layers, the feature input dimension, and the network architecture type; the network architecture type is a convolutional neural network, a long short-term memory network, or a combination of a convolutional neural network and a long short-term memory network; the feature input dimension is the speech feature dimension based on short-time Fourier transform; The experimental objectives include performance constraints, lightweight requirements for speech task processing models, low-latency operation standards, and performance improvement requirements in pre-defined low signal-to-noise ratio scenarios.
3. The method for training large models based on speech tasks according to claim 1, characterized in that, The step of using the task description data and the experimental objective as joint retrieval query information, and retrieving technical knowledge data related to the speech task from a preset knowledge base based on retrieval enhancement generation technology, specifically involves: The task description data and experimental objectives are vector-encoded based on a preset embedding model to generate a joint query vector; wherein, the preset embedding model is the qwen3-embedding-4b model. The joint query vector is input into a preset vector database to match and obtain several related technical documents whose vector similarity meets a preset threshold; wherein, the preset vector database is the Milvus database; The aforementioned related technical documents are input into a preset semantic understanding model, so that the semantic understanding model extracts technical knowledge data, including network structure optimization, data augmentation schemes, and hyperparameter configuration ranges, from the related technical documents according to preset extraction rules. The preset semantic understanding model is the qwen3-plus model.
4. The method for training large models based on speech tasks according to claim 1, characterized in that, The structured prompt words are constructed based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data. The structured prompts are input into a pre-set large model, which generates speech task model training code and experimental planning scheme that meet the experimental objectives. Specifically: Based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data, structured prompt words are constructed that include objective dimensions, reference dimensions, and historical experimental record dimensions. Among them, the target dimension is filled with task description data and experimental objectives, the reference dimension is filled with technical knowledge data, and the historical experiment record dimension is filled with code modifications to completed experiments and corresponding experimental results. The preset generation constraint instruction is embedded into the structured prompt word to obtain a structured prompt word carrying the generation constraint instruction; The structured cue words carrying generation constraints are input into a pre-set large model, so that the large model generates speech task model training code and experimental planning scheme that meet the experimental objectives based on the content of the structured cue words carrying generation constraints.
5. The method for training large models based on speech tasks according to claim 1, characterized in that, The training code is executed based on the experimental planning scheme to iteratively train the benchmark model, resulting in a first speech task model. Based on the preset test script and the performance indicators in the experimental objectives, the first speech task model is tested to obtain test result data, specifically: The training code is executed based on either a modular restricted update mode or a global adaptive update mode. Combined with a preset dataset, the baseline model is iteratively trained to obtain the first speech task model. The modular restricted update mode only modifies the optimizer, loss function, and learning rate scheduler. The global adaptive update mode can reconstruct the overall training framework. The preset dataset is a dataset obtained by mixing preset clean speech with preset environmental noise at a random signal-to-noise ratio, superimposing the room impact response, and applying random gain scaling. Based on the preset test script and combined with the speech performance index requirements in the experimental objectives, the first speech task model was tested, and the corresponding performance index values were statistically analyzed and compiled into test result data. The voice performance indicators include a voice quality index not lower than a preset quality threshold and a voice intelligibility index not lower than a preset intelligibility threshold.
6. The method for training large models based on speech tasks according to claim 1, characterized in that, The process continues until a preset iteration termination condition is met, resulting in an optimized speech task processing model, specifically: Real-time determination of whether the current iteration state meets the preset iteration termination condition; the iteration termination condition includes the number of experimental rounds reaching a preset round threshold, the model performance index reaching the experimental target requirement, or the number of consecutive error reports reaching a preset threshold; If the iteration termination condition is met, stop the model iteration training and determine the speech task model obtained in the last round of training as the optimized speech task processing model.
7. A large-scale model training device based on speech tasks, characterized in that, include: The acquisition module is used to acquire task description data, baseline model parameters, and experimental objectives for the speech task; wherein the baseline model is a preset initial model trained based on a speech network architecture. The retrieval module is used to use the task description data and the experimental target as joint retrieval query information, and retrieve technical knowledge data related to the speech task from a preset knowledge base based on retrieval enhancement generation technology. The module is used to construct structured prompt words based on the task description data, experimental objectives, technical knowledge data, and preset historical experimental record data; and input the structured prompt words into a preset large model so that the large model generates speech task model training code and experimental planning scheme that meet the requirements of the experimental objectives. The training module is used to execute the training code based on the experimental planning scheme to iteratively train the benchmark model to obtain the first speech task model; and to test the first speech task model based on the preset test script and the performance index requirements in the experimental objectives to obtain test result data. The generation module is used to feed back the test result data or error information that occurs during iterative training to the large model and update the historical experimental record data until the preset iteration termination condition is met, so as to obtain the optimized speech task processing model.
8. The large model training device based on speech tasks according to claim 7, characterized in that, The step of using the task description data and the experimental objective as joint retrieval query information, and retrieving technical knowledge data related to the speech task from a preset knowledge base based on retrieval enhancement generation technology, specifically involves: The task description data and experimental objectives are vector-encoded based on a preset embedding model to generate a joint query vector; wherein, the preset embedding model is the qwen3-embedding-4b model. The joint query vector is input into a preset vector database to match and obtain several related technical documents whose vector similarity meets a preset threshold; wherein, the preset vector database is the Milvus database; The aforementioned related technical documents are input into a preset semantic understanding model, so that the semantic understanding model extracts technical knowledge data, including network structure optimization, data augmentation schemes, and hyperparameter configuration ranges, from the related technical documents according to preset extraction rules. The preset semantic understanding model is the qwen3-plus model.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the large model training method for speech tasks as described in any one of claims 1 to 6.
10. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the large model training method for speech tasks as described in any one of claims 1 to 6.