Sound model training method and sound model training device
By dynamically determining the target training method of training data, and using dynamic distributed parameters to classify training data, efficient training of zero-sample learning and single-sample learning is achieved, solving the complex problems of training methods in the existing technology, simplifying the model training process and improving efficiency.
Patent Information
- Application Number
- CN202510903497.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-07-01
AI Technical Summary
The prior art has the problem of complex and inefficient training methods in the zero-sample learning and single-sample learning modes for training deep learning models, and different training data and strategies need to be constructed separately.
By obtaining dynamic distribution parameters, the target training method of each training data is dynamically determined to be zero-sample learning or single-sample learning, and the training data is classified based on the dynamic distribution parameters, and the same training data set and strategy are used to achieve training of the two learning modes.
The model training process is simplified and the training efficiency is improved, so that the model can perform well in both zero-sample learning and single-sample learning modes without complex feature decoupling or additional data preparation.
Smart Images

Figure CN120544545A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of model training technology, and in particular to a sound model training method and a sound model training device. Background Art
[0002] With the rapid development of artificial intelligence (AI), deep learning models have made significant progress in fields such as image recognition and natural language processing. Zero-shot learning and one-shot learning, emerging machine learning paradigms, can address the data scarcity and dynamic category expansion issues of traditional supervised learning. In zero-shot learning, a model can directly understand and reason about new inputs based on its generalization capabilities, even though it has never been exposed to a specific category or task during training. In one-shot learning, a model quickly adapts to new tasks using a single reference sample, which encodes or provides information about key features of the target task.
[0003] To achieve dual capabilities in both zero-shot and one-shot learning, two mainstream training approaches are commonly used: one involves obtaining independent training data for each learning mode, constructing different model architectures, and conducting phased training; the other involves decomposing input features into multiple attributes (such as duration, rhythm, and timbre in speech signals) and then controlling the performance of each attribute during data generation. However, both approaches have significant limitations, resulting in complex training methods and low training efficiency. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a sound model training method and a sound model training device, which only requires classifying the training method of each training data based on dynamic distribution parameters, so that the model can perform well in both zero-sample learning and single-sample learning modes. The training of the model in each mode is more efficient and easy to implement, and there is no need to construct different training data and training strategies for zero-sample learning and single-sample learning respectively, which simplifies the model training process.
[0005] In a first aspect, an embodiment of the present application provides a method for training a sound model, the method comprising: Obtaining dynamic distribution parameters, and determining a target training mode for each piece of training data based on the dynamic distribution parameters; Based on each piece of training data and a set of target training methods corresponding to each piece of training data, the initial model is trained to obtain a target model; wherein the target training methods include at least a single-sample learning mode and a zero-sample learning mode.
[0006] Furthermore, determining the target training mode for each piece of training data includes: For each piece of training data, a random number is obtained; wherein the dynamic distribution parameter and the random number are both between 0 and 1; When the random number corresponding to the training data is greater than or equal to the dynamic distribution parameter, determining that the target training mode is a single-sample learning mode; When the random number corresponding to the training data is smaller than the dynamic distribution parameter, the target training mode is determined to be a zero-sample learning mode.
[0007] Furthermore, when the target training mode is a single-sample learning mode, the initial model is trained based on a set of each training data and its corresponding target training mode, including: For each piece of training data, obtaining at least one target prompt word, and constructing a sound feature vector of the training data based on the at least one target prompt word; reconstructing the training data based on the sound feature vector to obtain target training data; The target training data is input into the initial model to perform single-sample learning on the initial model.
[0008] Furthermore, reconstructing the training data based on the sound feature vector to obtain target training data includes: The sound feature vector is concatenated at the beginning of the sentence of the training data to obtain the target training data; wherein the sound feature vector of each training data has the same dimension.
[0009] Furthermore, when the target training mode is a zero-sample learning mode, the initial model is trained based on a set of each training data and its corresponding target training mode, including: For each piece of training data, extract the sound representation information of the training data as the sound feature vector corresponding to the training data; The training data and the sound feature vector corresponding to the training data are simultaneously input into the initial model to perform zero-sample learning on the initial model.
[0010] Furthermore, the sound representation information of the training data is extracted from the training data, or the sound representation information of the training data is extracted from the Mel spectrum of the training data.
[0011] Furthermore, before determining the target training mode for each piece of training data, the sound model training method further includes: At least one prompt word is extracted from each piece of training data; wherein each prompt word describes the characteristic performance of the training data from different dimensions.
[0012] Furthermore, after obtaining the target model, the sound model training method further includes: Evaluate the single-shot learning mode performance and the zero-shot learning mode performance of the target model, and adjust the dynamic distribution parameters based on the performance evaluation results.
[0013] Furthermore, adjusting the dynamic distribution parameters based on the performance evaluation results includes: If the performance evaluation result of the target model in the single-sample learning mode does not reach a first preset evaluation threshold, reducing the value of the dynamic distribution parameter; If the performance evaluation result of the target model in the zero-sample learning mode does not reach a second preset evaluation threshold, the value of the dynamic distribution parameter is increased.
[0014] In a second aspect, an embodiment of the present application further provides a sound model training device, the sound model training device comprising: A training mode determination module is used to obtain dynamic distribution parameters and determine a target training mode for each piece of training data based on the dynamic distribution parameters; The model training module is used to train the initial model based on each training data and the set of target training methods corresponding to each training data to obtain a target model; wherein the target training method includes at least a single-sample learning mode and a zero-sample learning mode.
[0015] An embodiment of the present application provides a sound model training method and a sound model training device. First, dynamic distribution parameters are obtained, and the target training mode for each training data is determined based on the dynamic distribution parameters; then, based on each training data and a set of target training modes corresponding to each training data, the initial model is trained to obtain a target model; wherein the target training mode includes at least a single-sample learning mode and a zero-sample learning mode.
[0016] This application only needs to classify the training methods of each training data based on dynamic distribution parameters, and can use the same training data set and training strategy to implement the training of zero-sample learning mode and single-sample learning mode on the traditional generative model structure, so that the model can adapt to the zero-sample learning mode and the single-sample learning mode at the same time, so that the model can perform well in both zero-sample learning and single-sample learning modes without the need for complex feature decoupling or additional data preparation. Compared with traditional model training methods, the sound model training method provided by this application is more efficient and easy to implement for training in various modes of the model. There is no need to construct different training data and training strategies for zero-sample learning and single-sample learning respectively, which simplifies the model training process.
[0017] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A flow chart of a sound model training method provided in an embodiment of the present application; Figure 2 This is one of the structural diagrams of a sound model training device provided in an embodiment of the present application; Figure 3 This is a second structural diagram of a sound model training device provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, each other embodiment obtained by those skilled in the art without making creative work falls within the scope of protection of the present application.
[0021] First, the application scenarios to which this application is applicable are introduced. This application can be applied in the field of model training technology.
[0022] With the rapid development of artificial intelligence (AI), deep learning models have made significant progress in fields such as image recognition and natural language processing. Zero-shot learning and one-shot learning, emerging machine learning paradigms, can address the data scarcity and dynamic category expansion issues of traditional supervised learning. In zero-shot learning, a model can directly understand and reason about new inputs based on its generalization capabilities, even though it has never been exposed to a specific category or task during training. In one-shot learning, a model quickly adapts to new tasks using a single reference sample, which encodes or provides information about key features of the target task.
[0023] Research has found that to achieve dual capabilities in zero-shot and one-shot learning, two mainstream training approaches are commonly used: one is to obtain independent training data for each learning mode, construct different model architectures, and conduct phased training; the other is to decompose input features into multiple attributes (such as duration, rhythm, and timbre in speech signals) and then control the performance of each attribute during data generation. However, both methods have significant limitations, complex training methods, and low training efficiency.
[0024] Based on this, the embodiment of the present application provides a sound model training method, which enables the model to perform well in both zero-sample learning and single-sample learning modes. The training of the model is more efficient and easy to implement. There is no need to design different training strategies for zero-sample learning and single-sample learning respectively, which simplifies the model training process.
[0025] See also Figure 1 , Figure 1 This is a flow chart of a sound model training method provided in an embodiment of the present application. Figure 1 As shown in , the sound model training method provided in the embodiment of the present application includes: S101, obtaining dynamic distribution parameters, and determining a target training mode for each piece of training data according to the dynamic distribution parameters.
[0026] Regarding the above step S101, during specific implementation, a dynamic distribution parameter P is obtained, and a target training method for each piece of training data is determined based on the dynamic distribution parameter P. The training data here is an audio training sample used to train the initial sound model.
[0027] Specifically, with respect to the above step S101, determining the target training mode for each piece of training data includes: Step 1011: Obtain a random number for each piece of training data.
[0028] Regarding step 1011, during specific implementation, for each piece of training data, a random number R corresponding to the training data is obtained. Here, the dynamic distribution parameter P and the random number R are both between 0 and 1, that is, R, P∈[0,1].
[0029] Step 1012: When the random number corresponding to the training data is greater than or equal to the dynamic distribution parameter, the target training mode is determined to be a single-sample learning mode.
[0030] Step 1013: When the random number corresponding to the training data is smaller than the dynamic distribution parameter, it is determined that the target training mode is the zero-sample learning mode.
[0031] Here, one-shot learning refers to a model that may or may not have seen the target timbre during training, but can quickly adapt to new tasks or categories using a single reference sample during inference. For example, given a timbre sample, the model can adjust its output style based on that sample. Zero-shot learning refers to a model that has never been exposed to certain tasks or categories during training, but can still handle these new tasks or produce new results during inference. For example, it can generate voices with a similar timbre even though it has never been exposed to speech data of a certain timbre.
[0032] Regarding steps 1012 and 1013, in a specific implementation, when it is determined that the random number R corresponding to the training data is greater than or equal to the dynamic distribution parameter P, the target training mode for the training data is determined to be the single-shot learning mode. When it is determined that the random number R corresponding to the training data is less than the dynamic distribution parameter P, the target training mode for the training data is determined to be the zero-shot learning mode.
[0033] According to the sound model training method provided in this application, before determining the target training mode for each piece of training data, the sound model training method further includes: Extract at least one prompt word from each training data.
[0034] Regarding the above steps, in specific implementations, each piece of training data includes, in addition to the data portion, a qualification portion consisting of one or more prompt words. At least one prompt word is extracted from each piece of training data. Here, the prompt words are used to define the characteristic performance of the training data, and each prompt word describes the characteristic performance of the training data from different dimensions. For example, when the model to be trained is a TTS (Text-to-Speech) model, the dimensions of the prompt words in the training data may include: timbre, pitch, speaking rate, language, etc.
[0035] S102: Based on each piece of training data and a set of target training methods corresponding to each piece of training data, the initial model is trained to obtain a target model.
[0036] Here, the target model may be a sound model used for speech generation and / or music generation.
[0037] Regarding the above step S102, in specific implementation, after the target training method of each training data is determined, the initial model is trained based on each training data and the set of target training methods corresponding to each training data to obtain the target model.
[0038] Here, according to the embodiment provided by the present application, the trained model includes at least the following structures in sequence: an encoder, a decoder, and a waveform reconstruction module, wherein the encoder selects a learnable speaker encoder. The waveform reconstruction module can select a neural vocoder to convert several vectors output by the decoder into audible acoustic waveforms. The learnable speaker encoder is different from other audio encoders. Common audio encoders are usually pre-trained, but the learnable speaker encoder can be trained together with the model. In training and inference, its input is a segment of audio, and the output is a fixed-dimensional speaker embedding vector (embedding vector) used to represent the identity of the speaker.
[0039] In this way, according to the above steps S101-S102, it is only necessary to classify the training method of each training data based on the dynamic distribution parameters, so that the training of zero-sample learning mode and single-sample learning mode can be realized on the traditional generative model structure, so that the model can adapt to the zero-sample learning mode and the single-sample learning mode at the same time, so that the model can perform well in both zero-sample learning and single-sample learning modes without the need for complex feature decoupling or additional data preparation.
[0040] According to the embodiment provided by the present application, for the above step S102, when the target training mode is a single-sample learning mode, the initial model is trained based on a set of each training data and its corresponding target training mode, including: I: For each piece of training data, obtain at least one target prompt word, and construct a sound feature vector of the training data based on the at least one target prompt word.
[0041] Regarding step I above, when the single-sample learning mode is triggered, for each piece of training data, at least one target prompt word is obtained from at least one prompt word in the training data, and a sound feature vector for the training data is constructed based on the at least one target prompt word. In this way, when the training data is input into the initial model for training, the initial model can be trained based on both the sound feature vector and the audio data.
[0042] II: Reconstruct the training data based on the sound feature vector to obtain target training data.
[0043] Regarding step II above, in order to maintain the training effect, it is generally necessary to keep all different parts of the training data intact during implementation. Specifically, the sound feature vector can be concatenated at the beginning or end of the training data, and the concatenated data can be used to replace the original training data to obtain the target training data.
[0044] Furthermore, with respect to the above step II, reconstructing the training data based on the sound feature vector to obtain target training data includes: The sound feature vector is concatenated to the beginning of the sentence of the training data to obtain the target training data.
[0045] Preferably, the sound feature vector is spliced at the beginning of the sentence of the training data to obtain the target training data. Here, the sound feature vector of each training data has the same dimension. Specifically, in response to triggering the single-sample learning mode, a fixed-length prompt word is extracted from a limited part of the training data, and the training data is re-spliced. During splicing, the sound feature vector is spliced at the beginning of the sentence of the training data, rather than the end of the sentence. The uniform prompt word length and the uniform training data length can reduce the training overhead of the model. Specifically, if the various parts of each training data input to the model are kept in the same position and the lengths of the various parts are the same, the model does not need too much computing power to obtain the similarities of the training data, thereby reducing the learning difficulty of the model. The length of the data part of each training data is usually different. If the sound feature vector is spliced after the data part, it is easy to cause the position of the sound feature vector in each training data to be non-fixed. Relatively speaking, it is only necessary to extract sound feature vectors of the same length each time the training data is reconstructed to achieve the unification of the training data. Therefore, it is relatively easy to control the length of the sound feature vector. Setting the sound feature vector at the beginning of the sentence of the training data can ensure that the length and data position of the sound feature vector in each training data are relatively fixed.
[0046] III: Inputting the target training data into the initial model to perform single-sample learning on the initial model.
[0047] Regarding step III above, the target training data is input into the initial model for single-sample learning. This allows the initial model to understand the acoustic representation of the training data based on the cue word portion of the acoustic feature vector. Furthermore, because the actual acoustic features of the training data are directly used, the timbre information of the training data is more accurate.
[0048] According to the embodiment provided by the present application, for the above step S102, when the target training mode is the zero-sample learning mode, the training of the initial model based on the set of each training data and its corresponding target training mode includes: A: For each piece of training data, extract the sound representation information of the training data as the sound feature vector corresponding to the training data.
[0049] Regarding the above step A, during specific implementation, for each piece of training data, the sound representation information of the training data is extracted as the sound feature vector corresponding to the training data.
[0050] As an optional embodiment, the sound representation information of the training data is extracted from the training data, or the sound representation information of the training data is extracted from a mel-spectrogram of the training data, where the mel-spectrogram is converted from the training data. Alternatively, the sound representation information of the training data can be obtained from at least one prompt word in the training data.
[0051] B: The training data and the sound feature vector corresponding to the training data are simultaneously input into the initial model to perform zero-sample learning on the initial model.
[0052] With respect to the above step B, in specific implementation, when using the training data for training, the training data and the sound feature vector corresponding to the training data are simultaneously input into the initial model to perform zero-sample learning on the initial model. In this way, the model can also understand the relationship between sound features and audio performance based on the sound feature vector extracted from the training data. A representation vector representing the global sound representation is extracted from the training data. Although the model has not seen the timbre of the audio data during actual reasoning, the model has learned how to represent various sound features from the training data and the global sound representation, and can map the actual timbre obtained during reasoning to part of the sound feature vector for representing the timbre of the training data.
[0053] Preferably, in the single-sample learning mode and the zero-sample learning mode, the dimensions of the sound feature vectors are the same.
[0054] As an optional embodiment, after obtaining the target model, the sound model training method provided in the embodiment of the present application further includes: Evaluate the single-shot learning mode performance and the zero-shot learning mode performance of the target model, and adjust the dynamic distribution parameters based on the performance evaluation results.
[0055] For the above steps, in the specific implementation, after using the training data to train the initial model to obtain the target model, the single-sample learning mode performance and the zero-sample learning mode performance of the target model are evaluated to obtain the performance evaluation results, and the dynamic distribution parameters are dynamically adjusted based on the performance evaluation results.
[0056] Specifically, with respect to the above steps, adjusting the dynamic distribution parameters based on the performance evaluation results includes: (1) If the performance evaluation result of the target model in the single-sample learning mode does not reach a first preset evaluation threshold, the value of the dynamic distribution parameter is reduced.
[0057] (2) If the performance evaluation result of the target model in the zero-sample learning mode does not reach a second preset evaluation threshold, the value of the dynamic distribution parameter is increased.
[0058] For the above steps (1) to (2), in the specific implementation, if the performance evaluation result of the target model in the single-sample learning mode does not reach the first preset evaluation threshold, the value of the dynamic distribution parameter P is reduced so that the model is more inclined to the zero-sample learning mode when the training data is used for the next round of training, and the target model is retrained by returning to step S102. If the performance evaluation result of the target model in the zero-sample learning mode does not reach the second preset evaluation threshold, the value of the dynamic distribution parameter P is increased so that the model is more inclined to the single-sample learning mode when the training data is used for the next round of training, and the target model is retrained by returning to step S102. Specifically, the dynamic adjustment of the dynamic distribution parameter is based on the generation performance of the target model under the zero-sample learning mode and the single-sample learning mode. For example, in the field of TTS models, performance evaluation can be performed based on the performance parameters of the TTS model under the single-sample learning mode and the zero-sample learning mode, such as WER (word error rate) and SIM (speech similarity). For training modes with poor performance, the value of the dynamic distribution parameter can be adjusted to tilt more training samples for training, so as to flexibly control the performance of the model under the two modes.
[0059] Here, as an optional embodiment, when the performance evaluation results of the target model in both modes do not meet the standards, the value of the dynamic distribution parameter P can be adjusted to 0 or 1, and the two modes can be retrained separately. If the performance evaluation still does not meet the standards after retraining, it is necessary to adjust the training data, or adjust the training architecture or other model parameters of the target model and then retrain.
[0060] The sound model training method provided in the embodiment of the present application first obtains the dynamic distribution parameters and determines the target training mode for each training data based on the dynamic distribution parameters P; then, based on each training data and the set of target training modes corresponding to each training data, the initial model is trained to obtain the target model; wherein the target training mode includes at least a single-sample learning mode and a zero-sample learning mode.
[0061] This application only needs to classify the training methods of each training data based on dynamic distribution parameters, and can use the same training data set and training strategy to implement the training of zero-sample learning mode and single-sample learning mode on the traditional generative model structure, so that the model can adapt to the zero-sample learning mode and the single-sample learning mode at the same time, so that the model can perform well in both zero-sample learning and single-sample learning modes without the need for complex feature decoupling or additional data preparation. Compared with traditional model training methods, the sound model training method provided by this application is more efficient and easy to implement for training in various modes of the model. There is no need to construct different training data and training strategies for zero-sample learning and single-sample learning respectively, which simplifies the model training process.
[0062] See also Figure 2 、 Figure 3 , Figure 2 This is one of the structural diagrams of a sound model training device provided in an embodiment of the present application. Figure 3 This is a second structural diagram of a sound model training device provided in an embodiment of the present application. Figure 2 As shown in , the sound model training device 200 includes: The training mode determination module 201 is used to obtain dynamic distribution parameters and determine the target training mode for each training data according to the dynamic distribution parameters; The model training module 202 is used to train the initial model based on each training data and a set of target training methods corresponding to each training data to obtain a target model; wherein the target training method includes at least a single-sample learning mode and a zero-sample learning mode.
[0063] Furthermore, when the model training module 202 is used to determine the target training mode for each piece of training data, the model training module 202 is also used to: For each piece of training data, a random number is obtained; wherein the dynamic distribution parameter and the random number are both between 0 and 1; When the random number corresponding to the training data is greater than or equal to the dynamic distribution parameter, determining that the target training mode is a single-sample learning mode; When the random number corresponding to the training data is smaller than the dynamic distribution parameter, the target training mode is determined to be a zero-sample learning mode.
[0064] Furthermore, when the target training mode is a single-sample learning mode, the model training module 202 is further configured to: For each piece of training data, obtaining at least one target prompt word, and constructing a sound feature vector of the training data based on the at least one target prompt word; reconstructing the training data based on the sound feature vector to obtain target training data; The target training data is input into the initial model to perform single-sample learning on the initial model.
[0065] Furthermore, when the model training module 202 is used to reconstruct the training data based on the sound feature vector to obtain target training data, the model training module 202 is also used to: The sound feature vector is concatenated at the beginning of the sentence of the training data to obtain the target training data; wherein the sound feature vector of each training data has the same dimension.
[0066] Furthermore, when the target training mode is a zero-sample learning mode, the model training module 202 is further configured to: For each piece of training data, extract the sound representation information of the training data as the sound feature vector corresponding to the training data; The training data and the sound feature vector corresponding to the training data are simultaneously input into the initial model to perform zero-sample learning on the initial model.
[0067] Furthermore, the sound representation information of the training data is extracted from the training data, or the sound representation information of the training data is extracted from the Mel spectrum of the training data.
[0068] See also Figure 3 The sound model training device 200 further includes a prompt word acquisition module 203. Before determining the target training mode for each training data, the prompt word acquisition module 203 is used to: At least one prompt word is extracted from each piece of training data; wherein each prompt word describes the characteristic performance of the training data from different dimensions.
[0069] See also Figure 3The sound model training device 200 further includes a parameter adjustment module 204. After obtaining the target model, the parameter adjustment module 204 is used to: Evaluate the single-shot learning mode performance and the zero-shot learning mode performance of the target model, and adjust the dynamic distribution parameters based on the performance evaluation results.
[0070] Furthermore, when the parameter adjustment module 204 is used to adjust the dynamic distribution parameters based on the performance evaluation result, the parameter adjustment module 204 is further used to: If the performance evaluation result of the target model in the single-sample learning mode does not reach a first preset evaluation threshold, reducing the value of the dynamic distribution parameter; If the performance evaluation result of the target model in the zero-sample learning mode does not reach a second preset evaluation threshold, the value of the dynamic distribution parameter is increased.
[0071] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4 As shown in FIG, the electronic device 400 includes a processor 410 , a memory 420 and a bus 430 .
[0072] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, the above-mentioned Figure 1 The steps of the sound model training method in the method embodiment shown are specifically implemented in accordance with the method embodiment and will not be described in detail here.
[0073] The embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 1 The steps of the sound model training method in the method embodiment shown are specifically implemented in accordance with the method embodiment and will not be described in detail here.
[0074] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0075] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.
[0076] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0077] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0078] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0079] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. These modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A sound model training method, characterized in that: The sound model training method comprises: Obtaining dynamic distribution parameters, and determining a target training mode for each piece of training data based on the dynamic distribution parameters; Based on each piece of training data and a set of target training methods corresponding to each piece of training data, the initial model is trained to obtain a target model; wherein the target training methods include at least a single-sample learning mode and a zero-sample learning mode.
2. The sound model training method according to claim 1, characterized in that Determining the target training mode for each piece of training data includes: For each piece of training data, a random number is obtained; wherein the dynamic distribution parameter and the random number are both between 0 and 1; When the random number corresponding to the training data is greater than or equal to the dynamic distribution parameter, determining that the target training mode is a single-sample learning mode; When the random number corresponding to the training data is smaller than the dynamic distribution parameter, the target training mode is determined to be a zero-sample learning mode.
3. The sound model training method according to claim 2, characterized in that When the target training mode is a single-sample learning mode, the initial model is trained based on a set of each training data and its corresponding target training mode, including: For each piece of training data, obtaining at least one target prompt word, and constructing a sound feature vector of the training data based on the at least one target prompt word; reconstructing the training data based on the sound feature vector to obtain target training data; The target training data is input into the initial model to perform single-sample learning on the initial model.
4. The sound model training method according to claim 3, characterized in that The reconstructing the training data based on the sound feature vector to obtain target training data includes: The sound feature vector is concatenated at the beginning of the sentence of the training data to obtain the target training data; wherein the sound feature vector of each training data has the same dimension.
5. The sound model training method according to claim 1, characterized in that When the target training mode is a zero-sample learning mode, the initial model is trained based on a set of each training data and its corresponding target training mode, including: For each piece of training data, extract the sound representation information of the training data as the sound feature vector corresponding to the training data; The training data and the sound feature vector corresponding to the training data are simultaneously input into the initial model to perform zero-sample learning on the initial model.
6. The sound model training method according to claim 5, characterized in that: The sound representation information of the training data is extracted from the training data, or the sound representation information of the training data is extracted from the Mel spectrum of the training data.
7. The sound model training method according to claim 3, characterized in that: Before determining the target training mode for each piece of training data, the sound model training method further includes: At least one prompt word is extracted from each piece of training data; wherein each prompt word describes the characteristic performance of the training data from different dimensions.
8. The sound model training method according to claim 1, characterized in that After obtaining the target model, the sound model training method further includes: Evaluate the single-shot learning mode performance and the zero-shot learning mode performance of the target model, and adjust the dynamic distribution parameters based on the performance evaluation results.
9. The sound model training method according to claim 8, characterized in that: The adjusting the dynamic distribution parameters based on the performance evaluation result includes: If the performance evaluation result of the target model in the single-sample learning mode does not reach a first preset evaluation threshold, reducing the value of the dynamic distribution parameter; If the performance evaluation result of the target model in the zero-sample learning mode does not reach a second preset evaluation threshold, the value of the dynamic distribution parameter is increased.
10. A sound model training device, characterized in that: The sound model training device comprises: A training mode determination module is used to obtain dynamic distribution parameters and determine a target training mode for each piece of training data based on the dynamic distribution parameters; The model training module is used to train the initial model based on each training data and the set of target training methods corresponding to each training data to obtain a target model; wherein the target training method includes at least a single-sample learning mode and a zero-sample learning mode.
Citation Information
Patent Citations
Model training and content auditing management system based on intelligent data acquisition
CN119646601A
Zero-shot intent classification using a semantic similarity aware contrastive loss and large language model
US20250095638A1
Methods and systems for enhancing multimodal capabilities in large language models
US20250140238A1
Dynamic mode decomposition for evaluating early stop in training of machine learning models
WO2025106937A1