Model training method and device of task processing model, equipment and storage medium

CN119106135BActive Publication Date: 2026-09-25上海明胜品智人工智能科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411193999.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2026-09-25
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

这样,基于流水线式的学习模式,往往会使得任务模型在整体训练过程中容易产生识别错误的级联累积问题,导致每一级子任务模型产生的识别错误都会顺延到下一层级,进而造成流水线式的学习模式下的整体模型训练效果不理想

Benefits of technology

[0042]本申请实施例提供的一种任务处理模型的模型训练方法、装置、设备及存储介质,先获取训练语料,并将训练语料输入至共享特征提取模型中,通过共享特征提取模型提取训练语料中多种类别的共享特征信息;再按照预设的输入方式,将多种类别的共享特征信息和基于训练语料标注的训练文本信息输入至多个子任务模型中,并行对多个子任务模型进行训练,以使多个子任务模型的整体损失函数满足训练截止条件;在多个子任务模型独立训练的过程中,获取每一子任务模型的任务训练损失,根据每一子任务模型的任务训练损失的梯度变化情况,对该子任务模型的权重系数进行调整,以使多个子任务模型的训练率位于同一数值范围区间内,直至多个子任务模型的整体损失函数满足所述训练截止条件,将训练好的多个子任务模型作为训练好的任务处理模型。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119106135B_ABST
    Figure CN119106135B_ABST
Patent Text Reader

Abstract

The application provides a model training method and device of a task processing model, equipment and a storage medium. The method comprises: extracting shared feature information of multiple categories in training corpus through a shared feature extraction model; when task types to which different sub-task models to be executed belong are not distinguished, the shared feature information of the multiple categories and training text information labeled based on the training corpus are synchronously input into each sub-task model at a first layer model input node of each sub-task model in a second input mode of first layer input, and multiple sub-task models are trained in parallel until an overall loss function of the multiple sub-task models meets a training stop condition. In this way, while ensuring that each sub-task model can be independently trained, the application can provide different sub-task models with multiple shared feature information related to the sub-task executed by the sub-task model, thereby improving the overall model training effect of the task processing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and more specifically, to a model training method, apparatus, device, and storage medium for a task processing model. Background Technology

[0002] In the field of natural language processing, when using text information as model training data, models can be trained to perform various different types of training tasks on text information, such as text recognition tasks to identify named entities in specific business scenarios, and sentiment classification tasks to identify the emotions expressed in different sentences. Specifically, taking the aforementioned text recognition and sentiment classification tasks as examples, although the model training difficulty and objectives of these two types of tasks are different, considering that text recognition of named entities (equivalent to semantic recognition of word segmentation) is the foundation of sentence sentiment recognition, related technologies typically employ a holistic training approach for models with correlation between these training tasks.

[0003] Currently, in related technologies, a "pipeline-style learning model" is commonly used to train the aforementioned task models as a whole. Here, taking text recognition and sentiment classification tasks as examples, when using a pipeline-style learning model for overall training, the first-level sub-task model is first trained to perform named entity recognition. Then, using the recognition results of the first-level sub-task model, the second-level sub-task model is further trained to learn to recognize and classify the semantic sentiment and emotional expression themes of different named entities in the text information. This pipeline-style learning model often leads to a cascading accumulation of recognition errors during the overall training process, causing errors from each sub-task model to propagate to the next level, resulting in unsatisfactory overall model training performance under this pipeline-style learning model. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a model training method, apparatus, device and storage medium for a task processing model, so as to provide various shared feature information related to the sub-tasks they execute while ensuring that each sub-task model can be trained independently through the constructed multi-task learning model framework, which is conducive to improving the overall model training effect of the task processing model.

[0005] In a first aspect, embodiments of this application provide a model training method for a task processing model, applied to a multi-task learning model framework. The multi-task learning model framework includes a task processing model and a pre-trained shared feature extraction model. The task processing model includes multiple sub-task models. The model training method includes:

[0006] Acquire training corpus and input the training corpus into the shared feature extraction model, and extract shared feature information of multiple categories from the training corpus through the shared feature extraction model;

[0007] According to the preset input method, the shared feature information of the multiple categories and the training text information labeled based on the training corpus are input into the multiple sub-task models, and the multiple sub-task models are trained in parallel so that the overall loss function of the multiple sub-task models meets the training cutoff condition.

[0008] During the independent training of the multiple sub-task models, the task training loss of each sub-task model is obtained. Based on the gradient change of the task training loss of each sub-task model, the weight coefficient of the sub-task model is adjusted so that the training rates of the multiple sub-task models are within the same numerical range, until the overall loss function of the multiple sub-task models satisfies the training cutoff condition. The trained multiple sub-task models are then used as the trained task processing model.

[0009] In one optional implementation, the shared feature information of the multiple categories includes: character feature vectors after the training corpus is segmented into character sequences; word feature vectors in the training corpus representing syntactic dependencies between words; and sentence feature vectors in the training corpus.

[0010] In one alternative implementation, the multiple categories of shared feature information to be extracted by the shared feature extraction model are determined by the following method:

[0011] Based on the target task dependencies between the multiple subtasks to be executed in the multiple subtask models, multiple information categories corresponding to the target task dependencies are determined from a preset task dependency table as multiple categories of shared feature information to be extracted; wherein, the task dependency table pre-stores multiple information categories corresponding to multiple task dependencies.

[0012] In one alternative implementation, the multiple categories of shared feature information to be extracted by the shared feature extraction model are determined by the following method:

[0013] Based on the target task to be executed by the task processing model, the multiple sub-task models included in the task processing model are used as the first search space, and the ability to execute the target task is used as the first search strategy. The neural network structure search is performed on the sub-task model combination methods among different sub-task models in the first search space to obtain the optimal sub-task model combination method that conforms to the first search strategy.

[0014] Each subtask model included in the optimal subtask model combination method is taken as the first subtask model;

[0015] Based on the first task dependency relationship between the subtasks to be executed in each first subtask model, multiple information categories corresponding to the first task dependency relationship are determined from the preset task dependency relationship table as multiple categories of the shared feature information to be extracted.

[0016] In one alternative implementation, the multiple categories of shared feature information to be extracted by the shared feature extraction model are determined by the following method:

[0017] Based on the target task to be executed by the task processing model, obtain various text feature information related to completing the target task;

[0018] Using the aforementioned multiple text feature information as the second search space, and taking the ability of the multiple sub-task models to complete the target task based on the information combination of different text feature information as the second search strategy, a neural network structure search is performed on the information combination methods between different text feature information in the second search space to obtain the optimal information combination method that conforms to the second search strategy.

[0019] The information category to which each text feature information included in the optimal information combination method belongs is taken as multiple categories of the shared feature information to be extracted.

[0020] In one optional implementation, the step of inputting the shared feature information of the multiple categories and the training text information labeled based on the training corpus into the multiple sub-task models according to a preset input method includes:

[0021] At the first-level model input node of each sub-task model, the training text information is input into each sub-task model;

[0022] The shared feature information of the various categories is input hierarchically to different training nodes in each sub-task model according to the correspondence between information categories and training nodes, using a hierarchical input first input method; wherein, the different training nodes in each sub-task model are sorted according to the layers of the neural network in the sub-task model from shallow to deep.

[0023] In an optional implementation, the step of inputting the shared feature information of the multiple categories and the training text information labeled based on the training corpus into the multiple sub-task models according to a preset input method further includes:

[0024] At the first-level model input node of each sub-task model, the shared feature information of the multiple categories and the training text information are synchronously input into each sub-task model in the second input method of the first-level input.

[0025] In one optional implementation, the step of synchronously inputting the shared feature information of the multiple categories and the training text information into each of the sub-task models using the second input method of the first layer input includes:

[0026] When the task type of the subtask to be executed by different subtask models is not distinguished, the shared feature information of the multiple categories and the training text information are synchronously input into each subtask model using the second input method.

[0027] or,

[0028] When distinguishing the task types to which the subtasks to be executed in different subtask models belong, for each subtask model, based on the subtasks to be executed in that subtask model, the target shared feature information that matches the subtasks to be executed in the shared feature information of the multiple categories is determined.

[0029] Using the second input method, the shared feature information of the various categories, the training text information, and the target shared feature information are simultaneously input into the subtask model.

[0030] In one optional implementation, the overall loss function of the plurality of sub-task models is determined based on the product of the gradient of the task training loss of each sub-task model and the weight coefficient of that sub-task model in the multi-task learning model framework; adjusting the weight coefficient of the sub-task model according to the gradient change of the task training loss of each sub-task model includes:

[0031] For each subtask model, the gradient of the task training loss of that subtask model is used as the target gradient, and the periodic change amplitude of the target gradient within the gradient detection period is obtained.

[0032] When the periodic change of the target gradient is detected to be greater than or equal to the change of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted in a descending manner according to the gradient descent adjustment coefficient.

[0033] When the periodic change amplitude of the target gradient is detected to be less than the change amount of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted by increasing the gradient increase adjustment coefficient.

[0034] In one optional implementation, each subtask model in the task processing model is used to execute a corresponding subtask, and different subtask models cooperate with each other to process the target task to be executed by the task processing model; when the target task is related to the semantic sentiment expressed by the text information, the task processing model includes at least one named entity recognition model and one sentiment classification model; wherein, the named entity recognition model is used to perform a text recognition task for the named entities included in the training text information; the sentiment classification model is used to perform a sentiment classification task for the sentiment represented by each sentence in the training text information.

[0035] Secondly, embodiments of this application provide a model training apparatus for a task processing model, applied to a multi-task learning model framework. The multi-task learning model framework includes a task processing model and a pre-trained shared feature extraction model. The task processing model includes multiple sub-task models. The model training apparatus includes:

[0036] An extraction module is used to acquire training corpus and input the training corpus into the shared feature extraction model, and extract shared feature information of multiple categories from the training corpus through the shared feature extraction model;

[0037] The input module is used to input the shared feature information of the multiple categories and the training text information labeled based on the training corpus into the multiple sub-task models according to a preset input method, and train the multiple sub-task models in parallel so that the overall loss function of the multiple sub-task models meets the training cutoff condition.

[0038] The training module is used to obtain the task training loss of each of the multiple sub-task models during the independent training process of the multiple sub-task models, and adjust the weight coefficient of the sub-task model according to the gradient change of the task training loss of each sub-task model, so that the training rate of the multiple sub-task models is within the same numerical range, until the overall loss function of the multiple sub-task models meets the training cutoff condition, and the trained multiple sub-task models are used as the trained task processing model.

[0039] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the model training method described above.

[0040] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the model training method described above.

[0041] The technical solutions provided by the embodiments of this application may include the following beneficial effects:

[0042] This application provides a model training method, apparatus, device, and storage medium for a task processing model. First, training corpus is acquired and input into a shared feature extraction model. The shared feature extraction model extracts multiple categories of shared feature information from the training corpus. Then, according to a preset input method, the multiple categories of shared feature information and training text information labeled based on the training corpus are input into multiple sub-task models. These sub-task models are trained in parallel to ensure that the overall loss function of the multiple sub-task models meets the training cutoff condition. During the independent training of the multiple sub-task models, the task training loss of each sub-task model is acquired. Based on the gradient change of the task training loss of each sub-task model, the weight coefficients of that sub-task model are adjusted to ensure that the training rates of the multiple sub-task models are within the same numerical range, until the overall loss function of the multiple sub-task models meets the training cutoff condition. The trained multiple sub-task models are then used as the trained task processing model.

[0043] In this way, the multi-task learning model framework constructed in this application can ensure that each sub-task model can be trained independently, while providing different sub-task models with a variety of shared feature information related to the sub-tasks they perform, which is beneficial to improving the overall model training effect of the task processing model.

[0044] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0045] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 A schematic flowchart of a model training method for a task processing model provided in an embodiment of this application is shown.

[0047] Figure 2 This illustration shows a flowchart of a method for dynamically adjusting the weight coefficients of each sub-task model in a multi-task learning model framework, as provided in an embodiment of this application.

[0048] Figure 3A flowchart illustrating the first neural network structure search method provided in an embodiment of this application is shown;

[0049] Figure 4 A flowchart illustrating the second neural network structure search method provided in an embodiment of this application is shown;

[0050] Figure 5 The illustration shows a flowchart of a method for inputting shared feature information according to a first input method provided in an embodiment of this application;

[0051] Figure 6 This illustration shows a schematic diagram of the structure of a model training device for a task processing model provided in an embodiment of this application;

[0052] Figure 7 A schematic diagram of the structure of a computer device 700 provided in an embodiment of this application is shown. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0054] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0055] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0056] Currently, in related technologies, a "pipeline-style learning model" is commonly used to train the aforementioned task models as a whole. Here, taking text recognition and sentiment classification tasks as examples, when using a pipeline-style learning model for overall training, the first-level sub-task model is first trained to perform named entity recognition. Then, using the recognition results of the first-level sub-task model, the second-level sub-task model is further trained to learn to recognize and classify the semantic sentiment and emotional expression themes of different named entities in the text information. This pipeline-style learning model often leads to a cascading accumulation of recognition errors during the overall training process, causing errors from each sub-task model to propagate to the next level, resulting in unsatisfactory overall model training performance under this pipeline-style learning model.

[0057] Based on this, embodiments of this application provide a model training method, apparatus, device, and storage medium for a task processing model. First, training corpus is acquired and input into a shared feature extraction model. The shared feature extraction model extracts multiple categories of shared feature information from the training corpus. Then, according to a preset input method, the multiple categories of shared feature information and training text information labeled based on the training corpus are input into multiple sub-task models, and these models are trained in parallel to ensure that the overall loss function of the multiple sub-task models meets the training cutoff condition. During the independent training of the multiple sub-task models, the task training loss of each sub-task model is acquired. Based on the gradient change of the task training loss of each sub-task model, the weight coefficients of that sub-task model are adjusted to ensure that the training rates of the multiple sub-task models are within the same numerical range, until the overall loss function of the multiple sub-task models meets the training cutoff condition. The trained multiple sub-task models are then used as the trained task processing model.

[0058] In this way, the multi-task learning model framework constructed in this application can ensure that each sub-task model can be trained independently, while providing different sub-task models with a variety of shared feature information related to the sub-tasks they perform, which is beneficial to improving the overall model training effect of the task processing model.

[0059] The following is a detailed description of a model training method, apparatus, device, and storage medium for a task processing model provided in the embodiments of this application.

[0060] Reference Figure 1 As shown, Figure 1This illustration shows a flowchart of a model training method for a task processing model provided in an embodiment of this application. The model training method is applied to a multi-task learning model framework, which includes a task processing model and a pre-trained shared feature extraction model. The task processing model includes multiple sub-task models. The model training method includes steps S101-S103. Specifically:

[0061] S101, Obtain training corpus and input the training corpus into the shared feature extraction model, and extract shared feature information of multiple categories from the training corpus through the shared feature extraction model.

[0062] Here, the training processes of different sub-task models in the multi-task learning model framework are independent of each other. That is, the sub-tasks to be executed by different sub-task models can be the same or different. This application does not impose any limitations on this.

[0063] In this embodiment of the application, the training corpus obtained in step S101 represents: the corpus information of the same text type used by the above-mentioned multiple sub-task models during the training process; the information category of the above-mentioned shared feature information is determined according to the task dependency relationship between the sub-tasks to be executed by different sub-task models.

[0064] Specifically, in the embodiments of this application, the above-mentioned shared feature information of multiple categories may include: character feature vectors after the training corpus is segmented into character sequences; word feature vectors in the training corpus representing the syntactic dependencies between words; and sentence feature vectors in the training corpus.

[0065] Regarding the aforementioned multiple sub-task models, it should be noted that, in this embodiment of the application, the sub-task models in the multi-task learning model framework are not a collection of arbitrary models without any limitations; that is, the specific technical application scenario to which this embodiment of the application applies (i.e., the model scope of the aforementioned multiple sub-task models) is: multiple sub-task models under a multi-task learning model framework that "perform and process tasks related to text information" and "are trained using corpus information of the same text type" (equivalent to the aforementioned "multiple sub-task models are used to perform different sub-tasks based on corpus information of the same text type during the training process").

[0066] Here, regarding the character feature vectors, word feature vectors, and sentence feature vectors in the shared feature information mentioned above, it is necessary to clarify the following:

[0067] The character feature vector can be obtained by segmenting the training corpus with any fixed step size. For example, the character feature vector can be obtained by segmenting the training corpus with a fixed step size of 2 characters or with a fixed step size of 3 characters. The specific segmentation method of the above character feature vector is not limited in this application embodiment.

[0068] The word feature vector can be either a high-order word feature vector obtained by performing dependency parsing on the training corpus, or a simple word feature vector obtained by common word segmentation methods. This application does not limit the specific vector form of the word feature vector.

[0069] Sentence feature vectors can represent high-dimensional feature vectors that map different sentences in the training corpus under multiple dimensional features. This application does not limit the specific dimensional values ​​of the vectors in the sentence feature vectors.

[0070] Specifically, regarding the specific information categories of the shared feature information mentioned above, it should also be noted that in special cases, when the subtasks to be executed by two different subtask models are the same:

[0071] In a first optional implementation, the text feature information required to perform the subtask can be determined, which is the shared feature information corresponding to the two different subtask models.

[0072] In a second optional implementation, it can also be determined that there is no task dependency between the subtasks to be executed in the two different subtask models, that is, it can be determined that there is no shared feature information to be extracted between the two different subtask models.

[0073] S102, according to a preset input method, the shared feature information of the multiple categories and the training text information labeled based on the training corpus are input into the multiple sub-task models, and the multiple sub-task models are trained in parallel so that the overall loss function of the multiple sub-task models meets the training cutoff condition.

[0074] Here, the specific annotation method of the training corpus is determined according to the subtask to be executed by the subtask model to be input. For example, if subtask model A is used to perform a text recognition task for named entities included in the training corpus, then the named entities included in the training corpus are annotated, and the annotated training corpus is input into subtask model A as training text information.

[0075] Here, the aforementioned preset input methods include at least: a first input method for layered input and a second input method for first-layer input.

[0076] Specifically, in this embodiment of the application, when the preset input method is the first input method described above, step S102 can be executed according to the following method 1:

[0077] Method 1: At the first-level model input node of each sub-task model, the training text information is input into each sub-task model;

[0078] The shared feature information of the various categories is input hierarchically to different training nodes in each subtask model according to the correspondence between information categories and training nodes, using a hierarchical input first input method.

[0079] Here, the different training nodes in each subtask model are sorted according to the layers of the neural network in the model from shallow to deep; as an optional embodiment, for shared feature information of different information categories, it can be set that the lower the information content of the shared feature information, the shallower the neural network layer of the training node corresponding to the information category of the shared feature information.

[0080] For example, taking the shared feature information of sub-task model A and sub-task model B as character feature vectors and word feature vectors, if sub-task model A is a 3-layer neural network model composed of neural network a, neural network b, and neural network c, and the depth order of the neural networks in sub-task model A is: neural network a < neural network b < neural network c, then when neural network a is the first layer of neural network A, the character feature vector with lower information content and the training text information are input into sub-task model A from the input node where neural network a is located; the word feature vector with higher information content is input into sub-task model A from the input node where neural network b is located.

[0081] Specifically, in this embodiment of the application, when the preset input method is the second input method described above, step S102 can be executed according to the following method 2:

[0082] Method 2: At the first-layer model input node of each sub-task model, the shared feature information of the multiple categories and the training text information are synchronously input into each sub-task model using the second input method of the first-layer input.

[0083] For example, taking the shared feature information corresponding to sub-task model A and sub-task model B as character feature vectors and word feature vectors, according to the second input method described above, even if the specific model structure of sub-task model A (i.e., the number of neural network layers in the model) is uncertain, the character feature vectors, word feature vectors, and training text information can be directly input into sub-task model A from the first layer model input node (such as the input node of the lowest layer neural network).

[0084] S103, during the independent training of the multiple sub-task models, the task training loss of each sub-task model is obtained, and the weight coefficient of the sub-task model is adjusted according to the gradient change of the task training loss of each sub-task model so that the training rate of the multiple sub-task models is within the same numerical range, until the overall loss function of the multiple sub-task models satisfies the training cutoff condition, and the trained multiple sub-task models are used as the trained task processing model.

[0085] It should be noted that the model training process of each sub-task model within the multi-task learning model framework is independent of each other. Therefore, the task training loss function used by different sub-task models can be the same or different. This application does not impose any limitations on this.

[0086] In this embodiment, the overall loss function of the plurality of sub-task models is determined by multiplying the gradient of the task training loss of each sub-task model with the weight coefficient of the sub-task model in the multi-task learning model framework.

[0087] Here, considering that the subtasks to be executed by different subtask models may be different, and the more difficult the subtask model is, the more complex its specific model structure will be. Therefore, it is difficult to balance the gradient change rate of the task training loss of the subtask model (equivalent to the gradient magnitude of the backpropagation of the model training task) between subtask models with different model structures (it is difficult to guarantee that the model converges at a similar or the same training speed). In other words, it is difficult to keep the training rate of multiple subtask models synchronized (i.e., the training rate is the same or the training rate is within the same numerical range) under the same multi-task learning model framework.

[0088] Based on this, in order to ensure that sub-task models with different levels of structural complexity can be trained at similar training speeds (i.e., training rates within the same numerical range) within the multi-task learning model framework, and to improve the learning sufficiency of each sub-task model, in one optional implementation, it can also be done as follows: Figure 2 The dynamic adjustment method shown is used to dynamically adjust the weight coefficients of each sub-task model in the multi-task learning model framework. Specifically:

[0089] Reference Figure 2 As shown, Figure 2 This illustration shows a flowchart of a method for dynamically adjusting the weight coefficients of each sub-task model in a multi-task learning model framework, as provided in an embodiment of this application. The method includes steps S201-S203; specifically:

[0090] S201, for each sub-task model, the gradient of the task training loss of the sub-task model is used as the target gradient, and the periodic change amplitude of the target gradient within the gradient detection period is obtained.

[0091] Here, under the multi-task learning model framework, the gradient detection period is the same for different sub-task models. That is, within the same gradient detection period, the gradient of the task training loss of each sub-task model is detected. The specific time length of the gradient detection period is not limited in this embodiment.

[0092] S202, when the periodic change amplitude of the target gradient is detected to be greater than or equal to the change amount of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted in a descending manner according to the gradient descent adjustment coefficient.

[0093] Here, when the periodic change of the target gradient is detected to be greater than or equal to the change of the reference gradient, it can be determined that the gradient of the task training loss of the sub-task model has changed significantly (i.e. increased significantly) within the current gradient detection period. At this time, by dynamically adjusting the weight coefficients of the sub-task model in a decreasing manner, the proportion of the task training loss of the sub-task model in the overall loss of the multi-task learning model framework (or task processing model) can be kept in dynamic balance (equivalent to the product of the gradient of the task training loss of the sub-task model and the weight coefficients of the sub-task model). This ensures that sub-task models with different levels of model structure complexity can be trained at similar training speeds within the multi-task learning model framework, thereby improving the model learning sufficiency of each sub-task model.

[0094] S203, when the periodic change amplitude of the target gradient is detected to be less than the change amount of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted by increasing the gradient increase adjustment coefficient.

[0095] Here, corresponding to step S202 above, when the periodic change amplitude of the target gradient is detected to be less than the change amount of the reference gradient, it can be determined that within the current gradient detection period, the gradient of the task training loss of the sub-task model has also changed significantly (i.e., decreased significantly). At this time, by dynamically adjusting the weight coefficients of the sub-task model in an increasing manner, the proportion of the task training loss of the sub-task model in the overall loss of the multi-task learning model framework (equivalent to the product of the gradient of the task training loss of the sub-task model and the weight coefficients of the sub-task model) can also reach a dynamic balance. Thus, it is ensured that sub-task models with different model structure complexities can be trained at similar training speeds under the multi-task learning model framework, thereby improving the model learning sufficiency of each sub-task model.

[0096] To more clearly illustrate the implementation details of steps S101-S103 in the embodiments of this application, the following section uses the "semantic sentiment analysis task," which occurs frequently in text information processing, as the overall learning task of the above multi-task learning model framework, to provide a detailed description of the implementation details of steps S101-S103:

[0097] First, it should be noted that in the embodiments of this application, each sub-task model in the task processing model is used to execute the corresponding sub-task, and different sub-task models cooperate with each other to process the target task to be executed by the task processing model. The target task to be executed by the task processing model can be any learning task related to "text information processing", and is not limited to the above-mentioned "semantic sentiment analysis task". Here, a relatively complex text information processing task (i.e., "semantic sentiment analysis" task) is selected as an example to more clearly highlight the specific implementation details of the embodiment of this application in the process of "training multiple sub-task models". The embodiment of this application does not limit which type of "text information processing task" the target task to be executed by the task processing model belongs to.

[0098] Regarding the implementation process of step S101 above, when the target task to be executed by the task processing model is related to the semantic sentiment expressed by the text information, the task processing model (i.e., the multiple sub-task models) includes at least one named entity recognition model and one sentiment classification model; wherein, the named entity recognition model is used to perform a text recognition task for the named entities included in the training text information; the sentiment classification model is used to perform a sentiment classification task for the sentiment represented by each sentence in the training text information.

[0099] Based on this, in the embodiments of this application, when the pre-training methods for the shared feature extraction model are different, according to the different model training methods, in step S101, at least the following three different optional methods can be used to determine the specific information category of the shared feature information to be extracted by the shared feature extraction model, specifically:

[0100] Option (1): Based on the target task dependency relationship between the multiple sub-tasks to be executed in the multiple sub-task model, determine the multiple information categories corresponding to the target task dependency relationship from the preset task dependency relationship table as the multiple categories of the shared feature information to be extracted.

[0101] Here, the task dependency table pre-stores various information categories corresponding to multiple task dependencies. For example, the task dependency table pre-stores various information categories corresponding to task dependency a as: information categories x1, x2 and x3.

[0102] Regarding the implementation of the above optional method (1), here, taking multiple sub-task models in the task processing model as: Named Entity Recognition Model M and Sentiment Classification Model Q as an example, in the implementation process of the above optional method (1), Named Entity Recognition Model M needs to perform a text recognition task for the named entities included in the training text information; Sentiment Classification Model Q needs to perform a sentiment classification task for the sentiment represented by each sentence in the training text information; at this time, the sentiment classification task performed by Sentiment Classification Model Q is to process text using "sentence" as the smallest analysis unit, while the text recognition task performed by Named Entity Recognition Model M is to process text using the "character-word" text recognition order. Based on this, it can be determined that the target task dependency relationship between Named Entity Recognition Model M and Sentiment Classification Model Q is: the sub-task to be performed by Sentiment Classification Model Q (i.e., the text analysis task at the sentence level) depends on the text recognition result of the character / word of Named Entity Recognition Model M (i.e., the text recognition task at the character / word level); from the task dependency relationship table, the information category corresponding to this target task dependency relationship is determined: character feature vector and word feature vector are the information categories of shared feature information to be extracted.

[0103] Here, based on the above analysis, in this embodiment of the application, as an optional embodiment, when the target task of the task processing model is related to recognizing the semantic sentiment expressed in text information, the shared feature information of the multiple categories may at least include:

[0104] 1. Character-level shared feature information, that is, the character feature vector in step S101 above.

[0105] Here, the character-level shared feature information is used to characterize the character arrangement feature information that can form the target word segment.

[0106] By way of example, taking an adjacent character bigram (two-dimensional grammar) vector, the training corpus can be split into a bigram character sequence. For example, the sentence "Beijing today has strong north wind and clear blue sky" in the training text information will be split into the sequence: "Bei Jing / Jing Jin / Jin Tian / Tian Bei / Bei Feng / Feng Jin / Jin Chui / Chui Lan / Lan Tian / Tian Ba / Ba Ping", then, the word2vec (a group of related models used to generate word vectors) method is used for training to obtain a 50-dimensional adjacent character bigram vector.

[0107] It should be noted that for character-level shared feature information, in addition to the adjacent character bigram vector in the above example, tri-gram (that is, splitting a sentence by grouping every 3 characters) feature vectors, 4-gram (that is, splitting a sentence by grouping every 4 characters) feature vectors, ..., n-Gram (that is, splitting a sentence by grouping every n characters) feature vectors can also be used. The embodiment of the present application does not make any limitation on the specific acquisition method of the above character-level shared feature information.

[0108] 2. Word segmentation-level shared feature information, that is, the word feature vector in the above step S101.

[0109] Here, the word segmentation-level shared feature information is part-of-speech feature information used to characterize the dependency relationship between different word segments in the same context.

[0110] Here, as an optional embodiment, StandFordNLP (a natural language processing tool) can be used to perform dependency syntactic analysis on the input training text information, to obtain the syntactic dependency relationship between words in the training text information, thereby obtaining the above word segmentation-level shared feature information.

[0111] Here, as another alternative embodiment, an adjacency matrix of a directed graph can also be used to store and acquire the syntactic dependency relationship between words.

[0112] By way of example, first, ignore the specific types of dependency relationships between words (for example, whether it is a subject-predicate relationship or a verb-object relationship, it is considered that there is a "dependency relationship"), then, according to the preset direction of the dependency relationship (for example, i pointing to j in the matrix can indicate that i is a dependent word of j), take the word segmentation results of each sentence as the rows and columns of the matrix respectively to create an adjacency matrix; at this time, if there is the above-mentioned "dependency relationship" between word i and word j, the value of the corresponding adjacency matrix element aij is 1; if there is no the above-mentioned "dependency relationship" between word i and word j, the value of the corresponding adjacency matrix element aij is 0; finally, taking each row in the adjacency matrix as the word feature vector corresponding to each column of words, the above word segmentation-level shared feature information can be obtained.

[0113] 3. High-dimensional shared feature information at the sentence level, namely, the sentence feature vector in step S101 above.

[0114] Here, the statement-level high-dimensional shared feature information is used to characterize the high-dimensional feature vectors mapped by different statements in the text information under multiple dimensional features; the target word segmentation is used to characterize words with semantic meaning.

[0115] As an example, as an optional embodiment, a BERT pre-trained language model can be used to perform text embedding processing on the training text information to obtain the sentence vector of each sentence as the aforementioned sentence-level high-dimensional shared feature information.

[0116] Option (2): Training a shared feature extraction model. Using NAS (Neural Network Architecture Search), when the target task to be executed by the task processing model can be split, the model searches for the optimal combination of sub-task models required to complete the target task to determine and extract the information categories of shared feature information.

[0117] Reference Figure 3 As shown, Figure 3 This paper illustrates a flowchart of a first method for searching neural network structures provided in an embodiment of this application. The method includes steps S301-S303; specifically:

[0118] S301, based on the target task to be executed by the task processing model, using the multiple sub-task models included in the task processing model as the first search space, and using the ability to execute the target task as the first search strategy, perform neural network structure search on the sub-task model combination methods among different sub-task models in the first search space to obtain the optimal sub-task model combination method that conforms to the first search strategy.

[0119] Here, the target task in the first search strategy can be used to characterize either the highest-level learning task that the task processing model can perform, or the secondary learning task located below the highest-level learning task.

[0120] For example, if the highest-level learning task that the task processing model can perform is the above-mentioned "semantic sentiment analysis" task (named entity recognition task + sentence sentiment classification task), then in step S301, the target task in the first search strategy can be the "semantic sentiment analysis" task, or it can be a secondary learning task under the above-mentioned highest-level learning task: the named entity recognition task or the sentence sentiment classification task.

[0121] Here, the definition of "optimal" in the optimal subtask model combination method can be determined based on the user's actual model training needs. For example, when it is detected that the user's actual model training needs tend to improve the model training speed, it can be determined that the optimal subtask model combination method is the model combination method that meets the above first search strategy and has the fastest overall training speed of multiple subtask models. When it is detected that the user's actual model training needs tend to improve the model training accuracy, it can be determined that the optimal subtask model combination method is the model combination method that meets the above first search strategy and has the most accurate overall training result of multiple subtask models.

[0122] S302, each subtask model included in the optimal subtask model combination method is taken as the first subtask model.

[0123] S303, based on the first task dependency relationship between each of the subtasks to be executed in the first subtask model, determine multiple information categories corresponding to the first task dependency relationship from the preset task dependency relationship table as multiple categories of the shared feature information to be extracted.

[0124] Here, after determining the optimal subtask model combination method, the specific implementation process of steps S302-S303 is the same as the optional method (1) above, and the repeated parts will not be repeated here.

[0125] Here, regarding the implementation of steps S301-S303 above, it should also be noted that, specifically in the specific model structure of each sub-task model, each sub-task model can be regarded as a combination of multiple neural networks of different types / levels. Therefore, for a sub-task model, the optimal model structure of each sub-task model can be determined by using the structural combination of different neural networks to complete the sub-task to be executed by the sub-task model (i.e., the model training task of the sub-task model) as the first search strategy mentioned above.

[0126] Option (3): Train the shared feature extraction model by using NAS (Neural Network Architecture Search) to determine and extract the information categories of shared feature information when the framework structure of the multi-task learning model is fixed (i.e., the target task to be executed by the task processing model cannot be split and executed).

[0127] Reference Figure 4 As shown, Figure 4 This paper illustrates a flowchart of a second neural network structure search method provided in an embodiment of this application. The method includes steps S401-S403; specifically:

[0128] S401, based on the target task to be executed by the task processing model, obtain various text feature information related to completing the target task.

[0129] For example, taking the target task to be executed by the task processing model as the aforementioned "semantic sentiment analysis" task (named entity recognition task + sentence sentiment classification task), then without splitting the target task, we can obtain "character-level shared feature information a, b, c" (equivalent to text feature information with the information category of character feature vector), "word-level shared feature information d, e, f" (equivalent to text feature information with the information category of word feature vector), and "sentence-level shared feature information g, h" (equivalent to text feature information with the information category of sentence feature vector) as various text feature information related to completing the target task.

[0130] S402, using the multiple text feature information as the second search space, and using the ability of the multiple sub-task models to complete the target task based on the information combination of different text feature information as the second search strategy, a neural network structure search is performed on the information combination methods between different text feature information in the second search space to obtain the optimal information combination method that conforms to the second search strategy.

[0131] Here, similar to step S301 above, the definition of "optimal" in the optimal information combination method can also be determined based on the user's actual model training needs. For example, when it is detected that the user's actual model training needs tend to improve the model training speed, it can be determined that the optimal information combination method is the information combination method that conforms to the above second search strategy and has the fewest information categories of shared feature information to be extracted; when it is detected that the user's actual model training needs tend to improve the model training accuracy, it can be determined that the optimal information combination method is the information combination method that conforms to the above second search strategy and has the most accurate overall training result of multiple sub-task models.

[0132] S403, the information category to which each text feature information included in the optimal information combination method belongs is taken as multiple categories of the shared feature information to be extracted.

[0133] For example, taking the example in S401 above, if the text feature information included in the optimal information combination method is determined to be: character-level shared feature information a and word-level shared feature information d, then the multiple categories of the shared feature information to be extracted can be determined to be: character feature vector and word-level feature vector.

[0134] Regarding the implementation process of step S102 above, when inputting the shared feature information part according to the first input method of hierarchical input, in addition to method 1 in step S102 above, refer to Figure 5As shown, Figure 5 This illustration shows a flowchart of a method for inputting shared feature information according to a first input method, provided in an embodiment of this application. The method includes steps S501-S503; specifically:

[0135] S501, for each of the sub-task models, at the first training node of the sub-task model, the first shared feature information of the word feature vector as the information category is input into the sub-task model.

[0136] Here, the first training node is used to represent the input node of the shallow neural network in the subtask model.

[0137] S502, at the second training node of the subtask model, the second shared feature information, which is the word feature vector, is input into the subtask model.

[0138] Here, the second training node is used to represent the input node of the intermediate neural network in the subtask model.

[0139] S503, at the third training node of the subtask model, the third shared feature information of the sentence feature vector as the information category is input into the subtask model.

[0140] Here, the third training node is used to represent the input node of the deep neural network in the subtask model.

[0141] Regarding the implementation of steps S501-S503 above, it should be noted that when the number of layers in the neural network in the subtask model is less than 3, taking a 2-layer neural network model structure as an example, the "character-level shared feature information" (i.e., the first shared feature information whose information category is the character feature vector) and the "word-level shared feature information" (i.e., the second shared feature information whose information category is the word feature vector) can be input into the first layer (i.e., the shallowest layer) of the neural network, and the "sentence-level shared feature information" (i.e., the third shared feature information whose information category is the sentence feature vector) can be input into the second layer (i.e., the deepest layer) of the neural network.

[0142] In addition, "character-level shared feature information" can be input into the first layer of the neural network, and "word-level shared feature information" and "statement-level shared feature information" can be input into the second layer of the neural network. The specific number of layers of the neural network in the subtask model is not limited in this application embodiment.

[0143] Regarding the implementation process of step S102 above, when inputting the shared feature information part according to the second input method of the first layer input, the specific implementation of method 2 in step S102 can be further divided into the following two optional implementation schemes according to whether to distinguish the task type differences of the sub-tasks to be executed by different sub-task models:

[0144] Option 1: When there is no distinction between the task types of the subtasks to be executed by different subtask models, the shared feature information of the multiple categories and the training text information are synchronously input into each of the subtask models using the second input method.

[0145] Optional Implementation Scheme 2, consisting of steps a and b, specifically:

[0146] Step a: When distinguishing the task types to which the subtasks to be executed in different subtask models belong, for each subtask model, based on the subtasks to be executed in that subtask model, determine the target shared feature information that matches the subtasks to be executed in the shared feature information of the multiple categories.

[0147] Step b: Using the second input method, simultaneously input the shared feature information of the multiple categories, the training text information, and the target shared feature information into the subtask model.

[0148] The model training method for the above-mentioned task processing model provided in the embodiments of this application,

[0149] First, training corpora are acquired and input into a shared feature extraction model. This model extracts shared feature information of multiple categories from the training corpora. Then, following a preset input method, the shared feature information of multiple categories and training text information labeled based on the training corpora are input into multiple sub-task models. These sub-task models are trained in parallel to ensure that the overall loss function of the multiple sub-task models meets the training cutoff condition. During the independent training of the multiple sub-task models, the task training loss of each sub-task model is acquired. Based on the gradient change of the task training loss of each sub-task model, the weight coefficients of that sub-task model are adjusted to ensure that the training rates of the multiple sub-task models are within the same numerical range, until the overall loss function of the multiple sub-task models meets the training cutoff condition. The trained multiple sub-task models are then used as the trained task processing model.

[0150] In this way, the multi-task learning model framework constructed in this application can ensure that each sub-task model can be trained independently, while providing different sub-task models with a variety of shared feature information related to the sub-tasks they perform, which is beneficial to improving the overall model training effect of the task processing model.

[0151] Based on the same inventive concept, this application also provides a model training device corresponding to the model training method of the task processing model in the above embodiments. Since the principle of the model training device in this application is similar to that of the model training method in the above embodiments of this application, the implementation of the model training device can refer to the implementation of the aforementioned model training method, and the repeated parts will not be described again.

[0152] Reference Figure 6 As shown, Figure 6 This illustration shows a schematic diagram of a model training device for a task processing model according to an embodiment of this application; wherein, the model training device is applied to a multi-task learning model framework, the multi-task learning model framework including a task processing model and a pre-trained shared feature extraction model, the task processing model including multiple sub-task models; the model training device includes:

[0153] The extraction module 601 is used to acquire training corpus and input the training corpus into the shared feature extraction model, and extract shared feature information of multiple categories from the training corpus through the shared feature extraction model;

[0154] The input module 602 is used to input the shared feature information of the multiple categories and the training text information labeled based on the training corpus into the multiple sub-task models according to a preset input method, and train the multiple sub-task models in parallel so that the overall loss function of the multiple sub-task models meets the training cutoff condition.

[0155] The training module 603 is used to obtain the task training loss of each of the multiple sub-task models during the independent training process of the multiple sub-task models, and adjust the weight coefficient of the sub-task model according to the gradient change of the task training loss of each sub-task model, so that the training rate of the multiple sub-task models is within the same numerical range, until the overall loss function of the multiple sub-task models meets the training cutoff condition, and then use the trained multiple sub-task models as the trained task processing model.

[0156] In one optional implementation, the shared feature information of the multiple categories includes: character feature vectors after the training corpus is segmented into character sequences; word feature vectors in the training corpus representing syntactic dependencies between words; and sentence feature vectors in the training corpus.

[0157] In an optional implementation, the extraction module 601 is used to determine multiple categories of shared feature information to be extracted by the shared feature extraction model through the following method:

[0158] Based on the target task dependencies between the multiple subtasks to be executed in the multiple subtask models, multiple information categories corresponding to the target task dependencies are determined from a preset task dependency table as multiple categories of shared feature information to be extracted; wherein, the task dependency table pre-stores multiple information categories corresponding to multiple task dependencies.

[0159] In an optional implementation, the extraction module 601 is used to determine multiple categories of shared feature information to be extracted by the shared feature extraction model through the following method:

[0160] Based on the target task to be executed by the task processing model, the multiple sub-task models included in the task processing model are used as the first search space, and the ability to execute the target task is used as the first search strategy. The neural network structure search is performed on the sub-task model combination methods among different sub-task models in the first search space to obtain the optimal sub-task model combination method that conforms to the first search strategy.

[0161] Each subtask model included in the optimal subtask model combination method is taken as the first subtask model;

[0162] Based on the first task dependency relationship between the subtasks to be executed in each first subtask model, multiple information categories corresponding to the first task dependency relationship are determined from the preset task dependency relationship table as multiple categories of the shared feature information to be extracted.

[0163] In an optional implementation, the extraction module 601 is used to determine multiple categories of shared feature information to be extracted by the shared feature extraction model through the following method:

[0164] Based on the target task to be executed by the task processing model, obtain various text feature information related to completing the target task;

[0165] Using the aforementioned multiple text feature information as the second search space, and taking the ability of the multiple sub-task models to complete the target task based on the information combination of different text feature information as the second search strategy, a neural network structure search is performed on the information combination methods between different text feature information in the second search space to obtain the optimal information combination method that conforms to the second search strategy.

[0166] The information category to which each text feature information included in the optimal information combination method belongs is taken as multiple categories of the shared feature information to be extracted.

[0167] In an optional implementation, when the shared feature information of the multiple categories and the training text information labeled based on the training corpus are input into the multiple sub-task models according to a preset input method, the input module 602 is specifically used for:

[0168] At the first-level model input node of each sub-task model, the training text information is input into each sub-task model;

[0169] The shared feature information of the various categories is input hierarchically to different training nodes in each sub-task model according to the correspondence between information categories and training nodes, using a hierarchical input first input method; wherein, the different training nodes in each sub-task model are sorted according to the layers of the neural network in the sub-task model from shallow to deep.

[0170] In an optional implementation, when the shared feature information of the multiple categories and the training text information labeled based on the training corpus are input into the multiple sub-task models according to a preset input method, the input module 602 is further configured to:

[0171] At the first-level model input node of each sub-task model, the shared feature information of the multiple categories and the training text information are synchronously input into each sub-task model in the second input method of the first-level input.

[0172] In an optional implementation, when the shared feature information of the multiple categories and the training text information are synchronously input into each of the sub-task models using the second input method with the first-layer input, the input module 602 is specifically used for:

[0173] When the task type of the subtask to be executed by different subtask models is not distinguished, the shared feature information of the multiple categories and the training text information are synchronously input into each subtask model using the second input method.

[0174] or,

[0175] When distinguishing the task types to which the subtasks to be executed in different subtask models belong, for each subtask model, based on the subtasks to be executed in that subtask model, the target shared feature information that matches the subtasks to be executed in the shared feature information of the multiple categories is determined.

[0176] Using the second input method, the shared feature information of the various categories, the training text information, and the target shared feature information are simultaneously input into the subtask model.

[0177] In an optional implementation, the overall loss function of the plurality of sub-task models is determined based on the product of the gradient of the task training loss of each sub-task model and the weight coefficient of that sub-task model in the multi-task learning model framework; when adjusting the weight coefficient of the sub-task model according to the gradient change of the task training loss of each sub-task model, the training module 603 is specifically used for:

[0178] For each subtask model, the gradient of the task training loss of that subtask model is used as the target gradient, and the periodic change amplitude of the target gradient within the gradient detection period is obtained.

[0179] When the periodic change of the target gradient is detected to be greater than or equal to the change of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted in a descending manner according to the gradient descent adjustment coefficient.

[0180] When the periodic change amplitude of the target gradient is detected to be less than the change amount of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted by increasing the gradient increase adjustment coefficient.

[0181] In one optional implementation, each subtask model in the task processing model is used to execute a corresponding subtask, and different subtask models cooperate with each other to process the target task to be executed by the task processing model; when the target task is related to the semantic sentiment expressed by the text information, the task processing model includes at least one named entity recognition model and one sentiment classification model; wherein, the named entity recognition model is used to perform a text recognition task for the named entities included in the training text information; the sentiment classification model is used to perform a sentiment classification task for the sentiment represented by each sentence in the training text information.

[0182] like Figure 7 As shown, this application provides a computer device 700 for executing the model training method of the task processing model in this application. The device includes a memory 701, a processor 702, and a computer program stored in the memory 701 and executable on the processor 702. When the processor 702 executes the computer program, it implements the steps of the model training method of the task processing model.

[0183] Specifically, the memory 701 and processor 702 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 702 runs the computer program stored in the memory 701, it can execute the model training method of the task processing model mentioned above.

[0184] Corresponding to the model training method of the task processing model in this application, this application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the steps of the above-described model training method of the task processing model.

[0185] Specifically, the storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the storage medium is run, it can execute the model training method of the above-mentioned task processing model.

[0186] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0187] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0188] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0189] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0190] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0191] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A model training method for a task processing model, characterized in that, An application is made in a multi-task learning model framework, which includes a task processing model and a pre-trained shared feature extraction model, wherein the task processing model includes multiple sub-task models; the model training method includes: The training corpus is input into the shared feature extraction model, which extracts multiple categories of shared feature information from the training corpus. These multiple categories of shared feature information include: character feature vectors after the training corpus is segmented into character sequences; word feature vectors representing syntactic dependencies between words in the training corpus; and sentence feature vectors in the training corpus. When the task types of the subtasks to be executed by different subtask models are not distinguished, at the first-layer model input node of each subtask model, the shared feature information of the multiple categories and the training text information labeled based on the training corpus are synchronously input into each subtask model in the second input method of the first-layer input, and the multiple subtask models are trained in parallel so that the overall loss function of the multiple subtask models meets the training cutoff condition. The model training method also includes: During the independent training of the multiple sub-task models, the task training loss of each sub-task model is obtained. The multi-task learning model framework adjusts the weight coefficients of the sub-task model in the multi-task learning model framework according to the gradient change of the task training loss of each sub-task model, so that the training rates of the multiple sub-task models are within the same numerical range, until the overall loss function of the multiple sub-task models meets the training cutoff condition, and the trained multiple sub-task models are used as the trained task processing model. The method for inputting the shared feature information of the multiple categories and the training text information into the multiple sub-task models further includes: At the first-level model input node of each sub-task model, the training text information is input into each sub-task model; The shared feature information of the various categories is input hierarchically to different training nodes in each sub-task model according to the correspondence between information categories and training nodes, using a hierarchical input first input method; wherein, the different training nodes in each sub-task model are sorted according to the layers of the neural network in the sub-task model from shallow to deep.

2. The model training method according to claim 1, characterized in that, The following method is used to determine the various categories of shared feature information to be extracted by the shared feature extraction model: Based on the target task dependencies between the multiple subtasks to be executed in the multiple subtask models, multiple information categories corresponding to the target task dependencies are determined from a preset task dependency table as multiple categories of shared feature information to be extracted; wherein, the task dependency table pre-stores multiple information categories corresponding to multiple task dependencies.

3. The model training method according to claim 1, characterized in that, The following method is used to determine the various categories of shared feature information to be extracted by the shared feature extraction model: Based on the target task to be executed by the task processing model, the multiple sub-task models included in the task processing model are used as the first search space, and the ability to execute the target task is used as the first search strategy. The neural network structure search is performed on the sub-task model combination methods among different sub-task models in the first search space to obtain the optimal sub-task model combination method that conforms to the first search strategy. Each subtask model included in the optimal subtask model combination method is taken as the first subtask model; Based on the first task dependency relationship between the subtasks to be executed in each first subtask model, multiple information categories corresponding to the first task dependency relationship are determined from the preset task dependency relationship table as multiple categories of the shared feature information to be extracted.

4. The model training method according to claim 1, characterized in that, The following method is used to determine the various categories of shared feature information to be extracted by the shared feature extraction model: Based on the target task to be executed by the task processing model, obtain various text feature information related to completing the target task; Using the aforementioned multiple text feature information as the second search space, and taking the ability of the multiple sub-task models to complete the target task based on the information combination of different text feature information as the second search strategy, a neural network structure search is performed on the information combination methods between different text feature information in the second search space to obtain the optimal information combination method that conforms to the second search strategy. The information category to which each text feature information included in the optimal information combination method belongs is taken as multiple categories of the shared feature information to be extracted.

5. The model training method according to claim 1, characterized in that, The method of inputting the shared feature information of the multiple categories and the training text information into the multiple sub-task models further includes: When distinguishing the task types to which the subtasks to be executed in different subtask models belong, for each subtask model, based on the subtasks to be executed in that subtask model, the target shared feature information that matches the subtasks to be executed in the shared feature information of the multiple categories is determined. Using the second input method, the shared feature information of the various categories, the training text information, and the target shared feature information are simultaneously input into the subtask model.

6. The model training method according to claim 1, characterized in that, The overall loss function of the multiple sub-task models is determined by multiplying the gradient of the task training loss of each sub-task model with the weight coefficient of that sub-task model in the multi-task learning model framework. The step of adjusting the weight coefficients of the sub-task model in the multi-task learning model framework based on the gradient change of the task training loss of each sub-task model includes: For each subtask model, the gradient of the task training loss of that subtask model is used as the target gradient, and the periodic change amplitude of the target gradient within the gradient detection period is obtained. When the periodic change of the target gradient is detected to be greater than or equal to the change of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted in a descending manner according to the gradient descent adjustment coefficient. When the periodic change amplitude of the target gradient is detected to be less than the change amount of the reference gradient, the weight coefficients of the subtask model are dynamically adjusted by increasing the gradient increase adjustment coefficient.

7. The model training method according to claim 1, characterized in that, Each subtask model in the task processing model is used to execute a corresponding subtask, and different subtask models cooperate with each other to process the target task to be executed by the task processing model; when the target task is related to the semantic sentiment expressed by the text information, the task processing model includes at least one named entity recognition model and one sentiment classification model; wherein, the named entity recognition model is used to perform a text recognition task for the named entities included in the training text information; the sentiment classification model is used to perform a sentiment classification task for the sentiment represented by each sentence in the training text information.

8. A model training device for a task processing model, characterized in that, An application is made in a multi-task learning model framework, the multi-task learning model framework including a task processing model and a pre-trained shared feature extraction model, the task processing model including multiple sub-task models; the model training device includes: An extraction module is used to input the training corpus into the shared feature extraction model, and extract multiple categories of shared feature information from the training corpus through the shared feature extraction model; wherein, the multiple categories of shared feature information include: character feature vectors after the training corpus is segmented into character sequences; word feature vectors representing syntactic dependencies between words in the training corpus; and sentence feature vectors in the training corpus; The input module is used to simultaneously input the shared feature information of the multiple categories and the training text information labeled based on the training corpus into each sub-task model at the first-layer model input node of each sub-task model when there is no distinction between the task types of the sub-tasks to be executed by different sub-task models. This allows for parallel training of the multiple sub-task models so that the overall loss function of the multiple sub-task models meets the training cutoff condition. The training module is used to obtain the task training loss of each sub-task model during the independent training of the multiple sub-task models, and adjust the weight coefficients of the sub-task model in the multi-task learning model framework according to the gradient change of the task training loss of each sub-task model, so that the training rates of the multiple sub-task models are within the same numerical range, until the overall loss function of the multiple sub-task models meets the training cutoff condition, and the trained multiple sub-task models are used as trained task processing models. The input module is further configured to: At the first-level model input node of each sub-task model, the training text information is input into each sub-task model; The shared feature information of the various categories is input hierarchically to different training nodes in each sub-task model according to the correspondence between information categories and training nodes, using a hierarchical input first input method; wherein, the different training nodes in each sub-task model are sorted according to the layers of the neural network in the sub-task model from shallow to deep.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the model training method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the model training method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text processing model training method and device and text processing method

    CN110209817A

  • Multi-task interaction enhanced electronic text event extraction method

    CN112069811A