Task fusion training method, device and equipment and storage medium thereof
By employing a unified encoding and joint loss calculation method, the problem of the separation between visual language models and visual action models is solved, achieving high availability and broad task processing capabilities for multi-task processing models, applicable to medical and financial task processing fields.
Patent Information
- Application Number
- CN202511065273.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-12-16
AI Technical Summary
In existing technologies, the development paths of visual language models and visual action models are fragmented, resulting in difficulties in knowledge sharing, high redundancy of model parameters, significantly increased training costs, and a lack of comprehensive processing models with high availability and broad task processing capabilities.
A unified encoding mode is used to encode the training datasets of multiple target tasks. The encoding results are input into the corresponding task training network through the same encoder, and a joint loss calculation method is used until the comprehensive loss value and the individual task loss value meet the threshold to complete the task fusion training.
It achieves high availability and broad task processing capabilities for multi-task processing models, enabling unified management and maintenance for different tasks and reducing training costs.
Smart Images

Figure CN121147653A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of research and development design technology, and is applied to multiple processing task fusion training scenarios. It relates to a task fusion training method, device, equipment and its storage medium. Background Technology
[0002] In current multimodal artificial intelligence research, visual language models and visual action models have each made significant progress, but their development paths are almost entirely separate. Visual language models focus primarily on tasks such as image-text alignment and image-text generation, capable of understanding image semantics and outputting language results that conform to the context, but lacking action modeling capabilities, they cannot further translate language understanding or visual perception into physical-level behavioral decisions. On the other hand, visual action models focus on predicting specific action trajectories from visual observation, and are often used in robot control and physical interaction. They typically employ supervised learning methods, require a large number of action labels, and have fixed input formats and limited modal types, lacking flexible language interaction capabilities.
[0003] Under the existing technological framework, these two types of models employ completely different pre-training strategies, network architectures, and optimization objectives, resulting in difficulties in knowledge sharing, high redundancy of model parameters, and a significant increase in training costs. This makes it extremely difficult to build a unified model structure for processing text-image and action-image tasks. Consequently, relevant business parties need to train task models separately for different sub-tasks when performing task processing, lacking a comprehensive processing model with high availability and broad task processing capabilities. For example, in the field of medical task processing, text-image generation tasks and action execution tasks require separate model training, and in the field of financial task processing, claim risk prediction and claim risk adjustment based on text-image data require separate model training. Often, when performing a joint task, multiple small task models need to be processed in conjunction, which undoubtedly increases the difficulty of task processing. Summary of the Invention
[0004] The purpose of this application is to propose a task fusion training method, apparatus, device and storage medium to overcome the technical problem of the lack of a comprehensive processing model with high availability and wide task processing scope.
[0005] Firstly, embodiments of this application provide a task fusion training method, which employs the following technical solution:
[0006] A task fusion training method includes the following steps:
[0007] Obtain training datasets for multiple target tasks;
[0008] The training datasets of the multiple target tasks are input into a preset encoder and uniformly encoded using the same encoding mode to obtain an encoding result set.
[0009] The encoding results of each training sample in the encoding result set are input into the corresponding task training network according to the different target task identifiers, and training is performed for multiple target tasks respectively.
[0010] Calculate the loss values corresponding to the multiple target tasks respectively, and use a joint loss calculation method to obtain the comprehensive loss value of the encoder;
[0011] Task fusion training is completed when the overall loss value meets the preset overall loss threshold and the loss values corresponding to the multiple target tasks meet their respective loss thresholds.
[0012] Secondly, embodiments of this application also provide a task fusion training device, which adopts the following technical solution:
[0013] A task fusion training device, comprising:
[0014] The training dataset acquisition module is used to acquire training datasets for multiple target tasks;
[0015] The training data encoding module is used to input the training datasets of the multiple target tasks into a preset encoder and perform unified encoding processing using the same encoding mode to obtain an encoding result set.
[0016] The task training execution module is used to input the encoding results of each training sample in the encoding result set into the corresponding task training network according to the different target task identifiers, and to train multiple target tasks respectively.
[0017] The loss value calculation module is used to calculate the loss value corresponding to the multiple target tasks respectively, and to obtain the comprehensive loss value of the encoder by using a joint loss calculation method;
[0018] The training completion determination module is used to determine that the task fusion training is completed when the comprehensive loss value meets the preset comprehensive loss threshold and the loss values corresponding to the multiple target tasks meet the corresponding loss thresholds.
[0019] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below:
[0020] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the task fusion training method described above.
[0021] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below:
[0022] A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the task fusion training method described above.
[0023] Compared with the prior art, the embodiments of this application have the following main advantages:
[0024] The task fusion training method described in this application involves acquiring training datasets for multiple target tasks; uniformly encoding these training datasets using the same encoding pattern; inputting the encoding results into the corresponding task training network for training on multiple target tasks; and completing task fusion training until the relevant loss values meet a preset loss threshold, thus obtaining a multi-task processing model. Applying this task fusion training method to the field of medical task processing, such as in tasks like medical image descriptive text matching or lesion development trend prediction based on medical images, it can fuse and train a large-scale integrated processing model suitable for different medical processing tasks using training datasets from different medical processing tasks compiled by hospitals. Similarly, applying this task fusion training method to the field of financial task processing, it can also fuse and train a large-scale integrated processing model suitable for different financial processing tasks using training datasets from different financial processing tasks compiled by financial institutions, such as training datasets for tasks like investment risk prediction and insurance risk prediction combining text and images. This avoids training different processing models for different processing tasks separately, facilitating unified management and maintenance of subsequent processing models. Attached Figure Description
[0025] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;
[0027] Figure 2 This is a flowchart of an embodiment of a task fusion training method according to this application;
[0028] Figure 3 This is a flowchart of a specific embodiment of the task fusion training method described in this application, which involves preprocessing the training dataset.
[0029] Figure 4 yes Figure 3 A flowchart of a specific embodiment of step 302 shown;
[0030] Figure 5 This is a flowchart of a specific embodiment of the task fusion training method described in this application, which involves extracting reference data;
[0031] Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 202 shown;
[0032] Figure 7 yes Figure 2 A flowchart of a specific embodiment of step 203 shown;
[0033] Figure 8 This is a schematic diagram of one embodiment of a task fusion training device according to this application;
[0034] Figure 9 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0036] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0038] like Figure 1As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0039] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0040] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptop computer 1011, tablet computer 1012 or mobile phone 1013, terminal device 101 can also be e-book reader, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer and desktop computer, etc.
[0041] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0042] It should be noted that the task fusion training method provided in this application embodiment is generally executed by a server, and correspondingly, a task fusion training device is generally set in the server.
[0043] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0044] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of a task fusion training method according to this application. The task fusion training method includes the following steps:
[0045] Step 201: Obtain training datasets for multiple target tasks.
[0046] In this embodiment, the multiple target tasks refer to multiple tasks to be trained, including matching tasks, generation tasks, and prediction tasks.
[0047] More specifically, the target tasks include image-text matching tasks, image-text generation tasks, action prediction tasks, and speech synthesis tasks.
[0048] By acquiring training datasets for multiple target tasks, it is possible to subsequently use these training datasets for target task processing training.
[0049] Step 202: Input the training datasets of the multiple target tasks into a preset encoder and perform unified encoding processing using the same encoding mode to obtain an encoding result set.
[0050] In this embodiment, the training datasets of the multiple target tasks are input into a preset encoder and uniformly encoded using the same encoding mode. The purpose is to uniformly acquire the encoding features of the multiple target tasks.
[0051] Specifically, the preset encoder includes an encoder capable of performing multimodal feature encoding. By specifying the preset encoder as an encoder capable of performing multimodal feature encoding, it is ensured that when the multiple target tasks involve multimodal data encoding, the preset encoder can implement the corresponding multimodal encoding function.
[0052] By inputting the training datasets of the multiple target tasks into a preset encoder and using the same encoding mode for unified encoding processing, the need to set up multiple or complex encoders for different target tasks is avoided, making the encoding relatively simple and saving feature encoding processing resources.
[0053] Step 203: Input the encoding results of each training sample in the encoding result set into the corresponding task training network according to the different target task identifiers, and train multiple target tasks respectively.
[0054] Specifically, although the same encoder can be used for different target tasks during the task feature encoding stage, there are obvious differences in task execution when training different target tasks. This is because multiple tasks may involve different processing branches such as matching, prediction, and decision-making. Therefore, different task training networks are pre-built for different target tasks to achieve refined learning and training for different processing tasks when actually training target tasks.
[0055] For example, when the target tasks include image-text matching, image-text generation, action prediction, and speech synthesis, a unified encoder is used to encode the task features of the training datasets for different target tasks, resulting in task feature encoding results. Subsequently, during the learning and training of image-text matching, image-text generation, action prediction, and speech synthesis tasks, the corresponding task feature encoding results are input into the task training networks corresponding to each task for learning and training, ultimately resulting in the trained task processing networks corresponding to each task for image-text matching, image-text generation, action prediction, and speech synthesis.
[0056] In this embodiment, taking into account the actual task processing situation, a unified encoding processing mode is adopted in the feature encoding stage where features can be processed jointly. Then, in the actual task processing training stage, a mode of separate learning and training for different target tasks is adopted, which achieves high availability of the task processing network obtained by learning and training while reducing the complexity of feature encoding.
[0057] Step 204: Calculate the loss values corresponding to the multiple target tasks respectively, and use a joint loss calculation method to obtain the comprehensive loss value of the encoder.
[0058] Specifically, based on the actual input-output mapping results of the training dataset and the learning-training input-output mapping results of all target tasks after fusion learning training, the loss values corresponding to each target task are calculated.
[0059] By combining the loss values corresponding to all target tasks, a joint calculation of the loss values is performed to obtain the joint calculation result, which is used as the comprehensive loss value of the encoder.
[0060] By obtaining the comprehensive loss value of the encoder, the entire task processing model obtained after learning and training, namely the feature encoding processing component and the specific task decision processing component, can be applied to the task processing of all target tasks, thereby improving the usability and breadth of use of the final task processing model.
[0061] Step 205: Task fusion training is completed when the overall loss value meets the preset overall loss threshold and the loss values corresponding to the multiple target tasks meet the corresponding loss thresholds.
[0062] Specifically, the encoder is considered suitable for encoding task features for all target tasks until the overall loss value meets a preset overall loss threshold; and the loss values corresponding to the multiple target tasks meet their respective loss thresholds, meaning that the task processing decision networks for all target tasks can support decision processing for their respective target tasks. This ensures that the finally trained task fusion processing model has high availability and broad task processing capabilities, meaning it can be applied to decision processing for multiple processing tasks.
[0063] Applying the task fusion training method described in this application to the field of medical task processing, such as in tasks like matching medical descriptions to text based on medical images or predicting lesion development trends based on medical images, allows for the fusion training of training datasets from different medical processing tasks compiled by hospitals. This enables the creation of a single integrated large-scale processing model suitable for various tasks, avoiding the need to train separate processing models for each medical task and facilitating unified management and maintenance of subsequent processing models. Similarly, applying the task fusion training method to the field of financial task processing allows for the fusion training of training datasets from different financial processing tasks compiled by financial institutions, such as training datasets for tasks like predicting financial risks and insurance risks by combining text and images. This also avoids the need to train separate processing models for different financial tasks and facilitates unified management and maintenance of subsequent processing models.
[0064] In this embodiment, training datasets for multiple target tasks are acquired; the training datasets are uniformly encoded using the same encoding mode; the encoding results are input into the corresponding task training networks for training multiple target tasks; until the relevant loss values meet the preset loss threshold, the task fusion training is completed, resulting in a multi-task processing model. Applying this task fusion training method to the field of medical task processing, for example, in tasks such as medical image descriptive text matching or lesion development trend prediction based on medical images, a comprehensive large-scale processing model suitable for different medical processing tasks can be trained by fusing training datasets from different medical processing tasks compiled by hospitals. Similarly, applying this task fusion training method to the field of financial task processing, a comprehensive large-scale processing model suitable for different financial processing tasks can be trained by fusing training datasets from different financial processing tasks compiled by financial institutions, such as training datasets corresponding to tasks combining text and images for investment risk prediction and insurance risk prediction, avoiding the need to train different processing models for different processing tasks and facilitating unified management and maintenance of subsequent processing models.
[0065] Continue to refer to Figure 3In some optional implementations, after step 201, the method further includes a step of preprocessing the training dataset. Figure 3 This is a flowchart of a specific embodiment of the task fusion training method described in this application, which involves preprocessing the training dataset, and includes the following steps:
[0066] Step 301: According to the different target tasks, perform differential labeling processing on the training datasets of the multiple target tasks;
[0067] Specifically, different task identifiers are set for the training datasets corresponding to different target tasks, depending on the target task.
[0068] Step 302: According to the difference labeling processing results and the preset unified data structure encapsulation method, the training datasets of the multiple target tasks are encapsulated into task data structures to obtain task data encapsulation results.
[0069] Specifically, the training datasets of the multiple target tasks are encapsulated using a pre-defined unified data structure method to facilitate subsequent parsing based on the unified encapsulation method to uniformly obtain the target data.
[0070] In this embodiment, the multiple target tasks include an image-text matching task, an image-text generation task, and an action prediction task. Specifically, the image-text matching task is used to train the mapping correspondence between descriptive text and images. The training dataset for the image-text matching task includes an image set and the descriptive text applied to the image-text matching for each image. The image-text generation task is used to generate target images based on the descriptive text. The training dataset for the image-text generation task includes the original image, the image-generated descriptive text, and the generated image. The action prediction task is used to predict the next action of the target operator based on the action instruction text and the action in the original image. The training dataset for the action prediction task includes the original image, the action instruction text, and the subsequently generated image containing the next action.
[0071] Continue to refer to Figure 4 , Figure 4 yes Figure 3 A flowchart of a specific embodiment of step 302 shown includes the following steps:
[0072] Step 401: Based on the difference labeling processing results, select the training datasets corresponding to the image-text matching task, image-text generation task, and action prediction task respectively.
[0073] Specifically, based on the task identifiers corresponding to different training datasets, the training datasets corresponding to the image-text matching task, image-text generation task, and action prediction task are distinguished.
[0074] Step 402: For each sample in the training dataset corresponding to the image-text matching task, encapsulate it according to the task data structure where the input is an image and the descriptive text pair, and the output is the descriptive text.
[0075] Specifically, assuming that the image and descriptive text pair contained in the current sample in the image-text matching task are "orange" image and "orange" text respectively, then the sample is encapsulated as training sample data according to the encapsulation structure of input data pair consisting of "orange" image and "orange" text and output item consisting of "orange" text.
[0076] Step 403: For each sample in the training dataset corresponding to the image and text generation task, encapsulate the task data structure according to the input items being the existing image and the target image description text pair, and the output item being the generated image.
[0077] Similarly, for each sample in the training dataset corresponding to the image-text generation task, the training sample data is encapsulated according to the input data pair consisting of the prior image and the target generated image description text, and the output is the encapsulation structure of the generated image.
[0078] Step 404: For each sample in the training dataset corresponding to the action prediction task, encapsulate the task data structure with the input being the prior action image and the action instruction text pair, and the output being the predicted action.
[0079] Similarly, for each sample in the training dataset corresponding to the action prediction task, the training sample data is encapsulated according to the input data pair consisting of the prior image and the action instruction text, and the output is the encapsulated structure of the predicted action.
[0080] Step 405: Obtain the task data structure encapsulation result corresponding to each sample in all training datasets, and assign a corresponding task identifier to each sample according to the different target tasks.
[0081] Specifically, when performing structured encapsulation for each sample, for the task data structure encapsulation result corresponding to each sample, a task identifier is set according to the task identifier of the training dataset in which the sample is located. This facilitates subsequent processing steps by acquiring data according to a unified data encapsulation structure and also makes it easier to count the number of training samples included in different target tasks.
[0082] Continue to refer to Figure 5 In some optional implementations, after step 302, the method further includes a step of extracting reference data. Figure 5This is a flowchart of a specific embodiment of the task fusion training method described in this application, which involves extracting reference data, and includes the following steps:
[0083] Step 501: Encapsulate the task data structure of all samples corresponding to the image-text matching task, extract the output items, and generate a reference text set composed of the output items;
[0084] Step 502: Encapsulate the task data structure of all samples corresponding to the image and text generation task, extract the output items, and generate a reference generated image set composed of the output items;
[0085] Step 503: Encapsulate the task data structure of all samples corresponding to the action prediction task, extract the output items, and generate a reference predicted action set composed of the output items.
[0086] Specifically, since all sample data were encapsulated in a unified manner, we can extract only the output items corresponding to each sample data from the encapsulation results. Then, based on the task identifiers corresponding to each sample data, we can summarize the extracted output items and finally obtain the reference text set, reference generated image set, and reference predicted action set.
[0087] By obtaining the reference text set, reference generated image set, and reference predicted action set, the training loss can be calculated by combining the reference text set, reference generated image set, and reference predicted action set at the end of subsequent fusion learning training.
[0088] In this embodiment, the preset encoder includes a vision-language encoder. The vision encoding part of the vision-language encoder includes an image Vision Transformer (ViT) encoding part. The Vision Transformer is a deep learning model that introduces the Transformer architecture into the field of computer vision. It processes image data through a self-attention mechanism and can efficiently capture global information and long-distance dependencies. The language encoding part of the vision-language encoder includes word embedding encoding or a bag-of-words model-based encoding method.
[0089] Continue to refer to Figure 6 , Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 202 shown includes the following steps:
[0090] Step 601: Input the task data structure encapsulation result corresponding to each sample in all training datasets into the vision-language encoder;
[0091] Step 602: Using the visual encoding part in the visual-language encoder, the images in the encapsulation results of all task data structures are divided into small blocks according to a preset segmentation scale and converted into vector representations, and mapped to preset fixed-dimensional visual features.
[0092] Step 603: Using the language encoding part of the vision-language encoder, map the text content in the encapsulation result of all task data structures to a text embedding vector representation, and insert the corresponding task identifier;
[0093] Step 604: Using a multimodal fusion method, visual features and text embedding vector representations under the same task identifier are concatenated into a unified feature vector representation sequence, and then fused through a preset multi-layer Transformer Encoder structure.
[0094] Step 605: Output the multimodal representation vectors containing semantic alignment information generated after the fusion processing of all samples, and generate the encoding result set.
[0095] By employing a pre-defined visual-language encoder, image and text features are extracted from the image and text data in the training dataset and then fused. This yields multimodal representation vectors containing semantic alignment information for each sample after fusion processing. These multimodal representation vectors can then be directly used to train task-oriented networks for different target tasks. Compared to using image or text features alone, the multimodal representation vectors contain richer features, ensuring the high availability of the trained task-oriented networks.
[0096] In this embodiment, before performing the step of inputting the encoding results of each training sample in the encoding result set into the corresponding task training network according to the different target task identifiers, and training multiple target tasks respectively, the method further includes: pre-constructing task training networks corresponding to different tasks according to the different target tasks. Specifically, pre-constructing task training networks corresponding to image-text matching task, image-text generation task, and action prediction task respectively, and setting different task identifiers for different task training networks.
[0097] Specifically, different task training network entry points can be set for the pre-built task training networks corresponding to image-text matching, image-text generation, and action prediction tasks. Then, when the training data encoding results corresponding to different tasks, namely image-text matching, image-text generation, and action prediction tasks, are obtained, the corresponding training data encoding results are sent to the corresponding task training network entry points for separate training processing through the preset task scheduler and in combination with the different task training network entry points.
[0098] Continue to refer to Figure 7 , Figure 7 yes Figure 2 A flowchart of a specific embodiment of step 203 shown includes the following steps:
[0099] Step 701: Identify the encoding results of samples contained in different target tasks based on the different task identifiers;
[0100] Step 702: Input the encoding results of different samples into the corresponding task training network according to the recognition results;
[0101] Step 703: The task training network for the image-text matching task is used to perform text decoding processing on the input encoding result, and the generated decoded text set is output based on the text decoding processing result.
[0102] Step 704: The task training network for the image and text generation task is used to perform image decoding processing on the input encoding result, and the generated decoded image set is output based on the image decoding processing result.
[0103] Step 705: Use the task training network of the action prediction task to perform action trajectory prediction processing on the input encoding results, and output the generated action trajectory set based on the action trajectory prediction results.
[0104] Specifically, based on different task identifiers, the encoding results of samples contained in different target tasks are identified, and the encoding results of different samples are input into the corresponding task training networks according to the identification results. Then, the training processing results output by the different task training networks are obtained, namely the decoded text set, the decoded image set, and the action trajectory set. The decoding processing here can be performed on the multimodal representation vector containing semantic alignment information generated after fusion processing. For example, when obtaining the decoded text set, an autoregressive language decoder is used; when obtaining the decoded image set, an image decoder is used; and when obtaining the action trajectory set, an autoregressive action head method is used for prediction.
[0105] By acquiring the training output sets of different target tasks during fusion training, namely the decoded text set, decoded image set, and action trajectory set, it is possible to compare and identify them with the previously prepared reference text set, reference generated image set, and reference predicted action set, thereby calculating the loss value of the network trained for different tasks.
[0106] In this embodiment, the step of calculating the loss values corresponding to the multiple target tasks specifically includes: using a text recognition comparison method to identify the number of elements inconsistent between the decoded text set and the reference text set, calculating the ratio of the number of inconsistent elements to the number of elements in the reference text set, and setting the ratio as the loss value corresponding to the image-text matching task; using an image difference recognition method to identify the differences between the decoded image set and the reference generated image set, and setting the differences as the loss value corresponding to the image-text generation task; using a comparison recognition method to identify the number of differing actions between the action trajectory set and the reference predicted action set, calculating the ratio of the number of differing actions to the number of elements in the reference predicted action set, and setting the ratio as the loss value corresponding to the action prediction task.
[0107] In this embodiment, the step of obtaining the comprehensive loss value of the encoder by using the joint loss calculation method specifically includes: calculating the proportion of samples included in different target tasks in the total training samples; setting the proportion as the loss weight corresponding to different target tasks; and performing multiplication and summation processing based on the loss weight and loss value corresponding to different target tasks to calculate the comprehensive loss value of the encoder.
[0108] Specifically, assuming the training samples for the image-text matching task, image-text generation task, and action prediction task are 10,000, 30,000, and 60,000 respectively, then the proportions of the training samples for the image-text matching task, image-text generation task, and action prediction task are 0.1, 0.3, and 0.6 respectively. Assuming the loss values for the image-text matching task, image-text generation task, and action prediction task have been calculated in the previous steps as 0.03, 0.05, and 0.1 respectively, then the calculation method for the comprehensive loss value of the encoder is: 0.1×0.03+0.3×0.05+0.6×0.1=0.078.
[0109] By calculating the loss values corresponding to the multiple target tasks and the comprehensive loss value of the encoder, it is ensured that the loss values of the encoder and the task training network obtained by fusion training are both less than the corresponding loss thresholds, thereby ensuring the high availability and task processing breadth of the image and text task processing model with multi-task processing function obtained after fusion training.
[0110] In this embodiment, after performing the steps of calculating the loss values corresponding to the multiple target tasks respectively and using a joint loss calculation method to obtain the comprehensive loss value of the encoder, the method further includes: if the comprehensive loss value does not meet a preset comprehensive loss threshold, or if the loss value corresponding to at least one of the multiple target tasks does not meet the corresponding loss threshold, then the encoding parameters of the encoder and the network parameters of the task training networks corresponding to the multiple target tasks are adjusted, and task fusion training is performed again; until the comprehensive loss value meets the preset comprehensive loss threshold, and the loss values corresponding to the multiple target tasks meet the corresponding loss thresholds, the parameter adjustment is stopped, and a multi-task processing image and text task processing model is obtained.
[0111] In this embodiment, after performing the step of completing the task fusion training until the comprehensive loss value meets the preset comprehensive loss threshold and the loss values corresponding to the multiple target tasks meet the corresponding loss thresholds, the method further includes: recording the encoding parameters in the encoder, and recording the network parameters of the task training networks corresponding to the multiple target tasks, and deploying a multi-task processing model for text and image tasks based on the encoding parameters and the network parameters.
[0112] Specifically, by recording the encoding parameters in the encoder and the network parameters of the task training networks corresponding to the multiple target tasks, the image and text task processing model with integrated training functions can be flexibly deployed.
[0113] It should be understood that the multi-task processing image and text task processing model ultimately trained by the task fusion training method described in this application includes the trained encoder and the trained task processing networks corresponding to different target tasks.
[0114] In this embodiment, training datasets for multiple target tasks are acquired; the training datasets are uniformly encoded using the same encoding mode; the encoding results are input into the corresponding task training networks for training multiple target tasks; until the relevant loss values meet the preset loss threshold, the task fusion training is completed, resulting in a multi-task processing model. Applying this task fusion training method to the field of medical task processing, for example, in tasks such as medical image descriptive text matching or lesion development trend prediction based on medical images, a comprehensive large-scale processing model suitable for different medical processing tasks can be trained by fusing training datasets from different medical processing tasks compiled by hospitals. Similarly, applying this task fusion training method to the field of financial task processing, a comprehensive large-scale processing model suitable for different financial processing tasks can be trained by fusing training datasets from different financial processing tasks compiled by financial institutions, such as training datasets corresponding to tasks combining text and images for investment risk prediction and insurance risk prediction, avoiding the need to train different processing models for different processing tasks and facilitating unified management and maintenance of subsequent processing models.
[0115] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0116] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0117] In this embodiment, training datasets for multiple target tasks are acquired; the training datasets are uniformly encoded using the same encoding mode; the encoding results are input into the corresponding task training networks for training multiple target tasks; until the relevant loss values meet the preset loss threshold, the task fusion training is completed, resulting in a multi-task processing model. Applying this task fusion training method to the field of medical task processing, for example, in tasks such as medical image descriptive text matching or lesion development trend prediction based on medical images, a comprehensive large-scale processing model suitable for different medical processing tasks can be trained by fusing training datasets from different medical processing tasks compiled by hospitals. Similarly, applying this task fusion training method to the field of financial task processing, a comprehensive large-scale processing model suitable for different financial processing tasks can be trained by fusing training datasets from different financial processing tasks compiled by financial institutions, such as training datasets corresponding to tasks combining text and images for investment risk prediction and insurance risk prediction, avoiding the need to train different processing models for different processing tasks and facilitating unified management and maintenance of subsequent processing models.
[0118] Further reference Figure 8 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a task fusion training device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0119] like Figure 8 As shown, the task fusion training device 800 described in this embodiment includes: a training dataset acquisition module 801, a training data encoding module 802, a task training execution module 803, a loss value calculation module 804, and a training completion determination module 805. Wherein:
[0120] The training dataset acquisition module 801 is used to acquire training datasets for multiple target tasks;
[0121] The training data encoding module 802 is used to input the training datasets of the multiple target tasks into a preset encoder and perform unified encoding processing using the same encoding mode to obtain an encoding result set.
[0122] The task training execution module 803 is used to input the encoding results of each training sample in the encoding result set into the corresponding task training network according to the different target task identifiers, and to train multiple target tasks respectively.
[0123] The loss value calculation module 804 is used to calculate the loss values corresponding to the multiple target tasks respectively, and to obtain the comprehensive loss value of the encoder by using a joint loss calculation method;
[0124] The training completion determination module 805 is used to determine that the task fusion training is completed when the comprehensive loss value meets the preset comprehensive loss threshold and the loss values corresponding to the multiple target tasks meet the corresponding loss thresholds.
[0125] This application obtains training datasets for multiple target tasks; performs unified encoding processing on the training datasets using the same encoding mode; inputs the encoding results into the corresponding task training networks for training on multiple target tasks; and completes task fusion training until the relevant loss values meet a preset loss threshold, thus obtaining a multi-task processing model. Applying this task fusion training method to the field of medical task processing, such as in tasks like medical image descriptive text matching or lesion development trend prediction based on medical images, it can fuse and train a large-scale integrated processing model suitable for different medical processing tasks using training datasets compiled by hospitals for different medical processing tasks. Similarly, applying this task fusion training method to the field of financial task processing, it can also fuse and train a large-scale integrated processing model suitable for different financial processing tasks using training datasets compiled by financial institutions for different financial processing tasks, such as training datasets for tasks like investment risk prediction and insurance risk prediction combining text and images, avoiding the need to train different processing models for different processing tasks and facilitating unified management and maintenance of subsequent processing models.
[0126] In this embodiment, the task fusion training device 800 further includes a differentiation labeling processing module and a sample data unified encapsulation processing module. Wherein:
[0127] The differential labeling processing module is used to perform differential labeling processing on the training datasets of the multiple target tasks according to the different target tasks;
[0128] The sample data unified encapsulation processing module is used to encapsulate the training datasets of the multiple target tasks according to the distinguishing label processing results and the preset data structure unified encapsulation method, so as to obtain the task data encapsulation results.
[0129] In this embodiment, the unified sample data encapsulation processing module includes: a data filtering unit, a first encapsulation execution unit, a second encapsulation execution unit, a third encapsulation execution unit, and an encapsulation result acquisition and marking unit. Wherein:
[0130] The data filtering unit is used to filter out the training datasets corresponding to the image-text matching task, the image-text generation task, and the action prediction task based on the difference label processing results.
[0131] The first encapsulation execution unit is used to encapsulate each sample in the training dataset corresponding to the image-text matching task according to the task data structure where the input item is an image and the descriptive text pair, and the output item is the descriptive text.
[0132] The second encapsulation execution unit is used to encapsulate each sample in the training dataset corresponding to the image and text generation task according to the task data structure with the input item being a pair of description texts of the prior image and the target image, and the output item being the generated image.
[0133] The third encapsulation execution unit is used to encapsulate each sample in the training dataset corresponding to the action prediction task according to the task data structure with the input items being the prior action image and action instruction text pair, and the output item being the predicted action.
[0134] The encapsulation result acquisition and labeling unit is used to obtain the task data structure encapsulation result corresponding to each sample in all training datasets, and assign a corresponding task identifier to each sample according to the different target tasks.
[0135] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0136] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0137] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 9 , Figure 9This is a basic structural block diagram of the computer device in this embodiment.
[0138] The computer device 9 includes a memory 9a, a processor 9b, and a network interface 9c that are interconnected via a system bus. It should be noted that... Figure 9 Only a computer device 9 with component memory 9a, processor 9b, and network interface 9c is shown. However, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0139] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0140] The memory 9a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 9a may be an internal storage unit of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 9a may also be an external storage device of the computer device 9, such as a plug-in hard disk, smart memory card (SMC), secure digital (SD) card, flash memory card, etc. Of course, the memory 9a may include both internal storage units and external storage devices of the computer device 9. In this embodiment, the memory 9a is typically used to store the operating system and various application software installed on the computer device 9, such as computer-readable instructions for a task fusion training method. In addition, the memory 9a can also be used to temporarily store various types of data that have been output or will be output.
[0141] In some embodiments, the processor 9b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 9b is typically used to control the overall operation of the computer device 9. In this embodiment, the processor 9b is used to execute computer-readable instructions stored in the memory 9a or to process data, for example, to execute computer-readable instructions for the task fusion training method described above.
[0142] The network interface 9c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 9 and other electronic devices.
[0143] The computer device proposed in this embodiment belongs to the field of research and development design technology and is applied to a multi-task fusion training scenario. This application obtains training datasets for multiple target tasks; performs unified encoding processing on the training datasets using the same encoding mode; inputs the encoding results into the corresponding task training network for training multiple target tasks; until the relevant loss values meet a preset loss threshold, the task fusion training is completed, resulting in a multi-task processing model. Applying this task fusion training method to the field of medical task processing, for example, in tasks such as medical image descriptive text matching or lesion development trend prediction based on medical images, it can fuse and train an integrated large-scale processing model suitable for different medical processing tasks using training datasets of different medical processing tasks compiled by hospitals. Similarly, applying this task fusion training method to the field of financial task processing, it can also fuse and train an integrated large-scale processing model suitable for different financial processing tasks using training datasets of different financial processing tasks compiled by financial institutions, such as training datasets corresponding to tasks combining text and images for investment risk prediction and insurance risk prediction, avoiding the need to train different processing models for different processing tasks and facilitating unified management and maintenance of subsequent processing models.
[0144] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the task fusion training method described above.
[0145] The computer-readable storage medium proposed in this embodiment belongs to the field of research and development design technology and is applied in a multi-task fusion training scenario. This application obtains training datasets for multiple target tasks; performs unified encoding processing on the training datasets using the same encoding mode; inputs the encoding results into the corresponding task training network to train multiple target tasks; and completes task fusion training until the relevant loss values meet the preset loss threshold, thereby obtaining a multi-task processing model. Applying the aforementioned task fusion training method to the field of medical task processing, such as in tasks like matching medical descriptions to text based on medical images or predicting lesion development trends based on medical images, allows for the fusion training of training datasets from different medical processing tasks compiled by hospitals to create a single integrated large-scale processing model suitable for various medical processing tasks. Similarly, applying the same task fusion training method to the field of financial task processing also allows for the fusion training of training datasets from different financial processing tasks compiled by financial institutions, such as training datasets for tasks like predicting financial risks and insurance risks by combining text and images, to create a single integrated large-scale processing model suitable for various financial processing tasks. This avoids the need to train different processing models for different tasks and facilitates unified management and maintenance of subsequent processing models.
[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0147] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
Claims
1. A task fusion training method, characterized in that, Includes the following steps: Obtain training datasets for multiple target tasks; The training datasets of the multiple target tasks are input into a preset encoder and uniformly encoded using the same encoding mode to obtain an encoding result set. The encoding results of each training sample in the encoding result set are input into the corresponding task training network according to the different target task identifiers, and training is performed for multiple target tasks respectively. Calculate the loss values corresponding to the multiple target tasks respectively, and use a joint loss calculation method to obtain the comprehensive loss value of the encoder; Task fusion training is completed when the overall loss value meets the preset overall loss threshold and the loss values corresponding to the multiple target tasks meet their respective loss thresholds.
2. The task fusion training method according to claim 1, characterized in that, After performing the step of obtaining the training dataset for multiple target tasks, the method further includes: Based on the different target tasks, the training datasets for the multiple target tasks are distinguished and labeled accordingly; Based on the distinguishing labeling processing results and the preset unified data structure encapsulation method, the training datasets of the multiple target tasks are encapsulated into task data structures to obtain task data encapsulation results.
3. The task fusion training method according to claim 2, characterized in that, The multiple target tasks include image-text matching, image-text generation, and action prediction. The step of encapsulating the training datasets for these multiple target tasks into task data structures according to the distinguishing label processing results and a preset unified data structure encapsulation method, to obtain the task data encapsulation results, specifically includes: Based on the results of the distinguishing labeling process, training datasets corresponding to the image-text matching task, image-text generation task, and action prediction task are selected respectively. For each sample in the training dataset corresponding to the image-text matching task, the task data structure is encapsulated according to the input item being an image and a descriptive text pair, and the output item being the descriptive text; For each sample in the training dataset corresponding to the image-text generation task, the task data structure is encapsulated with the input being a pair of description texts for the prior image and the target image, and the output being the generated image. For each sample in the training dataset corresponding to the action prediction task, the task data structure is encapsulated with the input items being the prior action image and the action instruction text pair, and the output item being the predicted action. Obtain the task data structure encapsulation results corresponding to each sample in all training datasets, and assign a corresponding task identifier to each sample according to the different target tasks.
4. The task fusion training method according to any one of claims 2 or 3, characterized in that, After performing the step of encapsulating the training datasets of the multiple target tasks into task data structures according to the distinguishing label processing results and the preset data structure unified encapsulation method, and obtaining the task data encapsulation results, the method further includes: The task data structure of all samples corresponding to the image-text matching task is encapsulated, the output items are extracted, and a reference text set composed of the output items is generated. The task data structure encapsulation results of all samples corresponding to the image and text generation task are encapsulated, the output items are extracted, and a reference generated image set composed of the output items is generated. The task data structure of all samples corresponding to the action prediction task is encapsulated, output items are extracted, and a reference predicted action set composed of output items is generated.
5. The task fusion training method according to claim 3, characterized in that, The preset encoder includes a vision-language encoder. The step of inputting the training datasets of the multiple target tasks into the preset encoder and performing unified encoding processing using the same encoding mode to obtain an encoding result set specifically includes: The task data structure encapsulation result corresponding to each sample in all the training datasets is input into the vision-language encoder; Using the visual encoding part of the visual-language encoder, the images in the encapsulation results of all task data structures are segmented into small blocks according to a preset segmentation scale and converted into vector representations, which are then mapped to preset fixed-dimensional visual features. Using the language encoding part of the vision-language encoder, the text content in the encapsulation result of all task data structures is mapped to a text embedding vector representation, and the corresponding task identifier is inserted. A multimodal fusion approach is adopted to concatenate visual features and text embedding vector representations under the same task identifier into a unified feature vector representation sequence, which is then fused through a pre-defined multi-layer Transformer Encoder structure. Output the multimodal representation vectors containing semantic alignment information generated after the fusion processing of all samples, and generate the encoding result set.
6. The task fusion training method according to claim 4, characterized in that, Before performing the step of inputting the encoding results of each training sample in the encoding result set into the corresponding task training network according to the different target task identifiers, and training multiple target tasks respectively, the method further includes: Based on the different target tasks, task training networks corresponding to different tasks are pre-constructed. Specifically, task training networks corresponding to image-text matching tasks, image-text generation tasks, and action prediction tasks are pre-constructed, and different task labels are set for different task training networks. The step of inputting the encoding results of each training sample in the encoding result set into the corresponding task training network according to the different target task identifiers, and training multiple target tasks respectively, specifically includes: Based on the different task identifiers, the encoding results of samples contained in different target tasks are identified; Based on the recognition results, the encoding results of different samples are input into the corresponding task training network; A task-trained network for image-text matching is used to decode the input encoded results and output the generated decoded text set based on the text decoding results. A task-trained network for image and text generation is used to perform image decoding on the input encoded results and output the generated decoded image set based on the image decoding results. A task-trained network for action prediction is used to perform action trajectory prediction on the input encoded results, and the generated action trajectory set is output based on the action trajectory prediction results.
7. The task fusion training method according to claim 6, characterized in that, The step of calculating the loss value corresponding to each of the multiple target tasks specifically includes: A text recognition comparison method is used to identify the number of elements in the decoded text set that are inconsistent with the reference text set, calculate the ratio of the number of inconsistent elements to the number of elements in the reference text set, and set the ratio as the loss value corresponding to the image-text matching task. An image difference recognition method is used to identify the differences between the decoded image set and the reference generated image set, and the differences are set as the loss value corresponding to the image and text generation task. A comparative recognition method is used to identify the number of differing actions between the action trajectory set and the reference predicted action set, calculate the ratio of the number of differing actions to the number of elements in the reference predicted action set, and set the ratio as the loss value corresponding to the action prediction task. The step of obtaining the comprehensive loss value of the encoder using a joint loss calculation method specifically includes: Calculate the proportion of samples included in the total training samples for each different target task; The aforementioned percentages are set as the loss weights corresponding to different target tasks; Based on the loss weights and loss values corresponding to different target tasks, the loss weights and loss values are multiplied and summed to calculate the comprehensive loss value of the encoder.
8. A task fusion training device, characterized in that, include: The training dataset acquisition module is used to acquire training datasets for multiple target tasks; The training data encoding module is used to input the training datasets of the multiple target tasks into a preset encoder and perform unified encoding processing using the same encoding mode to obtain an encoding result set. The task training execution module is used to input the encoding results of each training sample in the encoding result set into the corresponding task training network according to the different target task identifiers, and to train multiple target tasks respectively. The loss value calculation module is used to calculate the loss value corresponding to the multiple target tasks respectively, and to obtain the comprehensive loss value of the encoder by using a joint loss calculation method; The training completion determination module is used to determine that the task fusion training is completed when the comprehensive loss value meets the preset comprehensive loss threshold and the loss values corresponding to the multiple target tasks meet the corresponding loss thresholds.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the task fusion training method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the task fusion training method as described in any one of claims 1 to 7.