Task processing method and device, electronic equipment and storage medium

By performing feature mapping and parameter adjustment on the task processing model, the limitations of the task processing model in a single scenario are resolved, and task processing generalization in different scenarios is achieved, enhancing logical reasoning and common sense judgment capabilities.

CN116975696BActive Publication Date: 2026-04-07HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing task processing models can only be applied to process specific types of tasks in a single scenario, lacking generalization ability.

Method used

After acquiring the task to be processed and extracting features, the initial features are mapped to the feature space supported by the large language model using a pre-trained language-like alignment model. The parameters of the language-like alignment model are then adjusted by combining the loss value of the large language model. Finally, the task instructions and features are input into the large language model to obtain the processing results.

Benefits of technology

It improves the generalization of task processing, enabling it to handle logical reasoning and common sense judgment tasks in different scenarios, and enhances task awareness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975696B_ABST
    Figure CN116975696B_ABST
Patent Text Reader

Abstract

This application provides a task processing method, apparatus, electronic device, and storage medium, relating to the field of artificial intelligence technology. The method includes: acquiring a task to be processed; wherein the task to be processed includes: a task instruction to be processed and at least one modality of data to be processed; extracting features from the data to be processed to obtain initial features to be processed; processing the initial features to be processed based on a pre-trained language-like alignment model to obtain features supported by a large language model, which are used as language-like features to be processed; wherein, when processing a first sample task based on the initial structure of the language-like alignment model and the large language model, the loss value of the large language model is used to adjust the model parameters of the initial structure of the language-like alignment model; inputting the task instruction to be processed and the language-like features to be processed into the large language model to obtain the processing result of the task to be processed. This improves the generalization of task processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a task processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, task processing models are widely used in various scenarios. In related technologies, technicians can train task processing models to complete specific types of tasks in specified scenarios according to actual needs. For example, for speech recognition scenarios, a task processing model can be pre-trained to convert user-input speech data into text data. Or, for image processing scenarios, a task processing model can be pre-trained to detect target objects in images.

[0003] However, the task processing model obtained based on the above method can only be applied to process specific types of tasks in a single scenario. Summary of the Invention

[0004] The purpose of this application is to provide a task processing method, apparatus, electronic device, and storage medium to improve the versatility of task processing. The specific technical solution is as follows:

[0005] A first aspect of this application provides a task processing method, the method comprising:

[0006] Obtain a task to be processed; wherein the task to be processed includes: a task instruction to be processed and at least one modality of data to be processed;

[0007] Feature extraction is performed on the data to be processed to obtain the initial features to be processed;

[0008] The initial features to be processed are processed based on a pre-trained language-like alignment model to obtain features supported by the large language model, which are used as language-like features to be processed. When processing the first sample task based on the language-like alignment model with the initial structure and the large language model, the loss value of the large language model is used to adjust the model parameters of the language-like alignment model with the initial structure.

[0009] The task instruction to be processed and the language features to be processed are input into the large language model to obtain the processing result of the task to be processed.

[0010] In some embodiments, the process of processing the initial features to be processed based on a pre-trained language-like alignment model to obtain features supported by a large language model, which are used as language-like features to be processed, includes:

[0011] Based on a pre-trained self-alignment model, the initial features to be processed are aligned to a preset unified feature space to obtain the self-aligned features of the data to be processed, which are used as the self-aligned features to be processed; wherein, the self-alignment model is: trained on data with multiple modalities and consistent semantics;

[0012] The self-aligned features to be processed are processed based on a pre-trained language-like alignment model to obtain the features supported by the large language model, which are then used as the language-like features to be processed.

[0013] In some embodiments, the task to be processed includes data to be processed in multiple modalities;

[0014] The self-alignment features to be processed are processed based on a pre-trained language-like alignment model to obtain features supported by the large language model, which are used as language-like features to be processed, including:

[0015] Based on the pre-trained fusion model, feature fusion is performed on the self-aligned features of the data to be processed in each modality to obtain the fused features to be processed; wherein, when the second sample task is processed based on the fusion model of the initial structure, the language-like alignment model and the large language model, the loss value of the large language model is used to adjust the model parameters of the fusion model of the initial structure.

[0016] The fusion features to be processed are input into a pre-trained language alignment model to obtain the features supported by the large language model, which are used as the language features to be processed.

[0017] In some embodiments, the training steps of the language-like alignment model include:

[0018] Obtain a first sample task and a first sample result; wherein, the first sample task includes: a first sample task instruction and first sample data of at least one modality; the first sample result represents: the expected result obtained by processing the first sample data according to the first sample task instruction;

[0019] Feature extraction is performed on the first sample data to obtain the initial features of the first sample;

[0020] The initial features of the first sample are input into the language alignment model of the initial structure to obtain the predicted language features;

[0021] The first sample task instruction and the predicted language features are input into the large language model to obtain the first prediction result of the first sample task.

[0022] Based on the first prediction result and the first sample result, calculate the loss value;

[0023] The model parameters of the language-like alignment model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a trained language-like alignment model.

[0024] In some embodiments, the training step of the self-alignment model includes:

[0025] Obtain second sample data; wherein the second sample data contains multiple data of different modalities and with consistent semantics;

[0026] Feature extraction is performed on each of the second sample data to obtain the initial features of the respective second sample;

[0027] For any two initial features of the second sample, input the two initial features of the second sample into the self-alignment model of the initial structure to obtain the self-alignment features of the two initial features of the second sample respectively.

[0028] The loss value is calculated based on the feature distance between the self-aligned features of the two initial features of the second sample.

[0029] Adjust the model parameters of the self-aligned model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained self-aligned model.

[0030] In some embodiments, the training step of the fusion model includes:

[0031] Obtain a second sample task and a second sample result; wherein, the second sample task includes: a second sample task instruction, and third sample data of multiple different modalities; the second sample result represents: the expected result obtained by processing the third sample data according to the second sample task instruction;

[0032] Feature extraction is performed on each of the third sample data to obtain the initial features of the respective third sample;

[0033] Based on the initial features of each third sample, multiple feature combinations are obtained; wherein any feature combination contains at least two initial features of the third sample.

[0034] For each feature combination, the initial feature of the third sample in the feature combination is input into the fusion model of the initial structure to obtain the sample fusion feature corresponding to the feature combination;

[0035] The sample fusion features corresponding to the feature combination are input into the language alignment model to obtain the features supported by the large language model, and the sample language features corresponding to the feature combination are obtained.

[0036] The second sample task instruction and the sample class language features corresponding to the feature combination are input into the large language model to obtain the second prediction result corresponding to the feature combination.

[0037] Based on the second prediction result corresponding to the feature combination and the second sample result, the loss value is calculated;

[0038] The model parameters of the fusion model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a well-trained fusion model.

[0039] A second aspect of this application provides a task processing apparatus, the apparatus comprising:

[0040] The task acquisition module is used to acquire tasks to be processed; wherein, the tasks to be processed include: task instructions to be processed and at least one modality of data to be processed;

[0041] The feature extraction module is used to extract features from the data to be processed to obtain the initial features to be processed.

[0042] The feature processing module is used to process the initial features to be processed based on a pre-trained language-like alignment model to obtain the features supported by the large language model, which are used as the language-like features to be processed; wherein, when processing the first sample task based on the language-like alignment model with the initial structure and the large language model, the loss value of the large language model is used to adjust the model parameters of the language-like alignment model with the initial structure.

[0043] The result acquisition module is used to input the task instruction to be processed and the language features to be processed into the large language model to obtain the processing result of the task to be processed.

[0044] In some embodiments, the feature processing module includes:

[0045] The self-alignment submodule is used to align the initial features to be processed to a preset unified feature space based on a pre-trained self-alignment model, so as to obtain the self-aligned features of the data to be processed, which are used as the self-aligned features to be processed; wherein, the self-alignment model is obtained by training on data with multiple modalities and consistent semantics.

[0046] The feature processing submodule is used to process the self-aligned features to be processed based on a pre-trained language-like alignment model to obtain the features supported by the large language model, which are used as the language-like features to be processed.

[0047] In some embodiments, the task to be processed includes data to be processed in multiple modalities;

[0048] The feature processing submodule includes:

[0049] The fusion unit is used to perform feature fusion on the self-aligned features of the data to be processed in each modality based on a pre-trained fusion model to obtain the fused features to be processed. In the second sample task, when the fusion model based on the initial structure, the language-like alignment model and the large language model are used to process the fusion model based on the initial structure, the loss value of the large language model is used to adjust the model parameters of the fusion model based on the initial structure.

[0050] The language alignment unit is used to input the fusion features to be processed into a pre-trained language alignment model to obtain the features supported by the large language model, which are used as the language features to be processed.

[0051] In some embodiments, the training steps of the language-like alignment model include:

[0052] Obtain a first sample task and a first sample result; wherein, the first sample task includes: a first sample task instruction and first sample data of at least one modality; the first sample result represents: the expected result obtained by processing the first sample data according to the first sample task instruction;

[0053] Feature extraction is performed on the first sample data to obtain the initial features of the first sample;

[0054] The initial features of the first sample are input into the language alignment model of the initial structure to obtain the predicted language features;

[0055] The first sample task instruction and the predicted language features are input into the large language model to obtain the first prediction result of the first sample task.

[0056] Based on the first prediction result and the first sample result, calculate the loss value;

[0057] The model parameters of the language-like alignment model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a trained language-like alignment model.

[0058] In some embodiments, the training step of the self-alignment model includes:

[0059] Obtain second sample data; wherein the second sample data contains multiple data of different modalities and with consistent semantics;

[0060] Feature extraction is performed on each of the second sample data to obtain the initial features of the respective second sample;

[0061] For any two initial features of the second sample, input the two initial features of the second sample into the self-alignment model of the initial structure to obtain the self-alignment features of the two initial features of the second sample respectively.

[0062] The loss value is calculated based on the feature distance between the self-aligned features of the two initial features of the second sample.

[0063] Adjust the model parameters of the self-aligned model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained self-aligned model.

[0064] In some embodiments, the training step of the fusion model includes:

[0065] Obtain a second sample task and a second sample result; wherein, the second sample task includes: a second sample task instruction, and third sample data of multiple different modalities; the second sample result represents: the expected result obtained by processing the third sample data according to the second sample task instruction;

[0066] Feature extraction is performed on each of the third sample data to obtain the initial features of the respective third sample;

[0067] Based on the initial features of each third sample, multiple feature combinations are obtained; wherein any feature combination contains at least two initial features of the third sample.

[0068] For each feature combination, the initial feature of the third sample in the feature combination is input into the fusion model of the initial structure to obtain the sample fusion feature corresponding to the feature combination;

[0069] The sample fusion features corresponding to the feature combination are input into the language alignment model to obtain the features supported by the large language model, and the sample language features corresponding to the feature combination are obtained.

[0070] The second sample task instruction and the sample class language features corresponding to the feature combination are input into the large language model to obtain the second prediction result corresponding to the feature combination.

[0071] Based on the second prediction result corresponding to the feature combination and the second sample result, the loss value is calculated;

[0072] The model parameters of the fusion model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a well-trained fusion model.

[0073] A third aspect of this application provides an electronic device, including:

[0074] Memory, used to store computer programs;

[0075] The processor, when executing a program stored in memory, implements any of the task processing methods described above.

[0076] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the task processing methods described above.

[0077] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the task processing methods described above.

[0078] Beneficial effects of the embodiments in this application:

[0079] This application provides a task processing method, which includes: acquiring a task to be processed; wherein the task to be processed includes: a task instruction to be processed and at least one modality of data to be processed; extracting features from the data to be processed to obtain initial features to be processed; processing the initial features to be processed based on a pre-trained language alignment model to obtain features supported by a large language model, which are used as language features to be processed; wherein, when processing a first sample task based on the language alignment model and the large language model with respect to the initial structure, the loss value of the large language model is used to adjust the model parameters of the language alignment model with respect to the initial structure; and inputting the task instruction to be processed and the language features to be processed into the large language model to obtain the processing result of the task to be processed.

[0080] Based on the above processing, for any task in any scenario, the task instructions and at least one modality of data to be processed can be obtained. Feature extraction is then performed on the data to obtain initial features. These initial features can then be processed using a pre-trained language alignment model to obtain the language features to be processed. Since the loss value used to adjust the model parameters of the language alignment model during training is determined based on the processing result of the first sample task output by the large language model, the language features obtained from the language alignment model are effectively recognized by the large language model. Correspondingly, inputting the task instructions and the language features to be processed into the large language model allows the model to obtain the result of processing the data according to the task instructions, i.e., the processing result of the task in that scenario. Since large language models possess logical reasoning and common sense judgment capabilities, the task processing method provided in this application can enhance task perception capabilities, thereby enabling the processing of tasks requiring logical reasoning and common sense judgment. Thus, for different tasks in different scenarios, the task processing results can be obtained according to the task processing method provided in this application, thereby improving the generalization of task processing.

[0081] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0082] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0083] Figure 1 This is a first flowchart of a task processing method provided in an embodiment of this application;

[0084] Figure 2 A flowchart illustrating the training process of a language-like alignment model provided in this application embodiment;

[0085] Figure 3 A second flowchart of the task processing method provided in the embodiments of this application;

[0086] Figure 4 A flowchart illustrating the training process of a self-alignment model provided in this application embodiment;

[0087] Figure 5 A third flowchart of the task processing method provided in the embodiments of this application;

[0088] Figure 6 A training flowchart of a fusion model provided in an embodiment of this application;

[0089] Figure 7 An architecture diagram of a task processing system provided in this application embodiment;

[0090] Figure 8 An architectural diagram of a mode converter 01 provided in an embodiment of this application;

[0091] Figure 9 A structural diagram of a task processing device provided in an embodiment of this application;

[0092] Figure 10 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0093] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0094] Currently, task processing models are widely used in various scenarios. In related technologies, technicians can train task processing models to complete specific types of tasks in specified scenarios according to actual needs. For example, for speech recognition scenarios, a task processing model can be pre-trained to convert user-input speech data into text data. Or, for image processing scenarios, a task processing model can be pre-trained to detect target objects in images.

[0095] However, the task processing model obtained based on the above method can only be applied to process specific types of tasks in a single scenario.

[0096] This application provides a task processing method. This method can be executed by a physical task processing device implemented in hardware and / or a virtual task processing device implemented in software. Both types of task processing devices can be configured in an electronic device, such as a server or a terminal, where the terminal can be a mobile phone, computer, etc.

[0097] For example, a user can submit a task to be processed (i.e., a task to be processed) on the interactive page of the aforementioned terminal. The terminal can then obtain the processing result of the task to be processed according to the task processing method provided in this application. Subsequently, the terminal can provide feedback on the processing result of the task to the user. For instance, the terminal can display the processing result of the task on its interactive page.

[0098] See Figure 1 , Figure 1 A first flowchart of a task processing method provided in this application embodiment, the method includes the following steps:

[0099] S101: Get tasks to be processed.

[0100] The task to be processed includes: task instructions to be processed and at least one modality of data to be processed.

[0101] S102: Extract features from the data to be processed to obtain the initial features to be processed.

[0102] S103: Based on the pre-trained language-like alignment model, process the initial features to be processed to obtain the features supported by the large language model, which are used as the language-like features to be processed.

[0103] Specifically, when processing the first sample task using the language-like alignment model and the large language model based on the initial structure, the loss value of the large language model is used to adjust the model parameters of the language-like alignment model based on the initial structure.

[0104] S104: Input the task instructions and language features to be processed into the large language model to obtain the processing results of the task.

[0105] Based on the above processing, for any task in any scenario, the task instructions and at least one modality of data to be processed can be obtained. Feature extraction is then performed on the data to obtain initial features. These initial features can then be processed using a pre-trained language alignment model to obtain the language features to be processed. Since the loss value used to adjust the model parameters of the language alignment model during training is determined based on the processing result of the first sample task output by the large language model, the language features obtained from the language alignment model are effectively recognized by the large language model. Correspondingly, inputting the task instructions and the language features to be processed into the large language model allows the model to obtain the result of processing the data according to the task instructions, i.e., the processing result of the task in that scenario. Since large language models possess logical reasoning and common sense judgment capabilities, the task processing method provided in this application can enhance task perception capabilities, thereby enabling the processing of tasks requiring logical reasoning and common sense judgment. Thus, for different tasks in different scenarios, the task processing results can be obtained according to the task processing method provided in this application, thereby improving the generalization of task processing.

[0106] The task processing method provided in this application is particularly suitable for tasks with small amounts of data in open scenarios.

[0107] Regarding step S101, in this application, a task includes: task instructions and the data required.

[0108] The term "task to be processed" refers to the tasks that need to be processed currently. Correspondingly, "task to be processed instruction" refers to the task instructions contained within the task to be processed, and "data to be processed" refers to the data required by the task to be processed.

[0109] For example, in a scenario where the aforementioned electronic device serves as the terminal, a user can submit a task to be processed (i.e., a task to be processed) on the terminal's interactive page. For instance, a user can upload an audio clip and enter text containing "Convert this audio clip into text," where the uploaded audio clip represents the data to be processed for that task, and the entered text represents the task's instructions. Alternatively, a user can upload multiple images and enter text containing "How many cats are in the image?", where the uploaded images represent the data to be processed for that task, and the entered text represents the task's instructions.

[0110] Data modalities are predefined by technical personnel and can be categorized according to data type and / or format. For example, data can be categorized by type as text modality, video modality, image modality, audio modality, etc. Furthermore, image modality data can be further categorized by format as RGB (Red, Green, and Blue) image modality, NIR (Near Infrared) image modality, grayscale image modality, etc.

[0111] If the data to be processed contains only a single modality, it can be called single-modal data. If the data to be processed consists of two or more modalities, it can be called multimodal data.

[0112] Regarding step S102, if the data to be processed is unimodal data, feature extraction can be directly performed on the data to obtain the features of the data to be processed (i.e., the initial features to be processed). If the data to be processed is multimodal data, feature extraction can be performed on the data to be processed for each modality separately to obtain the features of the data to be processed for each modality (i.e., the initial features to be processed). It can be understood that for a given set of data to be processed, the modality of the initial features to be processed obtained by feature extraction on the data to be processed is consistent with the modality of the data to be processed.

[0113] In one implementation, for each modality of data to be processed, features of the data to be processed in that modality can be extracted using a method corresponding to that modality.

[0114] For example, for image modal data to be processed, features can be extracted from the data using a pre-defined image feature extraction model. This image feature extraction model can be a VGG (Visual Geometry Group) neural network model or an AlexNet neural network model, etc.

[0115] For text-based data to be processed, features can be extracted from the data using a pre-defined text feature extraction model. This model can be a Word2Vec (word vector) model, a GloVe model, or a FastText (fast text classification) model, among others.

[0116] For audio modal data to be processed, features can be extracted from the data using a pre-defined audio feature extraction model. This model can be a VGGish model or a PANN (Pretrained Audio Neural Networks) model, etc.

[0117] Regarding steps S103-S104, since the Large Language Model (LLM) can only recognize features within a specific feature space, the initial features to be processed can be processed based on a pre-trained language-like alignment model to obtain features mapped to the aforementioned specific feature space, thus obtaining the features supported by the Large Language Model (i.e., the language-like features to be processed). For example, the Large Language Model could be GPT-3 (Generative Pre-trained Transformer 3), BERT (Bidirectional Encoder Representations from Transformers), etc.

[0118] In this application, the process of aligning features to a specific feature space supported by a large language model based on a language-like alignment model can be called independent alignment. The language-like alignment model can be a multi-layer self-attention network based on a self-attention mechanism, a single linear layer, or a network model with multiple linear layers stacked together. The linear layer can be a convolutional layer or a fully connected layer. The process of processing the initial features to be processed based on the pre-trained language-like alignment model will be described in subsequent embodiments.

[0119] Since the language features to be processed can be recognized by the large language model, the task instructions and the language features to be processed can be input into the large language model to obtain the processing result of the task. For example, the large language model can output the processing result of the task in the form of human language.

[0120] In one implementation, the processing results of the task to be processed output by the large language model can also be displayed.

[0121] For example, in a scenario where the aforementioned electronic device is the terminal, and a user uploads an audio clip and inputs "convert this audio clip into text" on the terminal's interactive page, the large language model could output the following processing result: the audio content is "Good morning." Correspondingly, the terminal's interactive page could display this processing result. Alternatively, in a scenario where the aforementioned electronic device is the terminal, and a user uploads multiple images and inputs "How many cats are in the images?" on the terminal's interactive page, the large language model could output the following processing result: there are five cats in the three images. Correspondingly, the terminal's interactive page could display this processing result.

[0122] In some embodiments, see Figure 2 , Figure 2This document provides a flowchart for training a language-like alignment model, as illustrated in an embodiment of this application. Accordingly, the training steps for the language-like alignment model include:

[0123] S201: Obtain the first sample task and the first sample result.

[0124] The first sample task includes: a first sample task instruction and first sample data of at least one modality; the first sample result represents: the expected result obtained by processing the first sample data according to the first sample task instruction.

[0125] S202: Extract features from the first sample data to obtain the initial features of the first sample.

[0126] S203: Input the initial features of the first sample into the language-aligned model of the initial structure to obtain the predicted language features.

[0127] S204: Input the first sample task instruction and the predicted language features into the large language model to obtain the first prediction result of the first sample task.

[0128] S205: Calculate the loss value based on the first prediction result and the first sample result.

[0129] S206: Adjust the model parameters of the language-like alignment model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained language-like alignment model.

[0130] It is understood that, since the modalities of the data contained in different tasks may be different, in order to improve the applicability of the task processing method provided in this application embodiment, the modalities of the sample data contained in each sample task in the first sample task can be enriched during the training of the language alignment model, so that the trained language alignment model can process the initial features to be processed in each modality.

[0131] In this embodiment, each first sample task corresponds to a first sample result, which represents the expected result obtained by processing the first sample data according to the instructions of the first sample task. For example, in a scenario where the electronic device is a terminal and a user uploads an audio clip containing "Good morning" and inputs "convert this audio clip into text" on the terminal's interactive page, the expected result of the task to be processed can be the text containing "Good morning". Alternatively, in a scenario where the electronic device is a terminal and a user uploads multiple images and inputs "How many cats are in the images?" on the terminal's interactive page, if there are three cats in the multiple images, the expected result of the task to be processed can be the text containing "three".

[0132] The description of steps S201-S204 can be found in the description of steps S101-S104 in the above embodiments, and will not be repeated here. The first prediction result of the first sample task is the actual result of the first sample task, obtained based on the above steps S201-S204 and output by the large language model.

[0133] Furthermore, based on the first prediction result and the first sample result, the loss value can be calculated according to a preset loss function. For example, the loss function can be the autoregressive cross-entropy loss function of a large language model. This loss value can represent the difference between the expected result and the actual result of the first sample task. Therefore, the model parameters of the language-like alignment model with the initial structure can be adjusted based on the calculated loss value until convergence is achieved, resulting in a trained language-like alignment model.

[0134] For example, a sample task T1 and its sample result R1 can be obtained. Sample task T1 includes a sample task instruction L1 and sample data D1 of a modality M1. Then, based on steps S202-S203 above, the predicted language features F1 of the sample data D1 can be obtained. The predicted language features F1 and L1 are input into the large language model to obtain the prediction result R1' of the sample task. Then, based on step S205 above, the loss value can be obtained, and the model parameters of the initial structured language alignment model can be adjusted based on the calculated loss value until convergence is achieved, resulting in a trained language alignment model.

[0135] Based on the above processing, during the training of the language-like alignment model, the model parameters can be adjusted by considering the difference between the predicted results of the large language model and the expected results of the sample task. This ensures that the features output by the trained language-like alignment model can be effectively recognized by the large language model.

[0136] In some embodiments, see Figure 3 , Figure 3 This is a second flowchart of a task processing method provided in an embodiment of this application. Accordingly, in Figure 1 Based on this, step S103 includes:

[0137] S1031: Based on the pre-trained self-alignment model, the initial features to be processed are aligned to a preset unified feature space to obtain the self-aligned features of the data to be processed, which are used as the self-aligned features to be processed.

[0138] The self-alignment model is obtained by training on data with multiple modalities and consistent semantics.

[0139] S1032: Based on the pre-trained language-like alignment model, the self-aligned features to be processed are processed to obtain the features supported by the large language model, which are used as the language-like features to be processed.

[0140] In real-world scenarios, since the modalities of the data to be processed in different tasks may differ, language-like alignment models may need to process the initial features of different modalities. To reduce the discrepancies in how language-like alignment models process the initial features of different modalities, the initial features of each modality can be pre-aligned to a predefined unified feature space.

[0141] In this embodiment, if the data to be processed is unimodal data, the obtained initial features to be processed are unimodal features. Furthermore, the initial features to be processed can be aligned to a preset unified feature space based on a pre-trained self-alignment model. That is, the initial features to be processed are input into the pre-trained self-alignment model, and the self-aligned features output by the self-alignment model are obtained as the self-aligned features to be processed.

[0142] Correspondingly, if the data to be processed is multimodal data, the resulting initial features are composed of features from multiple modalities. Furthermore, based on a pre-trained self-alignment model, the initial features to be processed can be aligned to a preset unified feature space. That is, the initial features to be processed for each modality are input into the pre-trained self-alignment model, and the self-aligned features of each modality output by the self-alignment model are obtained as the self-aligned features to be processed.

[0143] A self-alignment model is trained on data with multiple modalities and consistent semantics. The semantic representation of the data refers to the meaning of the concepts represented by the objects in the real world corresponding to the data, and the relationships between these meanings; it is the interpretation and logical representation of the data in a specific domain. Semantic consistency among multiple data sets means that the semantics of these data sets are the same or similar. For example, the semantics of an image containing a cup can be considered consistent with the semantics of text containing "water cup." Or, the semantics of an audio recording containing "good morning" can be considered consistent with the semantics of text containing "good morning."

[0144] In this application, the process of aligning features to a preset unified feature space based on a self-alignment model can be referred to as self-alignment. The self-alignment model can be a multi-layer self-attention network based on a self-attention mechanism, a single linear layer, or a network model with multiple linear layers stacked together. The linear layer can be a convolutional layer or a fully connected layer. The specific training process of the self-alignment model will be described in detail in subsequent embodiments.

[0145] Furthermore, the specific process of processing the self-aligned features based on the pre-trained language-like alignment model will be described in subsequent embodiments.

[0146] Based on the above processing, the initial features to be processed in different modalities can be processed using a self-alignment model to obtain features that map the initial features to be processed in each modality to the same preset feature space, i.e., the self-aligned features to be processed. In this way, the feature differences between different modalities can be reduced, the speed of language-like alignment models in processing features of different modalities can be improved, and thus, the perception effect of large language models on inputs of different modalities can be enhanced.

[0147] In addition, language-like alignment models may have a limited number of modalities in the sample data used for the same semantic during training.

[0148] For example, the first sample task includes sample data of the text modality related to "cat" and sample task instructions related to "cat". However, the first sample task may lack sample data of the image modality related to "cat" and sample task instructions related to "cat". Therefore, in a scenario where the task instruction to be processed is the text "How many cats are there in the image?" and the data to be processed is an image containing cats, directly inputting the initial features to be processed and the task instructions to be processed into the language alignment model and the large language model may result in inaccurate processing results and low efficiency.

[0149] However, since the self-alignment model is trained based on sample data of multiple modalities and semantic consistency. For example, during the training process of the self-alignment model, multiple images containing cats and text containing "cat" can be used as sample data (i.e., the second sample data in this application) for training, so that after the sample data of the two modalities are mapped to the same preset feature space, the self-alignment features of the sample data of the two modalities are aligned.

[0150] Therefore, in scenarios where the task to be processed contains the text "How many cats are there in the image?" and the data to be processed is an image containing cats, the initial features of the data to be processed can be first input into a self-alignment model for processing to obtain the self-aligned features of the data. Then, based on a pre-trained language-like alignment model, the self-aligned features are processed to obtain features supported by a large language model, which are used as the language-like features to be processed, thereby improving the accuracy and efficiency of obtaining the processing results of the task.

[0151] In this way, the self-alignment model can be used to generalize to data of a single modality. Furthermore, it can improve the generalization ability of language-like alignment models, compensate for the limitations of combining data and task instructions of different modalities, and enhance the generalization ability of task processing.

[0152] In some embodiments, see Figure 4 , Figure 4 A flowchart illustrating the training process of a self-alignment model is provided for embodiments of this application. Accordingly, the training steps of the self-alignment model include:

[0153] S401: Obtain the second sample data.

[0154] The second sample data contains multiple different modalities, but with consistent semantics.

[0155] S402: Extract features from each second sample data to obtain the initial features of each second sample.

[0156] S403: For any two initial features of the second sample, input the two initial features of the second sample into the self-alignment model of the initial structure to obtain the self-alignment features of the two initial features of the second sample respectively.

[0157] S404: Calculate the loss value based on the feature distance between the self-aligned features of the two initial features of the second sample.

[0158] S405: Adjust the model parameters of the self-aligned model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained self-aligned model.

[0159] Understandably, for sample data with the same semantic meaning, during multiple training iterations of the self-alignment model, different modalities of sample data can be acquired each time for training. This enriches the modalities of the sample data contained in the second set of sample data, enabling the trained self-alignment model to process the features of the data to be processed from various modalities. Alternatively, during multiple training iterations of the self-alignment model, different semantics of sample data can be acquired each time for training. This enriches the semantics of the sample data contained in the second set of sample data, enabling the trained self-alignment model to process the features of the data to be processed from various semantic perspectives.

[0160] In this embodiment, the second sample data includes multiple data with different modalities but consistent semantics. For each modality of the second sample data, features of the second sample data of that modality can be extracted using a method corresponding to that modality. The specific process of obtaining the initial features of the second samples for each modality can be referred to the description of step S102 in the above embodiment, and will not be repeated here.

[0161] Furthermore, from the initial features of the second samples from multiple modalities, any two initial features of the second samples (which can be called a feature pair) can be selected, and the self-aligned model of the initial structure can be trained once according to steps S403-S405 above. Correspondingly, if the number of modalities of the second sample data is greater than two, another feature pair can be selected from the initial features of the second samples, and the self-aligned model of the initial structure can be trained once according to steps S403-S405 above, until all feature pairs that can be obtained based on the initial features of the second samples from multiple modalities have been used for one training.

[0162] It is understandable that after completing one training iteration according to steps S401-S405, other second sample data can be acquired for further training. In this case, the semantics of the second sample data acquired in this training iteration can be consistent with or different from the semantics of the second sample data acquired in the previous training iteration. In this application, by selecting second sample data with different semantics during multiple training iterations of the self-alignment model, the effectiveness of the self-alignment model in processing features of data with different semantics can be improved, further enhancing the generalization of task processing.

[0163] The feature distance between the self-aligned features of the two initial features of the second sample can be represented by Euclidean distance or cosine distance.

[0164] For example, during a training process, sample data from two semantically consistent modalities (sample data D2 of modality M2 and sample data D3 of modality M3) can be acquired. Then, initial features F2 of sample data D2 and initial features F3 of sample data D3 can be obtained. These initial features F2 and F3 are input into the self-aligned model of the initial structure to obtain self-aligned features F2' of initial features F2 and F3' of initial features F3. The loss is calculated based on the feature distance between self-aligned features F2' and F3', and the model parameters of the self-aligned model of the initial structure are adjusted based on the calculated loss value.

[0165] Based on the above processing, since the self-alignment model is trained on data with multiple modalities and consistent semantics, it can effectively align features of different modalities to a preset unified feature space. It can also ensure that the features of data with consistent semantics in different modalities are highly consistent after being mapped to the same preset feature space. In this way, it can avoid the loss or error of semantics represented by the self-aligned features to be processed caused by feature mapping, thereby ensuring the effect of task perception.

[0166] In some embodiments, see Figure 5 , Figure 5 This is a third flowchart of a task processing method provided in an embodiment of this application. Accordingly, the task to be processed includes data to be processed in multiple modalities. Figure 3 Based on this, step S1032 includes:

[0167] S10321: Based on the pre-trained fusion model, feature fusion is performed on the self-aligned features of the data to be processed in each modality to obtain the fused features to be processed.

[0168] Specifically, when processing the second sample task using the fusion model based on the initial structure, the language-like alignment model, and the large language model, the loss value of the large language model is used to adjust the model parameters of the fusion model based on the initial structure.

[0169] S10322: Input the fusion features to be processed into a pre-trained language-like alignment model to obtain the features supported by the large language model, which are used as the language-like features to be processed.

[0170] In real-world scenarios, since the task to be processed may involve multimodal data, the resulting self-aligned features are also composed of features from multiple modalities. To enable the obtained language-like features to be processed to incorporate features from all modalities of the data, the self-aligned features from each modality can be fused.

[0171] In this embodiment, the self-aligned features to be processed for each modality can be input into a pre-trained fusion model to obtain the fused features output by the fusion model, which are then used as the fused features to be processed. Furthermore, the fused features to be processed can be input into a pre-trained language-like alignment model to obtain the features supported by the large language model, which are then used as the language-like features to be processed.

[0172] In this application, the process of processing features from multiple modalities based on a fusion model to obtain the fused features to be processed can be referred to as fusion alignment. The fusion model can be a multi-layer self-attention network based on a self-attention mechanism, a single linear layer, or a network model with multiple linear layers stacked together. The linear layer can be a convolutional layer or a fully connected layer. The specific process of training the fusion model will be described in subsequent embodiments.

[0173] Based on the above processing, when the task to be processed includes multimodal data, the features of the multimodal data can be fused based on the fusion model. In this way, the fused features to be processed are obtained by combining the data to be processed from all modes, ensuring the full use of the data to be processed. This also allows for data supplementation through multimodal data, thereby improving the accuracy of task processing.

[0174] In some embodiments, see Figure 6 , Figure 6 This is a flowchart illustrating the training process of a fusion model provided in an embodiment of this application.

[0175] Accordingly, the training steps for the fusion model include:

[0176] S601: Obtain the second sample task and the second sample result.

[0177] The second sample task includes: the second sample task instruction and multiple third sample data of different modalities; the second sample result represents: the expected result obtained by processing the third sample data according to the second sample task instruction.

[0178] S602: Extract features from each third sample data to obtain the initial features of the respective third sample.

[0179] S603: Based on the initial features of each third sample, multiple feature combinations are obtained.

[0180] Each feature combination contains at least two initial features of the third sample.

[0181] S604: For each feature combination, input the initial feature of the third sample in the feature combination into the fusion model of the initial structure to obtain the sample fusion feature corresponding to the feature combination.

[0182] S605: Input the sample fusion features corresponding to the feature combination into the language alignment model to obtain the features supported by the large language model, and obtain the sample language features corresponding to the feature combination.

[0183] S606: Input the second sample task instruction and the sample class language features corresponding to the feature combination into the large language model to obtain the second prediction result corresponding to the feature combination.

[0184] S607: Calculate the loss value based on the second prediction result and the second sample result corresponding to the feature combination.

[0185] S608: Adjust the model parameters of the fusion model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained fusion model.

[0186] It is understood that, since the modalities of the data contained in different tasks may be different, in order to improve the applicability of the task processing method provided in this application embodiment, the modalities of the sample data contained in the second sample task can be enriched during the training of the fusion model, so that the trained fusion model can perform fusion processing on the features of each modality.

[0187] In this embodiment, the second sample task includes: a second sample task instruction, and third sample data of multiple different modalities. The semantics of the data from each modality included in the third sample data may be the same or different. Furthermore, for each modality's third sample data, features of that modality's third sample data can be extracted using a method corresponding to that modality. The specific process of obtaining the initial features of each modality's third sample data can be found in the description of step S102 in the above embodiments, and will not be repeated here.

[0188] Accordingly, at least two different modal initial features can be selected from the initial features of the third samples of each modality to obtain a feature combination. For example, initial features of the third samples of all modalities can be selected to obtain a feature combination. Then, the fusion model can be trained once according to the above steps S604-S608. Correspondingly, other feature combinations can be selected from the initial features of the third samples of each modality, and the fusion model of the initial structure can be trained once according to the above steps S604-S608, until all feature combinations that can be obtained based on the initial features of the third samples of multiple modalities are used for training. For specific implementation methods, please refer to the relevant descriptions of steps S203-S206 in the above embodiments.

[0189] For example, if the third sample data contains data from three modes: mode 1, mode 2, and mode 3, that is, there are third sample data for mode 1, mode 2, and mode 3. Accordingly, we can obtain the initial features F1 for the third sample of mode 1, F2 for the third sample of mode 2, and F3 for the third sample of mode 3. Furthermore, we can obtain the following feature combinations: Feature combination 1: F1, F2; Feature combination 2: F1, F3; Feature combination 3: F2, F3; Feature combination 4: F1, F2, F3.

[0190] In one implementation, during the training of the fusion model, a pre-trained language-like alignment model can be used, or an initial structured language-like alignment model can be used.

[0191] In this application, when using the initial structure of the language-like alignment model, the model parameters of both the initial structure's fusion model and the initial structure's language-like alignment model can be adjusted simultaneously. This allows for the joint training of the fusion model and the language-like alignment model, improving the overall performance of task processing based on the method provided in this application.

[0192] Based on the above processing, during the training process of the fusion model, it can be trained based on the feature combination of different modalities, so that the obtained fusion model can better fuse the features of different modalities and improve the fusion effect of features of different modalities.

[0193] In some embodiments, after step S104, the method further includes:

[0194] S105: Process the processing results of the task to be processed, and output the processing results of the task to be processed according to the preset output method.

[0195] In one implementation, the processing result of the task to be processed can be transformed into information with a preset structure.

[0196] In this embodiment of the application, in a scenario where the large language model outputs the processing result of the task to be processed in the form of human language, the value of the target data can be determined from the processing result of the task to be processed according to the user's actual needs, so as to generate information with a preset structure.

[0197] For example, if the data to be processed in the task is an image containing a target object, and the task instruction is text containing "get the position of the target object in the image", the value (x, y) of the target position can be determined from the processing result of the task, and the information of the generated preset structure is: position is (x, y).

[0198] For example, if the data to be processed in the task is two images, and the task instruction is text containing "similarity between the two images", the target similarity value can be determined from the processing result of the task, and the information of the generated preset structure is: similarity is 80%.

[0199] In another implementation, the modality of the processing result of the task to be processed can also be transformed.

[0200] For example, when a large language model outputs the processing result of a task in the form of human language, if the modality of the processing result is text, the processing result in the audio modality can be obtained based on an audio generation model, according to the user's actual needs. Alternatively, the processing result in the image modality can be obtained based on an image generation model.

[0201] The aforementioned preset output method can be referred to as the output method of the traditional task processing model.

[0202] Based on the above processing, the output method of the processing results of the task to be processed can be transformed. In this way, the processing results of the large language model data can be output according to the output method of the traditional task processing model.

[0203] like Figure 7 As shown, Figure 7 This is an architecture diagram of a task processing system provided in an embodiment of this application.

[0204] Figure 7 In this context, a multimodal data source represents the data to be processed in at least one modality within the task to be processed. Feature extraction can be performed on the data to be processed in different modalities to obtain the features of each modality (i.e., the initial features to be processed). For example, feature extraction can be performed on image modal data to obtain image features, on video modal data to obtain video features, on speech modal data to obtain speech features, and on other modal data to obtain features for other modalities, which will not be listed here. Furthermore, the initial features to be processed can be input into the modality converter 01, which processes the initial features to obtain the language-like features supported by the large language model (i.e., the language-like features to be processed). Inputting the task instructions (i.e., the task instructions to be processed) and the language-like features to be processed into the large language model yields the processing result of the task. For example, the large language model can output the processing result of the task in the form of human language. Subsequently, the processing result of the task to be processed can be input to the modality converter 02. The modality converter 02 is used to output the processing result of the task to be processed according to a preset output method. For example, the preset output method can be the output method of a traditional task processing model.

[0205] like Figure 8 As shown, Figure 8 This is an architectural diagram of a mode converter 01 provided in an embodiment of this application.

[0206] For each modality corresponding to modality 1, modality 2, and modality 3, the modality converter 01 can process the features of each modality separately based on an independent alignment model (i.e., the language-like alignment model in this application). For example, the features of modality 1 can be input into model A (i.e., the independent alignment model) and aligned with the language model (i.e., the large language model in this application) to obtain the language-like features of modality 1. Similarly, the features of modality 2 can be input into model A and aligned with the language model to obtain the language-like features of modality 2. The features of modality 3 can be input into model A and aligned with the language model to obtain the language-like features of modality 3.

[0207] Furthermore, for the features of each of the three modes (modality 1, modality 2, and modality 3), the modality converter 01 can also process the features of each mode separately based on a self-alignment model. For example, when features of both modality 1 and modality 2 exist simultaneously, the features of modality 1 can be input into the B model (i.e., the self-alignment model) and aligned to a preset unified feature space to obtain the self-aligned features of modality 1; and the features of modality 2 can be input into the B model and aligned to the preset unified feature space to obtain the self-aligned features of modality 2. When features of both modality 1 and modality 3 exist simultaneously, the features of modality 1 can be input into the B model (i.e., the self-alignment model) and aligned to a preset unified feature space to obtain the self-aligned features of modality 1; and the features of modality 3 can be input into the B model and aligned to the preset unified feature space to obtain the self-aligned features of modality 3. When features of both mode 2 and mode 3 exist simultaneously, the features of mode 2 can be input into the B model and aligned to the preset unified feature space to obtain the self-aligned features of mode 2; and the features of mode 3 can be input into the B model and aligned to the preset unified feature space to obtain the self-aligned features of mode 3.

[0208] For the features of each modality in modality 1, modality 2, and modality 3, modality converter 01 can also fuse the features of each modality based on a fusion model to obtain fused features, and process the fused features based on an independent alignment model to obtain the language-like features corresponding to the fused features. For example, the features of modality 1, modality 2, and modality 3 can be input into model C (i.e., the fusion model) to obtain fused features, and then the fused features can be input into model A and aligned with the language model to obtain the language-like features of the fused features.

[0209] In the technical solution of this application, all operations such as acquisition, storage, use, processing, transmission, provision and disclosure of data of any modality (including but not limited to user-inputted images, videos, voice, text, etc.) are carried out with the user's authorization.

[0210] Based on the same inventive concept, embodiments of this application provide a task processing apparatus. See also Figure 9 , Figure 9 This application provides a structural diagram of a task processing device according to an embodiment of the present application. The device includes:

[0211] The task acquisition module 901 is used to acquire tasks to be processed; wherein, the tasks to be processed include: task instructions to be processed and at least one modality of data to be processed;

[0212] Feature extraction module 902 is used to extract features from the data to be processed to obtain initial features to be processed;

[0213] The feature processing module 903 is used to process the initial features to be processed based on a pre-trained language-like alignment model to obtain the features supported by the large language model, which are used as the language-like features to be processed; wherein, when processing the first sample task based on the language-like alignment model with the initial structure and the large language model, the loss value of the large language model is used to adjust the model parameters of the language-like alignment model with the initial structure.

[0214] The result acquisition module 904 is used to input the task instruction to be processed and the language features to be processed into the large language model to obtain the processing result of the task to be processed.

[0215] In some embodiments, the feature processing module 903 includes:

[0216] The self-alignment submodule is used to align the initial features to be processed to a preset unified feature space based on a pre-trained self-alignment model, so as to obtain the self-aligned features of the data to be processed, which are used as the self-aligned features to be processed; wherein, the self-alignment model is obtained by training on data with multiple modalities and consistent semantics.

[0217] The feature processing submodule is used to process the self-aligned features to be processed based on a pre-trained language-like alignment model to obtain the features supported by the large language model, which are used as the language-like features to be processed.

[0218] In some embodiments, the task to be processed includes data to be processed in multiple modalities;

[0219] The feature processing submodule includes:

[0220] The fusion unit is used to perform feature fusion on the self-aligned features of the data to be processed in each modality based on a pre-trained fusion model to obtain the fused features to be processed. In the second sample task, when the fusion model based on the initial structure, the language-like alignment model and the large language model are used to process the fusion model based on the initial structure, the loss value of the large language model is used to adjust the model parameters of the fusion model based on the initial structure.

[0221] The language alignment unit is used to input the fusion features to be processed into a pre-trained language alignment model to obtain the features supported by the large language model, which are used as the language features to be processed.

[0222] In some embodiments, the training steps of the language-like alignment model include:

[0223] Obtain a first sample task and a first sample result; wherein, the first sample task includes: a first sample task instruction and first sample data of at least one modality; the first sample result represents: the expected result obtained by processing the first sample data according to the first sample task instruction;

[0224] Feature extraction is performed on the first sample data to obtain the initial features of the first sample;

[0225] The initial features of the first sample are input into the language alignment model of the initial structure to obtain the predicted language features;

[0226] The first sample task instruction and the predicted language features are input into the large language model to obtain the first prediction result of the first sample task.

[0227] Based on the first prediction result and the first sample result, calculate the loss value;

[0228] The model parameters of the language-like alignment model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a trained language-like alignment model.

[0229] In some embodiments, the training step of the self-alignment model includes:

[0230] Obtain second sample data; wherein the second sample data contains multiple data of different modalities and with consistent semantics;

[0231] Feature extraction is performed on each of the second sample data to obtain the initial features of the respective second sample;

[0232] For any two initial features of the second sample, input the two initial features of the second sample into the self-alignment model of the initial structure to obtain the self-alignment features of the two initial features of the second sample respectively.

[0233] The loss value is calculated based on the feature distance between the self-aligned features of the two initial features of the second sample.

[0234] Adjust the model parameters of the self-aligned model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained self-aligned model.

[0235] In some embodiments, the training step of the fusion model includes:

[0236] Obtain a second sample task and a second sample result; wherein, the second sample task includes: a second sample task instruction, and third sample data of multiple different modalities; the second sample result represents: the expected result obtained by processing the third sample data according to the second sample task instruction;

[0237] Feature extraction is performed on each of the third sample data to obtain the initial features of the respective third sample;

[0238] Based on the initial features of each third sample, multiple feature combinations are obtained; wherein any feature combination contains at least two initial features of the third sample.

[0239] For each feature combination, the initial feature of the third sample in the feature combination is input into the fusion model of the initial structure to obtain the sample fusion feature corresponding to the feature combination;

[0240] The sample fusion features corresponding to the feature combination are input into the language alignment model to obtain the features supported by the large language model, and the sample language features corresponding to the feature combination are obtained.

[0241] The second sample task instruction and the sample class language features corresponding to the feature combination are input into the large language model to obtain the second prediction result corresponding to the feature combination.

[0242] Based on the second prediction result corresponding to the feature combination and the second sample result, the loss value is calculated;

[0243] The model parameters of the fusion model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a well-trained fusion model.

[0244] This application also provides an electronic device, such as... Figure 10 As shown, it includes:

[0245] Memory 1001 is used to store computer programs;

[0246] When the processor 1002 executes the program stored in the memory 1001, it implements the steps of any of the task processing methods in the above embodiments.

[0247] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 1002, the communication interface, and the memory 1001 communicating with each other via the communication bus.

[0248] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0249] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0250] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0251] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0252] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described task processing methods.

[0253] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the task processing methods described in the above embodiments.

[0254] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.

[0255] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0256] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for apparatus, electronic devices, and computer-readable storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0257] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A task processing method, characterized in that, The method includes: Obtain a task to be processed; wherein the task to be processed includes: a task instruction to be processed and at least one modality of data to be processed; any data to be processed is text, video, image or audio; Feature extraction is performed on the data to be processed to obtain the initial features to be processed; Each modality's initial features to be processed are input into a pre-trained self-alignment model. The features of each modality's initial features to be processed, which are then aligned to a preset unified feature space, are obtained from the output of the self-alignment model and used as the self-aligned features to be processed. The self-alignment features to be processed are processed based on a pre-trained language-like alignment model to obtain features supported by the large language model, which are used as language-like features to be processed. When processing the first sample task based on the language-like alignment model with the initial structure and the large language model, the loss value of the large language model is used to adjust the model parameters of the language-like alignment model with the initial structure. The task instruction to be processed and the language features to be processed are input into the large language model to obtain the processing result of the task to be processed; The training steps of the self-alignment model include: Obtain second sample data; wherein the second sample data contains multiple data of different modalities and with consistent semantics; Feature extraction is performed on each of the second sample data to obtain the initial features of the respective second sample; For any two initial features of the second sample, input the two initial features of the second sample into the self-alignment model of the initial structure to obtain the self-alignment features of the two initial features of the second sample respectively. The loss value is calculated based on the feature distance between the self-aligned features of the two initial features of the second sample. Adjust the model parameters of the self-aligned model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained self-aligned model.

2. The method according to claim 1, characterized in that, The task to be processed contains data in multiple modalities. The self-alignment features to be processed are processed based on a pre-trained language-like alignment model to obtain features supported by the large language model, which are used as language-like features to be processed, including: Based on the pre-trained fusion model, feature fusion is performed on the self-aligned features of the data to be processed in each modality to obtain the fused features to be processed; wherein, when the second sample task is processed based on the fusion model of the initial structure, the language-like alignment model and the large language model, the loss value of the large language model is used to adjust the model parameters of the fusion model of the initial structure. The fusion features to be processed are input into a pre-trained language alignment model to obtain the features supported by the large language model, which are used as the language features to be processed.

3. The method according to claim 1, characterized in that, The training steps of the language-like alignment model include: Obtain a first sample task and a first sample result; wherein, the first sample task includes: a first sample task instruction and first sample data of at least one modality; the first sample result represents: the expected result obtained by processing the first sample data according to the first sample task instruction; Feature extraction is performed on the first sample data to obtain the initial features of the first sample; The initial features of the first sample are input into the language alignment model of the initial structure to obtain the predicted language features; The first sample task instruction and the predicted language features are input into the large language model to obtain the first prediction result of the first sample task. Based on the first prediction result and the first sample result, calculate the loss value; The model parameters of the language-like alignment model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a trained language-like alignment model.

4. The method according to claim 2, characterized in that, The training steps of the fusion model include: Obtain a second sample task and a second sample result; wherein, the second sample task includes: a second sample task instruction, and third sample data of multiple different modalities; the second sample result represents: the expected result obtained by processing the third sample data according to the second sample task instruction; Feature extraction is performed on each of the third sample data to obtain the initial features of the respective third sample; Based on the initial features of each third sample, multiple feature combinations are obtained; wherein any feature combination contains at least two initial features of the third sample. For each feature combination, the initial feature of the third sample in the feature combination is input into the fusion model of the initial structure to obtain the sample fusion feature corresponding to the feature combination; The sample fusion features corresponding to the feature combination are input into the language alignment model to obtain the features supported by the large language model, and the sample language features corresponding to the feature combination are obtained. The second sample task instruction and the sample class language features corresponding to the feature combination are input into the large language model to obtain the second prediction result corresponding to the feature combination. Based on the second prediction result corresponding to the feature combination and the second sample result, the loss value is calculated; The model parameters of the fusion model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a well-trained fusion model.

5. A task processing device, characterized in that, The device includes: The task acquisition module is used to acquire tasks to be processed; wherein, the tasks to be processed include: task instructions to be processed and at least one modality of data to be processed; any data to be processed can be text, video, image or audio. The feature extraction module is used to extract features from the data to be processed to obtain the initial features to be processed. The self-alignment submodule is used to input the initial features to be processed for each modality into the pre-trained self-alignment model, and obtain the features of each modality after the initial features to be processed are aligned to the preset unified feature space, which are then used as the self-aligned features to be processed. The feature processing submodule is used to process the self-aligned features to be processed based on a pre-trained language-like alignment model to obtain the features supported by the large language model, which are used as the language-like features to be processed; wherein, when processing the first sample task based on the language-like alignment model with the initial structure and the large language model, the loss value of the large language model is used to adjust the model parameters of the language-like alignment model with the initial structure. The result acquisition module is used to input the task instruction to be processed and the language features to be processed into the large language model to obtain the processing result of the task to be processed; The training steps of the self-alignment model include: Obtain second sample data; wherein the second sample data contains multiple data of different modalities and with consistent semantics; Feature extraction is performed on each of the second sample data to obtain the initial features of the respective second sample; For any two initial features of the second sample, input the two initial features of the second sample into the self-alignment model of the initial structure to obtain the self-alignment features of the two initial features of the second sample respectively. The loss value is calculated based on the feature distance between the self-aligned features of the two initial features of the second sample. Adjust the model parameters of the self-aligned model of the initial structure based on the calculated loss value until convergence is achieved, and obtain the trained self-aligned model.

6. The apparatus according to claim 5, characterized in that, The task to be processed contains data in multiple modalities. The feature processing submodule includes: The fusion unit is used to perform feature fusion on the self-aligned features of the data to be processed in each modality based on a pre-trained fusion model to obtain the fused features to be processed. In the second sample task, when the fusion model based on the initial structure, the language-like alignment model and the large language model are used to process the fusion model based on the initial structure, the loss value of the large language model is used to adjust the model parameters of the fusion model based on the initial structure. The language alignment unit is used to input the fusion features to be processed into a pre-trained language alignment model to obtain the features supported by the large language model, which are used as the language features to be processed.

7. The apparatus according to claim 5, characterized in that, The training steps of the language-like alignment model include: Obtain a first sample task and a first sample result; wherein, the first sample task includes: a first sample task instruction and first sample data of at least one modality; the first sample result represents: the expected result obtained by processing the first sample data according to the first sample task instruction; Feature extraction is performed on the first sample data to obtain the initial features of the first sample; The initial features of the first sample are input into the language alignment model of the initial structure to obtain the predicted language features; The first sample task instruction and the predicted language features are input into the large language model to obtain the first prediction result of the first sample task. Based on the first prediction result and the first sample result, calculate the loss value; The model parameters of the language-like alignment model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a trained language-like alignment model.

8. The apparatus according to claim 6, characterized in that, The training steps of the fusion model include: Obtain a second sample task and a second sample result; wherein, the second sample task includes: a second sample task instruction, and third sample data of multiple different modalities; the second sample result represents: the expected result obtained by processing the third sample data according to the second sample task instruction; Feature extraction is performed on each of the third sample data to obtain the initial features of the respective third sample; Based on the initial features of each third sample, multiple feature combinations are obtained; wherein any feature combination contains at least two initial features of the third sample. For each feature combination, the initial feature of the third sample in the feature combination is input into the fusion model of the initial structure to obtain the sample fusion feature corresponding to the feature combination; The sample fusion features corresponding to the feature combination are input into the language alignment model to obtain the features supported by the large language model, and the sample language features corresponding to the feature combination are obtained. The second sample task instruction and the sample class language features corresponding to the feature combination are input into the large language model to obtain the second prediction result corresponding to the feature combination. Based on the second prediction result corresponding to the feature combination and the second sample result, the loss value is calculated; The model parameters of the fusion model with the initial structure are adjusted based on the calculated loss value until convergence is achieved, resulting in a well-trained fusion model.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Disease diagnosis method and device based on multi-modal data, equipment and medium

    CN116259407A