Multi-task processing method and device, computer equipment, storage medium and computer program product

By using the feature extraction network and task processing unit in the content processing model, the problems of training time and resource consumption in multimodal social platforms are solved, and a rapid improvement in multi-task processing efficiency is achieved.

CN121327451APending Publication Date: 2026-01-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410920454.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In multimodal social platforms, existing technologies require separate training for different modalities, resulting in long training times, large sample data volumes, and reduced model training efficiency and task processing efficiency.

Method used

The feature extraction network in the content processing model is used to extract modal features of the target interaction data. The task processing units trained with the same interaction data samples are used to process the modal features separately, sharing the feature extraction results, saving training costs and improving task processing speed.

Benefits of technology

In multi-task scenarios, it enables rapid feature extraction and processing of single-modal and multi-modal data, improving task processing efficiency and reducing training resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121327451A_ABST
    Figure CN121327451A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-task processing method and device, computer equipment, a storage medium and a computer program product. The method comprises: in an interaction scene supporting multi-modal interaction, in response to a multi-task processing request, determining target interaction data indicated by the multi-task processing request, the target interaction data having at least one data modal; based on a feature extraction network in the content processing model, according to a target data mode corresponding to the target interaction data, extracting modal features of the target interaction data; extracting a plurality of processing tasks carried in the multi-task processing request, and respectively determining target task processing units matched with the processing tasks from the content processing model; and based on each target task processing unit, performing task processing on the modal features to obtain a multi-task processing result of the target interaction data. By adopting the method, the multi-task processing accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for multitasking. Background Technology

[0002] With the continuous development of technology and the internet, more and more social media platforms are becoming ubiquitous in people's lives, facilitating online communication. However, these platforms are also frequented by sensitive organizations engaging in sensitive activities and spreading harmful information. To evade monitoring of such information, these organizations often employ various countermeasures, such as repeatedly probing the platform with multiple modalities including text and images. Therefore, identifying multimodal abnormal information can help prevent its display on various social media platforms.

[0003] Currently, different modalities can be trained to classify and identify different modal information to determine whether text or images are abnormal. However, since different modalities require separate training, and different abnormal information and sample recall are needed in different application scenarios, sample acquisition and full parameter training for different scenarios require a long training time and a large amount of sample data. This not only increases the consumption of training resources but also reduces the efficiency of model training, thereby reducing the efficiency of task processing when there are multiple tasks. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer device, storage medium, and computer program product for multitasking that can improve the efficiency of task processing when multiple tasks exist, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for multitasking. The method includes:

[0006] In an interactive scenario that supports multimodal interaction, in response to a multi-task processing request, the target interactive data indicated by the multi-task processing request is determined, and the target interactive data has at least one data modality.

[0007] Based on the feature extraction network in the content processing model, modal features of the target interaction data are extracted according to the target data modality corresponding to the target interaction data.

[0008] Extract multiple processing tasks carried in the multi-task processing request, and determine the target task processing unit that matches each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples.

[0009] Based on each target task processing unit, modal features are processed separately to obtain the multi-task processing results of target interaction data.

[0010] Secondly, this application also provides a multitasking processing apparatus. The apparatus includes:

[0011] The data determination module is used to determine the target interactive data indicated by the multi-task processing request in response to the multi-task processing request in an interactive scenario that supports multimodal interaction. The target interactive data has at least one data modality.

[0012] The feature extraction module is used to extract modal features of the target interaction data based on the feature extraction network in the content processing model, according to the target data modality corresponding to the target interaction data.

[0013] The processing unit determination module is used to extract multiple processing tasks carried in the multi-task processing request and determine the target task processing unit that matches each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples.

[0014] The task processing module is used to perform task processing on modal features based on each target task processing unit to obtain the multi-task processing results of the target interaction data.

[0015] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0016] In an interactive scenario that supports multimodal interaction, in response to a multi-task processing request, the target interactive data indicated by the multi-task processing request is determined, and the target interactive data has at least one data modality.

[0017] Based on the feature extraction network in the content processing model, modal features of the target interaction data are extracted according to the target data modality corresponding to the target interaction data.

[0018] Extract multiple processing tasks carried in the multi-task processing request, and determine the target task processing unit that matches each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples.

[0019] Based on each target task processing unit, modal features are processed separately to obtain the multi-task processing results of target interaction data.

[0020] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0021] In an interactive scenario that supports multimodal interaction, in response to a multi-task processing request, the target interactive data indicated by the multi-task processing request is determined, and the target interactive data has at least one data modality.

[0022] Based on the feature extraction network in the content processing model, modal features of the target interaction data are extracted according to the target data modality corresponding to the target interaction data.

[0023] Extract multiple processing tasks carried in the multi-task processing request, and determine the target task processing unit that matches each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples.

[0024] Based on each target task processing unit, modal features are processed separately to obtain the multi-task processing results of target interaction data.

[0025] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0026] In an interactive scenario that supports multimodal interaction, in response to a multi-task processing request, the target interactive data indicated by the multi-task processing request is determined, and the target interactive data has at least one data modality.

[0027] Based on the feature extraction network in the content processing model, modal features of the target interaction data are extracted according to the target data modality corresponding to the target interaction data.

[0028] Extract multiple processing tasks carried in the multi-task processing request, and determine the target task processing unit that matches each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples.

[0029] Based on each target task processing unit, modal features are processed separately to obtain the multi-task processing results of target interaction data.

[0030] The aforementioned multi-task processing methods, apparatuses, computer devices, storage media, and computer program products, in interactive scenarios supporting multimodal interaction, respond to multi-task processing requests, determine the target interactive data with at least one data modality indicated by the multi-task processing request, and then, based on the feature extraction network in the content processing model, extract the modal features of the target interactive data according to the target data modality corresponding to the target interactive data. That is, whether it is single-modal or multimodal data, the feature extraction network can be used to extract the corresponding modal features to ensure the reliability of feature extraction. The multiple processing tasks carried in the multi-task processing request are then extracted. Target task processing units matching each processing task are determined from the content processing model. This eliminates the need for repetitive feature extraction for different processing tasks. Instead, each target task processing unit processes the modal features separately to obtain the multi-task processing results of the target interactive data. Therefore, the content processing model shares the feature extraction results for the data by fixing the feature extraction parameters. Since the content processing model includes multiple task processing units trained using the same interactive data samples, it accelerates the task prediction speed of the content processing model for different tasks. While saving training costs, it can quickly extract features for both single-modal and multi-modal data and provide timely response processing for multiple tasks, thereby improving the efficiency of task processing when multiple tasks exist. Attached Figure Description

[0031] Figure 1 This is an application environment diagram of a multi-task processing method in one embodiment;

[0032] Figure 2 A simplified flowchart of the multitasking technique in one embodiment;

[0033] Figure 3 This is a simplified model structure diagram of the content processing model in one embodiment;

[0034] Figure 4 This is a flowchart illustrating a multi-task processing method in one embodiment;

[0035] Figure 5 This is a schematic diagram illustrating the data classification process performed by a target task processing unit for data classification in one embodiment.

[0036] Figure 6 This is a schematic diagram illustrating the data retrieval process performed by a target task processing unit for data retrieval in one embodiment.

[0037] Figure 7 This is a schematic diagram illustrating the target detection process performed by a target task processing unit for target detection in one embodiment.

[0038] Figure 8 This is a schematic diagram illustrating the process of acquiring a task processing unit for data classification in one embodiment;

[0039] Figure 9 This is a schematic diagram of the feature masking process in one embodiment;

[0040] Figure 10 This is a schematic diagram illustrating the process of acquiring a task processing unit for target detection in one embodiment;

[0041] Figure 11 This is a schematic diagram of the target detection process in one embodiment;

[0042] Figure 12 This is a complete flowchart of a multi-task processing method in one embodiment;

[0043] Figure 13 This is a structural block diagram of a multitasking device in one embodiment;

[0044] Figure 14 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0046] With the continuous development of technology and the internet, more and more social media platforms have become ubiquitous in people's lives, facilitating online communication. However, these platforms often harbor sensitive organizations that engage in sensitive activities and spread harmful information. To evade monitoring of abnormal information, these organizations often employ various countermeasures, such as repeatedly probing the platform with multiple modalities including text and images. Therefore, identifying multimodal abnormal information can prevent its display on various social media platforms. Currently, different modalities can be trained separately to classify and identify different types of information, determining whether text or images are abnormal. However, since different modalities require separate training, and different abnormal information and application scenarios necessitate sample recall, obtaining samples and performing full-parameter training for different scenarios requires significant time and a large amount of data. This increases training resource consumption and reduces model training efficiency, thereby decreasing task processing efficiency when multiple tasks are involved.

[0047] Therefore, embodiments of this application provide a multitasking method that can improve task processing efficiency when multiple tasks exist. The multitasking method provided in embodiments of this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another server.

[0048] Specifically, taking server 104 as an example, in an interactive scenario supporting multimodal interaction, server 104 responds to a multi-task processing request, determines the target interactive data indicated by the request, and the target interactive data has at least one data modality. Then, based on the feature extraction network in the content processing model, it extracts the modal features of the target interactive data according to the target data modality corresponding to the target interactive data, and extracts multiple processing tasks carried in the multi-task processing request. It then determines target task processing units matching each processing task from the content processing model, which includes multiple task processing units trained using the same interactive data samples. Finally, based on each target task processing unit, it performs task processing on the modal features to obtain the multi-task processing result of the target interactive data. Since both unimodal and multimodal data can have their corresponding modal features extracted through feature extraction networks, ensuring the reliability of feature extraction, and then the content processing model can share the feature extraction results for the data by using fixed feature extraction parameters, and since the content processing model includes multiple task processing units trained with the same interactive data samples, that is, different task processing units perform corresponding task processing, the task prediction speed of the content processing model for different tasks can be accelerated. While saving training costs, it can perform fast feature extraction for both unimodal and multimodal data, and perform timely response processing for multiple tasks, thereby improving the efficiency of task processing when multiple tasks exist.

[0049] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein. The multi-tasking method provided in the embodiments of this application can be applied to various scenarios, including but not limited to cloud technology and artificial intelligence.

[0050] This application primarily applies to multimodal interaction scenarios. These scenarios can include, but are not limited to, public interaction platforms (such as web forums and public comment sections), personal interaction platforms (such as personal comment sections), object-to-object interaction applications (such as object-to-object chat programs), and multimedia applications (such as bullet comments or comments in audio and video programs). No specific limitations are set here. Furthermore, the main purpose of this application is to filter and identify abnormal data in the interaction data, that is, to detect abnormal or risky data in multimodal interaction scenarios to prevent the repeated use and transmission of abnormal or risky data within the interaction context.

[0051] For ease of understanding, the technical process provided in the embodiments of this application is briefly described below, such as... Figure 2The diagram shows a simplified flowchart of the multi-task processing technique. First, the target interaction data is used as input to the content processing model. Then, based on the feature extraction network in the content processing model, modal features of the target interaction data can be extracted according to the target data modality corresponding to the target interaction data. Therefore, modal features can be one of the output components of the content processing model. Further, target task processing units matching each processing task need to be determined from the content processing model. Based on each target task processing unit, task processing is performed on the modal features to obtain the multi-task processing result of the target interaction data. The multiple processing tasks include at least two of the following: classification tasks, retrieval tasks, and object detection tasks. Taking multiple processing tasks including classification, retrieval, and object detection tasks as an example, modal features can be classified for the target task processing unit used for data classification to obtain the classification result for the target interaction data. Secondly, for the target task processing unit used for data retrieval, the feature similarity between the modal features and the data features of each candidate interaction data can be determined. Then, from the multiple candidate interaction data, retrieval interaction data whose feature similarity with the target interaction data meets the retrieval conditions is selected. Furthermore, it can determine target parameters that match the target detection task to indicate the target object. Based on modal features and target parameters, the target object is detected by the target task processing unit used for target detection, yielding target detection results for the target interaction data. Therefore, the output can consist of modal features, classification results, retrieval interaction data, and target detection results, enabling the detection and retrieval of target interaction data from different detection interactions based on different tasks, thereby improving the detection efficiency of abnormal or risky data in interactive scenarios.

[0052] Furthermore, the following is a brief description of the content processing model provided in the embodiments of this application, such as... Figure 3 The diagram shows a simplified model structure for the content processing model. The model includes a feature extraction network and multiple task processing units for different tasks. Therefore, text data, image-text data, and image data can be used as input to the content processing model. The feature extraction network then extracts features from these data to obtain text features, image features, and a combination of text and image features. Based on this, task processing units for data classification can perform data classification based on either text or image features. For example, abnormal data classification could be required, such as vulgar data, abnormal social group data, or inappropriate data. Different classification heads can be used to address these different classification needs. Abnormal social groups are defined as social groups that spread a large amount of repetitive information.

[0053] Secondly, the task processing unit used for data retrieval can determine the feature similarity between any of the text data, image data, and graphic data and the data features of the candidate interactive data. From multiple candidate interactive data, the retrieval interactive data whose feature similarity with the input data meets the retrieval conditions can be selected to complete the data retrieval. In other words, considering that in the initial stage of finding abnormal data in practical applications, a small number of abnormal data samples may result in low accuracy of training results, similarity retrieval can be performed on a small number of abnormal data to recall retrieval interactive data similar to the abnormal data, thereby preventing the large-scale spread of similar abnormal data and providing more reliable sample data for model training.

[0054] Furthermore, a task processing unit for object detection can be used to detect target objects and obtain object detection results. Object detection is only applicable to image data, that is, determining whether the object to be detected exists in the image data. In practical applications, object detection can also be performed on text data, that is, the presence of abnormal text in the text indicates that the text is abnormal. Figure 3 The examples provided are for understanding the content processing model and are a brief structure and process for multi-task processing of modal features of multimodal data. They should not be construed as specific limitations of this application.

[0055] As described above, the multi-task processing method provided in this application involves a content processing model and a feature extraction network, meaning that the multi-task processing method also involves Artificial Intelligence (AI). AI will be introduced below. AI is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making functions.

[0056] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0057] The content processing model and feature extraction network provided in this application specifically relate to Machine Learning (ML) technology under artificial intelligence. Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and pre-trained learning. Pre-trained models are the latest development in deep learning, integrating the above techniques.

[0058] Secondly, since the content processing model in this application can process both unimodal and multimodal data—that is, image data, text data, and image-text data—the processing of such data also involves computer vision (CV) technology under artificial intelligence. Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes for target recognition and measurement, and further performs image processing to create images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, attempting to establish artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the vision field, such as Swin-transformer, ViT, V-MOE, and MAE, can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0059] The following examples illustrate this in detail: In one example, as... Figure 4 As shown, a multi-task processing method is provided, which is applied to... Figure 1 Taking server 104 as an example, it can be understood that this method can also be applied to a system including terminal 102 and server 104, and implemented through the interaction between terminal 102 and server 104. In this embodiment, the method includes the following steps:

[0060] Step 402: In an interaction scenario that supports multimodal interaction, in response to a multi-task processing request, determine the target interaction data indicated by the multi-task processing request, wherein the target interaction data has at least one data modality.

[0061] Multimodal interaction specifically supports the interaction of data in multiple different modalities. These modalities can include image data modalities and text data modalities, and in practical applications, can also include audio data modalities and video data modalities, etc., without further limitation. Based on this, multimodal interaction scenarios can include, but are not limited to: public interaction platforms (such as web forums and public comment sections), personal interaction platforms (such as personal comment sections), object-to-object interaction applications (such as object-to-object chat programs), and multimedia applications (such as bullet comments or comments in audio and video programs).

[0062] Secondly, multitasking requests are used to instruct multiple different tasks to be performed on specified interactive data. Therefore, a multitasking request indicates target interactive data, which is the interactive data that requires task processing. For example, in a public interactive platform, comment data sent by a platform object is target interactive data; in an object-interaction application, chat data sent by a chat object is target interactive data; and in a multimedia application, bullet screen data or comment data sent by an application user object is target interactive data. Since multimodal interaction is supported, the target interactive data has at least one data modality, meaning it consists of interactive data of at least one data modality. For example, the target interactive data can be image data of the image data modality (unimodal data), or text data of the text data modality (unimodal data), or a combination of text and image data of the image-text data modality (multimodal data). In practical applications, audio and video data modalities may also be included; therefore, the target interactive data may also include audio and video data. No specific limitations are imposed on the target interactive data here.

[0063] Furthermore, as described above, a multi-task processing request instructs the processing of specified interactive data into multiple different tasks. To perform these tasks, the request must contain multiple processing tasks, each with a distinct objective. Therefore, these processing tasks consist of at least two of the following: classification, retrieval, and object detection. The classification task categorizes the interactive data. The classification type is determined based on the specific needs of the classification. For example, in a scenario requiring the identification of vulgar data, the classification header would be "vulgar data." Similarly, in a scenario requiring the identification of abnormal social groups, the classification header would be "abnormal social group data." And in a scenario requiring the identification of inappropriate data, the classification header would be "inappropriate data." The retrieval task retrieves interactive data that meets the similarity requirements to the interactive data. Finally, the object detection task detects the presence of a target object from the interactive data in the image data modality.

[0064] Specifically, in interactive scenarios that support multimodal interaction, if it is necessary to perform content recognition and classification of interactive data during multimodal interaction, then it is necessary to generate a multi-task processing request based on the task processing requirements for the target interactive data to be processed. For example, if it is necessary to perform data classification and target detection on the comment data sent by the platform object in the public interactive platform, then the comment data sent by the platform object in the public interactive platform is identified as the target interactive data, and the task processing for the aforementioned target interactive data is identified as data classification and target detection. Thus, a multi-task processing request is generated. At this time, the multi-task processing request indicates the comment data (i.e., the target interactive data) sent by the platform object in the public interactive platform, and carries the processing tasks of data classification and target detection.

[0065] Based on this, in response to a multi-task processing request generated in an interactive scenario that supports multimodal interaction, the server first determines the target interactive data indicated by the multi-task processing request. As in the previous example, the comment data sent by the platform object in the public interactive platform indicated in the multi-task processing request can be determined as the target interactive data.

[0066] Step 404: Based on the feature extraction network in the content processing model, extract the modal features of the target interaction data according to the target data modality corresponding to the target interaction data.

[0067] In this context, the modal features of the target interactive data correspond to the target data modality. For example, if the target interactive data is image data in the image data modality, then the modal features of the target interactive data are image features. Similarly, if the target interactive data is text data in the text data modality, then the modal features of the target interactive data are text features. Furthermore, the target interactive data can also consist of text data and image data in the image-text data modality, in which case the modal features of the target interactive data are image features and text features. Considering that in practical applications, audio data modalities and video data modalities may also be included, the modal features can also include modal features corresponding to other data modalities, such as audio features.

[0068] Secondly, as described above, the content processing model includes a feature extraction network and multiple task processing units for different tasks. The feature extraction network has fixed parameters and is not adjusted during model training. That is, the feature extraction network in this application has fixed parameters. Specifically, the feature extraction network is the encoder in the Contrastive Language-Image Pre-training (CLIP) model, and the encoder is specifically divided into a text encoder and an image encoder.

[0069] Specifically, the server, based on the feature extraction network in the content processing model, extracts modal features of the target interactive data according to the target data modality corresponding to the target interactive data. Since the target interactive data has at least one data modality, and different feature extraction methods should be used for different data modalities, the feature extraction network includes a feature extraction sub-network, and the feature extraction sub-network corresponds to the data modality. That is, the server extracts features according to the target data modality corresponding to the target interactive data, through the feature extraction sub-network corresponding to the data modality included in the target data modality, to obtain the modal features of the target interactive data.

[0070] For example, if feature extraction sub-network A1 is a text encoder, then feature extraction sub-network A1 is used to extract modal features of text data modality; that is, the modal features output by feature extraction sub-network A1 are text features. Similarly, if feature extraction sub-network A2 is an image encoder, then feature extraction sub-network A2 is used to extract modal features of image data modality; that is, the modal features output by feature extraction sub-network A2 are image features. Therefore, if the target interactive data is text and image data, then feature extraction sub-networks A1 and A2 need to work together to extract features from the target interactive data. Feature extraction sub-network A1 extracts text data belonging to the text data modality, while feature extraction sub-network A2 extracts image data belonging to the image data modality. The text features output by feature extraction sub-network A1 and the image features output by feature extraction sub-network A2 together constitute the modal features of the target interactive data; that is, the modal features of the target interactive data in this case include both text features and image features.

[0071] Step 406: Extract the multiple processing tasks carried in the multi-task processing request, and determine the target task processing unit that matches each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples.

[0072] As described above, a multi-task processing request instructs the processing of specified interactive data into multiple different tasks. To perform these tasks, the request must contain multiple processing tasks, each with a distinct objective. Therefore, these multiple processing tasks must consist of at least two of the following: classification, retrieval, and object detection. Furthermore, the content processing model comprises multiple task processing units trained using the same interactive data samples. In other words, each task processing unit is trained using the same interactive data samples, through a feature extraction network with fixed parameters, to extract features for its corresponding task.

[0073] Specifically, the server extracts multiple processing tasks from the multi-task processing request, such as extracting classification tasks, retrieval tasks, and object detection tasks, or extracting classification tasks and object detection tasks, or extracting retrieval tasks and object detection tasks, or extracting classification tasks and retrieval tasks. Further, the server determines the target task processing unit matching each processing task from the content processing model; that is, the server determines the target task processing unit for executing the processing task from the content processing model.

[0074] Since the specific examples in this application include classification tasks, retrieval tasks, and object detection tasks, the following will describe the methods for determining the target task processing units that include the aforementioned tasks in multiple processing tasks, and will further explain in detail the methods for executing different target task processing units in the next step. First, let's introduce the classification task:

[0075] In one optional embodiment, the multiple processing tasks include at least a classification task. The classification task is used to classify the target interaction data, and the classification type can be determined based on the actual classification requirements. For example, in a scenario requiring the identification of vulgar data, the classification head would be "vulgar data," meaning determining whether the target interaction data is vulgar. Similarly, in a scenario requiring the identification of abnormal social group data, the classification head would be "abnormal social group data," meaning determining whether the target interaction data is abnormal social group data. And, in a scenario requiring the identification of inappropriate data, the classification head would be "inappropriate data," meaning determining whether the target interaction data is inappropriate. In these cases, different classification heads are used for different classification requirements to complete the corresponding data classification.

[0076] Based on this, target task processing units matching each processing task are determined from the content processing model, including: for classification tasks, target task processing units for data classification are determined from the content processing model.

[0077] In the content processing model, task processing units used for different tasks can be labeled. Each task processing unit carries a unit identifier for performing the task, and each processing task has a corresponding task identifier. Therefore, there is a mapping relationship between the unit identifier that uniquely identifies a task processing unit and the task identifier corresponding to the processing task performed by that unit. For example, task identifier B1 indicates a classification task, task identifier B2 indicates a retrieval task, and task identifier B3 indicates a target detection task. Furthermore, task processing unit C1 (the unit performing the classification task) carries unit identifier D1, task processing unit C2 (the unit performing the retrieval task) carries unit identifier D2, and task processing unit C3 (the unit performing the target detection task) carries unit identifier D3. Thus, there is a mapping relationship between task identifier B1 and unit identifier D1, between task identifier B2 and unit identifier D2, and between task identifier B3 and unit identifier D3.

[0078] Specifically, for a classification task, the server determines the target task processing unit for data classification from the content processing model. That is, based on the task identifier indicating the classification task, the server searches for the target unit identifier that has an identifier mapping relationship with the task identifier indicating the classification task from the unit identifiers corresponding to the multiple task processing units included in the content processing model. Then, the server determines the task processing unit carrying the target unit identifier as the target task processing unit for data classification.

[0079] For ease of understanding, based on the aforementioned example, it can be seen that the task identifier indicating the classification task is specifically task identifier B1, and the unit identifier that has an identifier mapping relationship with task identifier B1 is unit identifier D1, that is, the target unit identifier is unit identifier D1, and the task processing unit carrying the target unit identifier (unit identifier D1) is task processing unit C1. Therefore, task processing unit C1 can be identified as the task processing unit that performs the classification task, that is, the target task processing unit used for data classification in the content processing model is task processing unit C1.

[0080] In one optional embodiment, the multiple processing tasks include at least a retrieval task. The retrieval task is used to retrieve retrieval interaction data that meets the requirements for similarity to the interaction data. That is, considering practical applications, in the initial stages of finding anomalous data, a small number of anomalous data samples may compromise the accuracy of the training results. In this case, similarity retrieval can be performed on the small number of anomalous data samples to recall retrieval interaction data similar to the anomalous data, thereby preventing the large-scale propagation of similar anomalous data and providing more reliable sample data for model training. Therefore, a retrieval task is necessary.

[0081] Based on this, target task processing units matching each processing task are determined from the content processing model, including: for retrieval tasks, target task processing units for data retrieval are determined from the content processing model.

[0082] Specifically, for a retrieval task, the server determines the target task processing unit for data retrieval from the content processing model. That is, based on the task identifier indicating the retrieval task, the server searches for the target unit identifier that has an identifier mapping relationship with the task identifier indicating the retrieval task from the unit identifiers corresponding to the multiple task processing units included in the content processing model. Then, the task processing unit carrying the target unit identifier is determined as the target task processing unit for data retrieval.

[0083] For ease of understanding, based on the aforementioned example, it can be seen that the task identifier indicating the retrieval task is specifically task identifier B2, and the unit identifier that has an identifier mapping relationship with task identifier B2 is unit identifier D2, that is, the target unit identifier is unit identifier D2, and the task processing unit carrying the target unit identifier (unit identifier D2) is task processing unit C2. Therefore, task processing unit C2 can be identified as the task processing unit that executes the retrieval task, that is, the target task processing unit used for data retrieval in the content processing model is task processing unit C2.

[0084] In one optional embodiment, the multiple processing tasks include at least an object detection task. The object detection task is used to perform object detection on a target object, obtaining the object detection result. Object detection is only applied to image data, that is, determining whether the object to be detected exists in the image data. For example, detecting whether there are abnormal objects or vulgar objects in the image data. In practical applications, object detection can also be performed on text data; that is, the presence of abnormal text in the text indicates that the text is abnormal.

[0085] Based on this, target task processing units matching each processing task are determined from the content processing model, including: for the target detection task, target task processing units for target detection are determined from the content processing model.

[0086] Specifically, for the target detection task, the server determines the target task processing unit for target detection from the content processing model. That is, based on the task identifier that indicates the target detection task, the server finds the target unit identifier that has an identifier mapping relationship with the task identifier that indicates the target detection task from the unit identifiers corresponding to the multiple task processing units included in the content processing model, and then determines the task processing unit carrying the target unit identifier as the target task processing unit for target detection.

[0087] For ease of understanding, based on the aforementioned example, it can be seen that the task identifier indicating the target detection task is specifically task identifier B3, and the unit identifier that has an identifier mapping relationship with task identifier B3 is unit identifier D3, that is, the target unit identifier is unit identifier D3, and the task processing unit carrying the target unit identifier (unit identifier D3) is task processing unit C3. Therefore, task processing unit C3 can be identified as the task processing unit that performs the target detection task, that is, the target task processing unit used for target detection in the content processing model is task processing unit C3.

[0088] Step 408: Based on each target task processing unit, perform task processing on the modal features to obtain the multi-task processing results of the target interaction data.

[0089] Specifically, the server processes the modal features based on each target task processing unit to obtain the multi-task processing result of the target interaction data. In other words, the server uses target task processing units to process different tasks to perform corresponding task processing on the modal features of the obtained target interaction data, thereby obtaining the multi-task processing result of the target interaction data. The multi-task processing result includes the task processing results output by each target task processing unit.

[0090] In other words, when processing tasks including classification, the target task processing unit is used for data classification. In this case, the target task processing unit will output classification results for the target interaction data, so the multi-task processing result must at least include the classification results. Similarly, when processing tasks including retrieval, the target task processing unit is used for data retrieval. In this case, the target task processing unit will output retrieval interaction data that meets the retrieval criteria, so the multi-task processing result must at least include the retrieval interaction data. Likewise, when processing tasks including object detection, the target task processing unit is used for object detection. In this case, the target task processing unit will output object detection results for the target interaction data, so the multi-task processing result must at least include the object detection results. The following will detail how the target task processing unit obtains the corresponding results for different task processing tasks:

[0091] In one optional embodiment, when the processing task includes a classification task, the target task processing unit is used for data classification. The classification task also needs to indicate the target category; for example, if the task is to detect vulgar data, then the target category is vulgar data. If the task is to detect abnormal social group data, then the target category is abnormal social group data. And if the task is to detect harmful data, then the target category is harmful data.

[0092] Based on this, based on each target task processing unit, the modal features are processed separately to obtain the multi-task processing results of the target interaction data, including: for the target task processing unit used for data classification, the modal features are classified to obtain the classification results for the target interaction data; the multi-task processing results include at least the classification results.

[0093] The classification result is determined based on the classification category required by the classification task, which means it needs to be determined based on the classification head in the target task processing unit. Therefore, the multi-task processing result includes at least the classification result. Secondly, the classification head refers to the Multilayer Perceptron (MLP) in the neural network. Since the feature extraction network remains unchanged when processing different tasks in different business scenarios, the MLP parameters need to be changed based on different task processing requirements, resulting in different classification heads. Therefore, the classification task in this application also needs to divide the CLIP based on different target categories. For example, the MLP parameters trained using vulgar data samples are the vulgar classification head, and the MLP parameters trained using abnormal social group data samples are the abnormal social group classification head. Compared to CLIP parameters, the classification head (i.e., MLP parameters) has a smaller number of parameters and is more like a convenient classification plugin, capable of more flexibly adapting to classification tasks.

[0094] Specifically, the server, using the target task processing unit for data classification, classifies the modal features to obtain the classification result for the target interaction data. That is, the server first needs to determine the target category indicated by the classification task, and then, through the target task processing unit for data classification, classifies the modal features according to the target category to output the probability that the data category of the target interaction data belongs to the target category, thus obtaining the classification result for the target interaction data. The classification result is either: the data category of the target interaction data belongs to the target category, or the data category of the target interaction data does not belong to the target category.

[0095] In other words, the target task processing unit, used for data classification, outputs the probability that the data category of the target interactive data belongs to the target category. This probability is then compared with a preset probability threshold. If the probability of belonging to the target category is less than the preset probability threshold, the classification result for the target interactive data is determined to be: the data category of the target interactive data does not belong to the target category. Conversely, if the probability of belonging to the target category reaches the preset probability threshold, the classification result for the target interactive data is determined to be: the data category of the target interactive data belongs to the target category.

[0096] To facilitate understanding of the classification task, we will use the target interaction data, which consists of text and images, as an example for illustration. Figure 5 As shown, since the target interactive data is text and image data, a text encoder extracts features from the text data to obtain text features, and an image encoder extracts features from the image data to obtain image features. This yields the modal features of the target interactive data, including both text and image features. The classification head used needs to be determined based on the target category indicated by the classification task. For example, if the target category indicated by the classification task is "vulgar," the modal features need to be classified using the corresponding classification head 1 to obtain the classification result for the target interactive data. The classification result can be either: the target interactive data is of the "vulgar" type, or the target interactive data is not of the "vulgar" type.

[0097] Similarly, if the target category indicated by the classification task is abnormal social groups, then the modal features need to be classified using the classification head 2 corresponding to abnormal social groups to obtain the classification result for the target interaction data. The classification result can be: the target interaction data is an abnormal social group type, or the target interaction data is not an abnormal social group type. If the target category indicated by the classification task is bad, then the modal features need to be classified using the classification head 3 corresponding to bad to obtain the classification result for the target interaction data. The classification result can be: the target interaction data is a bad type, or the target interaction data is not a bad type.

[0098] In an optional embodiment, when the processing task includes a retrieval task, the target task processing unit is used for data retrieval, and the data retrieval is used to retrieve retrieval interaction data whose feature similarity satisfies the retrieval conditions.

[0099] Based on this, based on each target task processing unit, the modal features are processed separately to obtain the multi-task processing results of the target interaction data, including: for the target task processing unit used for data retrieval, determining the feature similarity between the modal features and the data features of each candidate interaction data, and selecting the retrieval interaction data from multiple candidate interaction data whose feature similarity with the target interaction data meets the retrieval conditions; the multi-task processing results include at least the retrieval interaction data.

[0100] In this context, feature similarity refers to the similarity between modal features and data features. Since modal features can be text features, image features, or a combination of text and image features, the data features for candidate interactive data can also be text features, image features, or a combination of text and image features. Therefore, feature similarity can be the similarity between features of the same data modality or the similarity between features of different data modalities. Specifically, feature similarity can be any of the following: text feature similarity, image feature similarity, and text-image feature similarity. Text feature similarity is the similarity between text features, image feature similarity is the similarity between image features, and text-image feature similarity is the similarity between text features and images. Secondly, the retrieval criteria can be that the feature similarity reaches a similarity threshold, or that a predetermined number of similarities with the highest numerical values ​​are selected from multiple feature similarities. Therefore, the retrieved interactive data can be a single interactive data set, or it can consist of multiple interactive data sets, and the multi-task processing results must at least include the obtained retrieval interactive data.

[0101] Specifically, the server also uses a feature extraction network from the content processing model to extract features for each candidate interaction data, thereby obtaining separate data features for each candidate interaction data. Based on this, for the target task processing unit used for data retrieval, the feature similarity between the modal features and each data feature is determined. Then, data retrieval is performed based on feature similarity; that is, from multiple candidate interaction data, the server selects the retrieval interaction data whose feature similarity to the target interaction data meets the retrieval criteria. In other words, the server selects candidate interaction data whose feature similarity to the target interaction data reaches a similarity threshold as the retrieval interaction data. Alternatively, the server sorts the feature similarity to the target interaction data, then selects a preset number of feature similarities from largest to smallest, and uses the candidate interaction data corresponding to the preset number of feature similarities as the retrieval interaction data.

[0102] For example, given candidate interaction data E1, E2, E3, E4, and E5, if the feature similarity between the modal feature and the data features of candidate interaction data E1 is 50%, the feature similarity between the modal feature and the data features of candidate interaction data E2 is 30%, the feature similarity between the modal feature and the data features of candidate interaction data E3 is 72%, the feature similarity between the modal feature and the data features of candidate interaction data E4 is 88%, and the feature similarity between the modal feature and the data features of candidate interaction data E5 is 76%, and taking a feature similarity threshold of 60% as an example, feature similarities greater than 60% (88%, 76%, and 72%) can be selected. This allows us to determine the candidate interaction data E3, candidate interaction data E4, and candidate interaction data E5 as the target interaction data.

[0103] Taking the retrieval condition of selecting the highest similarity value from multiple feature similarities as an example, for the aforementioned multiple feature similarities, the highest similarity value is the one with the largest similarity value. Therefore, it can be determined that the feature similarity is 88%. The feature similarity of 88% is the feature similarity between the modal feature and the data feature of the candidate interaction data E5. Therefore, it can be determined that the retrieved interaction data is the candidate interaction data E5.

[0104] Furthermore, as described above, the similarity between text features and images (i.e., image-text feature similarity) can also be calculated and determined. The following details how to calculate this image-text feature similarity: In a specific embodiment, for a target task processing unit used for data retrieval, the feature similarity between modal features and data features of each candidate interactive data is determined, including: when the modal features include text features, determining the image-text feature similarity between text features and data features including image features; when the modal features include image features, determining the image-text feature similarity between image features and data features including text features; the feature similarity includes at least image-text feature similarity.

[0105] Specifically, when the target interactive data is text data or image-text data, the extracted modal features of the target interactive data include text features. That is, if the target interactive data is text data, the modal features are text features; if the target interactive data is image-text data, the modal features are both text features and image data. Therefore, when the modal features include text features, the similarity between the text features in the modal features and the data features that include image features is determined. The aforementioned data features that include image features can be: data features that only include image features, or data features that include both image features and text features.

[0106] The following is a description of the aforementioned situations:

[0107] 1. When the modal feature is a text feature. If the data feature is an image feature, then the image-text feature similarity between the modal feature (text feature) and the data feature (image feature) can be obtained; that is, the feature similarity is the obtained image-text feature similarity. Secondly, if the data feature is both an image feature and a text feature, then the image-text feature similarity between the modal feature (text feature) and the image feature in the data feature can be obtained, as well as the text feature similarity between the modal feature (text feature) and the text feature in the data feature. In this case, the feature similarity includes both image-text feature similarity and text feature similarity.

[0108] 2. When the modal features are text features and image features. If the data features are image features, then we can obtain the image-text feature similarity between the text features in the modal features and the data features (image features), as well as the image feature similarity between the image features in the modal features and the data features (image features). In this case, the feature similarity includes both image-text feature similarity and image feature similarity. Secondly, if the data features are both image features and text features, then we can obtain the text feature similarity between the text features in the modal features and the text features in the data, as well as the image feature similarity between the text features in the modal features and the image features in the data. In this case, the feature similarity includes image-text feature similarity, text feature similarity, and image feature similarity.

[0109] Similarly, when the target interaction data is image data or image-text data, the extracted modal features of the target interaction data include image features. That is, if the target interaction data is image data, the modal features are image features; if the target interaction data is image-text data, the modal features are both image features and text data. Therefore, when the modal features include image features, the similarity between the image features in the modal features and the data features that include text features is determined. The aforementioned data features that include text features can be: data features that only include text features, or data features that include both image features and text features.

[0110] The following is a description of the aforementioned situations:

[0111] 1. When the modal feature is an image feature. If the data feature is a text feature, then the image-text feature similarity between the modal feature (image feature) and the data feature (text feature) can be obtained; that is, the feature similarity is the obtained image-text feature similarity. Secondly, if the data feature consists of both image and text features, then the image feature similarity between the modal feature (image feature) and the image features within the data feature can be obtained, as well as the image-text feature similarity between the modal feature (image feature) and the text features within the data feature. In this case, the feature similarity includes both image feature similarity and image-text feature similarity.

[0112] 2. When the modal features are text features and image features. If the data features are image features, then we can obtain the image-text feature similarity between the text features in the modal features and the data features (image features), as well as the image feature similarity between the image features in the modal features and the data features (image features). In this case, the feature similarity includes both image-text feature similarity and image feature similarity. Secondly, if the data features are both image features and text features, then we can obtain the text feature similarity between the text features in the modal features and the text features in the data, as well as the image feature similarity between the text features in the modal features and the image features in the data. In this case, the feature similarity includes image-text feature similarity, text feature similarity, and image feature similarity.

[0113] To facilitate understanding of the aforementioned feature similarity, let's take image and text data as the target interaction data as an example for illustration. Figure 6 As shown, the modal features of the target interactive data include image feature 601 and text feature 602, while the data features of the candidate interactive data include text feature 603 and image feature 604. Therefore, the feature similarity between text feature 602 and text feature 603 is the text feature similarity between the target interactive data and the candidate interactive data. Similarly, the feature similarity between image feature 601 and image feature 604 is the image feature similarity between the target interactive data and the candidate interactive data. Likewise, the feature similarity between image feature 601 and text feature 603 constitutes the image-text feature similarity between the target interactive data and the candidate interactive data, and the feature similarity between text feature 602 and image feature 604 constitutes the image-text feature similarity between the target interactive data and the candidate interactive data.

[0114] In an optional embodiment, when the processing task includes an object detection task, the object task processing unit is used for object detection, that is, the object task processing unit can detect whether a target object exists from the interactive data.

[0115] Based on this, each target task processing unit performs task processing on the modal features to obtain the multi-task processing result of the target interaction data, including: determining the target parameters that match the target detection task; the target parameters are used to indicate the target object; for the target task processing unit used for target detection, the target object is detected based on the modal features and the target parameters to obtain the target detection result of the target interaction data; the multi-task processing result includes at least the target detection result.

[0116] The target parameter indicates the target object, which is the object to be detected. Specifically, the target object describes an element object of a target category, such as a vulgar or inappropriate element object. Secondly, the target detection result characterizes whether the target interaction data contains the target object, and if so, indicates the target object's location within the target interaction data. Therefore, the multi-task processing result includes at least the target detection result.

[0117] Specifically, since the target task processing unit can detect the existence of target objects in the interaction data, the server first needs to determine the target parameters that match the target detection task and indicate the target objects. That is, the target parameters determine the element objects of the target category to be detected. Based on this, the server, for the target task processing unit used for target detection, detects the target objects indicated by the target parameters based on modal features and the target parameters, and then obtains the target detection result for the target interaction data of the target object. If there are no objects of the target category in the target interaction data, the target detection result indicates that there are no target objects in the target interaction data. If there are objects of the target category in the target interaction data, the target detection result is used to indicate the location information of the target object in the target interaction data, that is, to determine the detection box where the target object is located in the target interaction data, so as to determine that the target object belongs to the target category.

[0118] To facilitate understanding of the aforementioned object detection task, we will use image data as the target interaction data as an example for explanation. Figure 7 As shown, feature extraction is performed on the target interaction data to obtain modal features including image features 701. Then, based on the modal features of image features 701 and target parameters 702 used to indicate the target object, the target task processing unit 703 for target detection outputs the target detection result 704 for the target interaction data.

[0119] It should be understood that the corresponding examples in the embodiments of this application are used to understand this solution, but should not be construed as specific limitations on this solution.

[0120] In the above-mentioned multi-task processing method, in an interactive scenario that supports multimodal interaction, in response to a multi-task processing request, the target interactive data with at least one data modality indicated by the multi-task processing request is determined. Then, based on the feature extraction network in the content processing model, the modal features of the target interactive data are extracted according to the target data modality corresponding to the target interactive data. That is, whether it is single-modal or multimodal data, the feature extraction network can be used to extract the corresponding modality features to ensure the reliability of feature extraction. The multiple processing tasks carried in the multi-task processing request are then extracted. Target task processing units matching each processing task are determined from the content processing model. This eliminates the need for repetitive feature extraction for different processing tasks. Instead, each target task processing unit processes the modal features separately to obtain the multi-task processing results of the target interactive data. Therefore, the content processing model shares the feature extraction results for the data by fixing the feature extraction parameters. Since the content processing model includes multiple task processing units trained using the same interactive data samples, it accelerates the task prediction speed of the content processing model for different tasks. While saving training costs, it can quickly extract features for both single-modal and multi-modal data and provide timely response processing for multiple tasks, thereby improving the efficiency of task processing when multiple tasks exist.

[0121] As described in the foregoing embodiments, the content processing model includes a task processing unit for data classification, a target task processing unit for data retrieval, and a target task processing unit for object detection. Since the similarity retrieval task primarily considers practical applications, in the initial stages of finding anomalous data, a small number of anomalous data samples may result in low accuracy of the training results. In this case, similarity retrieval can be performed on a small number of anomalous data samples to recall similar interactive data, thereby preventing the widespread propagation of similar anomalous data and providing more reliable sample data for model training. Therefore, as described above, the similarity-based data retrieval process provided in this application can calculate not only the similarity between text and text, and between images, but also the similarity between text and images, thereby increasing the richness of anomalous data recall. Based on this, the target task processing unit for data retrieval can be pre-trained based on similarity calculations using interactive data samples, which will not be described in detail here. The following sections describe the methods for obtaining the task processing unit for data classification and the target task processing unit for object detection:

[0122] First, we will introduce how to obtain the task processing unit used for data classification. In one embodiment, such as... Figure 8 As shown, the methods for obtaining task processing units used for data classification in the content processing model include:

[0123] Step 802: Obtain interactive data samples and extract sample features of the interactive data samples based on the feature extraction network; the feature extraction parameters of the feature extraction network are fixed parameters.

[0124] The interactive data samples possess at least one data modality, meaning they are similar to the target interactive data in the aforementioned embodiments. They can be single-modality interactive data, such as image or text interactive data, or multi-modality interactive data, such as a data sample composed of image and text interactive data. Secondly, the feature extraction parameters of the feature extraction network are fixed parameters, meaning they are not adjusted based on model training. Furthermore, as described in the aforementioned embodiments, the feature extraction network is specifically the Encoder in CLIP, and the Encoder is specifically divided into a Text Encoder and an Image Encoder.

[0125] Specifically, the server first acquires interaction data samples, including both single-modal and multi-modal interaction data. The server can retrieve historical interaction data sent by various objects in multi-modal interaction scenarios based on its integrated data storage system, and then filter the historical interaction data to obtain interaction data samples. Considering that in practical applications, a small number of abnormal data samples may lead to low accuracy in training results in the initial stages of finding abnormal data, a similarity retrieval task can be used to perform similarity retrieval on the abnormal data samples. This retrieves the interaction data whose feature similarity to the abnormal data samples meets the retrieval criteria and adds it to the historical interaction data, which is then further filtered to obtain interaction data samples. Secondly, interaction data samples can also be data directly input by the person triggering model training; therefore, this application does not limit the method of acquiring interaction data samples.

[0126] Furthermore, the server extracts sample features from the interactive data samples based on a feature extraction network. Specifically, the server extracts text features from the text data in the interactive data samples using a text encoder in the feature extraction network, and extracts image features from the image data in the interactive data samples using an image encoder in the feature extraction network. Therefore, when the interactive data sample is single-modal image interactive data, the sample features are image features. When the interactive data sample is single-modal text interactive data, the sample features are text features. And when the interactive data sample is multi-modal image-text interactive data, the sample features consist of both image features and text features. The aforementioned data modalities and feature composition are similar to the modal features described in the previous embodiments, and will not be repeated here.

[0127] Step 804: Based on the sample features, classification prediction is performed through the initial task processing unit to obtain the classification prediction results of the interactive data samples.

[0128] The interactive data samples have sample labels, which correspond to the classification heads described in the aforementioned embodiments. These labels characterize the target category. In other words, the sample label representing the interactive data sample indicates a different data category depending on the classification head. For example, in a scenario requiring the identification of vulgar data, the data category corresponding to the classification head is vulgar data, and the sample label representing the interactive data sample indicates whether the sample is vulgar data. Similarly, in a scenario requiring the identification of inappropriate data, the data category corresponding to the classification head is inappropriate data, and the sample label representing the interactive data sample indicates whether the sample is inappropriate data.

[0129] Specifically, based on sample features, the server uses the initial task processing unit to perform classification prediction on the target category represented by the sample label, obtaining the classification prediction result of the interactive data sample for the target category. In other words, the classification prediction result indicates whether the interactive data sample belongs to the target category, or whether the interactive data sample does not belong to the target category.

[0130] For the initial task processing unit, the output is the probability that the interactive data sample belongs to the target category. Therefore, the obtained classification prediction result is actually a classification probability result, which will be described in detail below: In a specific embodiment, the sample label of the interactive data sample is used to represent the target category. And the target category is consistent with the data category corresponding to the classification head.

[0131] Based on this, classification prediction is performed through the initial task processing unit based on sample features to obtain the classification prediction result of the interactive data sample, including: classification prediction is performed through the initial task processing unit based on sample features to obtain the classification probability result of the interactive data sample belonging to the target category; the classification prediction result is the classification probability result.

[0132] The classification probability result describes the probability that the data category of the interactive data sample belongs to the target category. Specifically, based on the sample features, the server performs classification prediction on the target category represented by the sample label through the initial task processing unit, obtaining the classification probability result of the interactive data sample belonging to the target category. That is, the classification probability result of the interactive data sample belonging to the target category is compared with a preset probability threshold. If the probability that the data category of the interactive data sample belongs to the target category is less than the preset probability threshold, it is determined that the interactive data sample does not belong to the target category. Conversely, if the probability that the data category of the interactive data sample belongs to the target category reaches the preset probability threshold, it is determined that the interactive data sample belongs to the target category.

[0133] As described above, feature extraction can be performed on image-text interaction data to obtain image features and text features. Considering the generalization ability of the model, a random mask can be added to the features corresponding to the image-text interaction data to replace the original image features or text features, thereby increasing the generalization ability of the model. This will be described in detail below: In an optional embodiment, the interaction data sample is image-text interaction data; the sample features include image features and text features.

[0134] In other words, for the case where the interactive data sample is text-image interactive data, the server extracts the text features of the text data in the text-image interactive data through the text encoder in the feature extraction network, and extracts the image features of the image data in the text-image interactive data through the image encoder in the feature extraction network. Therefore, the obtained sample features include both image features and text features.

[0135] Based on this, classification prediction is performed through the initial task processing unit based on the sample features to obtain the classification probability result of the interactive data sample belonging to the target category. This includes: performing random masking processing on image features and text features to obtain masked sample features; the masked sample features are: text masked sample features, or image masked sample features; based on the masked sample features, classification prediction is performed through the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category.

[0136] In this process, the sample features of each interactive data sample can only be masked once; that is, masking is performed on the sample features of a single interactive data sample. Specifically, the server performs random masking on either the image features or the text features in the sample features to obtain the masked sample features. Therefore, if the masking process specifically masks the text features, then the masked sample features are: text-masked sample features. Conversely, if the masking process specifically masks the image features, then the masked sample features are: image-masked sample features. Based on this, the server then uses a similar method to perform classification prediction on the target category represented by the sample label through the initial task processing unit, obtaining the classification probability result of the interactive data sample belonging to the target category.

[0137] In one optional embodiment, the interaction data sample is image interaction data or text interaction data; the sample features are image features corresponding to the image interaction data, or text features corresponding to the text interaction data.

[0138] In other words, for interactive data samples that are either image-based or text-based, the text encoder in the feature extraction network extracts the text features of the interactive data samples (i.e., text-based interactive data) as text data. Therefore, the sample features are the text features corresponding to the text-based interactive data. Similarly, the image encoder in the feature extraction network extracts the image features of interactive data samples (i.e., image-based interactive data) as image data. Therefore, the sample features are the image features corresponding to the image-based interactive data.

[0139] Based on this, classification prediction is performed through the initial task processing unit based on the sample features to obtain the classification probability result of the interactive data sample belonging to the target category. This includes: when the sample features are image features corresponding to image interactive data, the sample features are determined as sample features after text masking; when the sample features are text features corresponding to text interactive data, the sample features are determined as sample features after image masking; and based on the masked sample features, classification prediction is performed through the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category.

[0140] The masked sample features are either text-masked sample features or image-masked sample features. Specifically, when the sample features are image features corresponding to the image interaction data, i.e., when the interaction data sample is image interaction data, the interaction data sample does not include text data. In other words, the sample features of the interaction data sample do not include text features. Since they do not include text features, they can be considered as text features being masked. Therefore, the sample features that only include image features are determined as text-masked sample features. That is, when the sample features are image interaction data, the masked sample features are text-masked sample features.

[0141] Similarly, when the sample features are text features corresponding to the text interaction data, i.e., when the interaction data sample is text interaction data, the interaction data sample does not include image data. In other words, the sample features of the interaction data sample do not include image features. Since image features are not included, they can be considered as masked image features. Therefore, the sample features that only include text features are determined as the image-masked sample features. That is, for the case where the sample features are text interaction data, the masked sample features are the image-masked sample features.

[0142] Based on this, the server then uses a similar method to the above, based on the masked sample features, and through the initial task processing unit, performs classification prediction on the target category represented by the sample label to obtain the classification probability result of the interactive data sample belonging to the target category.

[0143] To facilitate understanding of the two masking methods mentioned above, such as Figure 9 As shown, the training dataset includes multiple interactive data samples, including image interactive data, image-text interactive data, and text interactive data. Feature extraction is performed on the image interactive data to obtain image features, which are then used as sample features after text masking. Similarly, feature extraction is performed on the text interactive data to obtain text features, which are then used as sample features after image masking. Furthermore, feature extraction is performed on the image-text interactive data to obtain both image and text features, which are then randomly masked to obtain masked sample features. These masked sample features can be either text-masked or image-masked.

[0144] Step 806: Based on the sample labels of the interactive data samples and the classification prediction results, adjust the task processing parameters of the initial task processing unit to obtain a task processing unit for data classification.

[0145] Specifically, the server adjusts the task processing parameters of the initial task processing unit based on the sample labels and classification prediction results of the interactive data samples to obtain a task processing unit for data classification. That is, the server calculates the binary classification model loss for the target category represented by the sample labels and the result belonging to the target category represented by the classification prediction results. Then, it fixes the feature extraction parameters of the feature extraction network and adjusts the task processing parameters of the initial task processing unit using the binary classification model loss, thus obtaining the task processing unit for data classification. In other words, after calculating the binary classification model loss, the server determines whether the loss function of the initial task processing unit has reached the convergence condition. If it has not reached the convergence condition, the server adjusts the task processing parameters of the initial task processing unit using the binary classification model loss. When the loss function of the initial task processing unit reaches the convergence condition, the task processing unit for data classification is obtained based on the task processing parameters saved after the last adjustment.

[0146] The convergence condition for the aforementioned loss function can be that the loss value is less than or equal to a first preset threshold. For example, the first preset threshold can be 0.005, 0.01, 0.02, or other values ​​close to 0. Alternatively, the convergence condition can be that the difference between two consecutive loss values ​​is less than or equal to a second preset threshold. The second threshold can be the same as or different from the first threshold. For example, the second preset threshold can be 0.005, 0.01, 0.02, or other values ​​close to 0. Another convergence condition can be that the number of updates to the task processing parameters of the initial task processing unit reaches an update iteration threshold. In practical applications, other convergence conditions can also be used, which are not limited here.

[0147] In one specific embodiment, the task processing parameters of the initial task processing unit are adjusted based on the sample labels and classification prediction results of the interactive data samples to obtain a task processing unit for data classification. This includes: adjusting the task processing parameters of the initial task processing unit based on the target category and classification probability results to obtain a task processing unit for data classification.

[0148] Specifically, the server calculates the binary classification model loss based on the target category represented by the sample label and the probability of belonging to the target category represented by the classification probability result. Then, the feature extraction parameters of the feature extraction network are fixed, and the task processing parameters of the initial task processing unit are adjusted using the binary classification model loss to obtain the task processing unit used for data classification. The specific parameter adjustment method is similar to that in the aforementioned embodiments and will not be repeated here.

[0149] It should be understood that the corresponding examples in the embodiments of this application are used to understand this solution, but should not be construed as specific limitations on this solution.

[0150] In this embodiment, by using training data with both unimodal and multimodal data, and masking features to enhance the model's generalization ability, different classification heads are added and trained according to different classification purposes after obtaining the masked features. This allows for data classification of different categories of images, text, and image-text pairs, thereby improving the generalization ability and accuracy of data classification. Furthermore, by employing a feature extraction and classification head separation approach, different classification requirements can be met during training by focusing only on the classification head, saving training resources and improving training efficiency, thus also enhancing the efficiency of data classification.

[0151] In one embodiment, such as Figure 10 As shown, the methods for obtaining task processing units used for object detection in the content processing model include:

[0152] Step 1002: Obtain interactive data samples and object annotation results for the interactive data samples; the interactive data samples are image interactive data; the object annotation results are used to characterize the annotated objects present in the interactive data samples.

[0153] The interactive data samples are image interactive data. The object annotation results are used to characterize the annotated objects existing in the interactive data samples, and the annotated objects have corresponding annotation categories. For example, interactive data sample F1 contains annotated objects G1, G2, and G3. Annotated object G1 corresponds to annotation category H1, annotated object G2 corresponds to annotation category H2, and annotated object G3 corresponds to annotation category H3.

[0154] In one optional embodiment, the object annotation result specifically indicates the annotation detection box of the annotated object in the interactive data sample, as well as the annotation category to which the annotation detection box belongs. The annotation category to which the annotation detection box belongs is the annotation category of the annotated object, and the annotation detection box is used to indicate the location information of the annotated object in the interactive data sample. For example, if there is an annotated object G1 in the interactive data sample F1, and the annotated object G1 is located in the annotation detection box I1 in the interactive data sample F1, then the annotation category to which the annotation detection box I1 belongs is the annotation category H1.

[0155] Specifically, the server obtains interactive data samples and object annotation results for the interactive data samples. The method by which the server obtains interactive data samples is similar to that in the previous embodiments, and will not be described again here. The object annotation results for the interactive data samples are obtained by annotators annotating the interactive data samples. Therefore, the method by which the object annotation results for the interactive data samples are obtained is similar to that in the previous embodiments, and will not be described again here.

[0156] Step 1004: Extract sample features from interactive data samples based on the feature extraction network; the feature extraction parameters of the feature extraction network are fixed parameters.

[0157] Secondly, the feature extraction parameters of the feature extraction network are fixed parameters, meaning they are not adjusted based on model training. As described in the preceding embodiments, the feature extraction network is specifically the Encoder in CLIP, and the Encoder is further divided into a Text Encoder and an Image Encoder. Specifically, the server extracts sample features from the interactive data samples based on the feature extraction network. Since the interactive data samples are image interactive data, the server extracts sample features from the interactive data samples based on the Image Encoder in the feature extraction network, and these sample features are specifically image features.

[0158] Step 1006: For the initial task processing unit, based on the sample features and the learnable query parameters, perform object detection on the query object represented by the learnable query parameters to obtain the object detection result of the query object in the interactive data sample.

[0159] Specifically, learnable query parameters are initialized and learnable query parameters used to characterize the query object. Specifically, for the initial task processing unit, the server first determines the query object requiring object detection using the learnable query parameters, and then performs object detection on the query object based on sample features, obtaining the object detection result of the query object in the interaction data sample. The object detection result can indicate whether the query object exists in the interaction data sample or not. Secondly, if the query object is not considered, object detection based on sample features can obtain the object detection result in the interaction data sample. The aforementioned object detection result describes the detected object present in the interaction data sample.

[0160] As described above, object annotation results can specifically indicate the bounding box containing the annotated object in the interactive data sample, as well as the annotation category to which the bounding box belongs. Therefore, when object annotation results specifically indicate the bounding box containing the annotated object in the interactive data sample, and the annotation category to which the bounding box belongs, object detection not only needs to detect the existence of the query object, but also needs to detect the bounding box containing the query object. This will be explained in detail below:

[0161] In one specific embodiment, for the initial task processing unit, object detection is performed on the query object represented by the learnable query parameters based on sample features and learnable query parameters to obtain the object detection result of the query object in the interactive data sample, including: for the initial task processing unit, object detection is performed on the query object represented by the learnable query parameters based on sample features and learnable query parameters to obtain the predicted detection box in the interactive data sample and the predicted category to which the predicted detection box belongs.

[0162] Specifically, for the initial task processing unit, the server first determines the query object to be detected using learnable query parameters. Then, based on sample features, it performs object detection on the query object, obtaining the predicted detection box of the query object in the interaction data sample and the predicted category of the query object within the predicted detection box. If there is no predicted detection box for the query object in the interaction data sample, it means that the query object does not exist in the interaction data sample, resulting in an empty predicted detection box and an empty predicted category. If there is a predicted detection box for the query object in the interaction data sample, the predicted category of the predicted detection box is also obtained, which is the predicted category of the query object within the predicted detection box.

[0163] Step 1008: For the initial task processing unit, based on the object annotation results and object detection results, adjust the task processing parameters and learnable query parameters of the initial task processing unit to obtain a task processing unit for object detection.

[0164] Specifically, the server adjusts the task processing parameters and learnable query parameters of the initial task processing unit based on the object annotation and object detection results to obtain a task processing unit for object detection. That is, the server calculates the model loss based on the object annotation and object detection results, then fixes the feature extraction parameters of the feature extraction network, and adjusts the task processing parameters and learnable query parameters of the initial task processing unit using the model loss to obtain the task processing unit for object detection. In other words, after calculating the model loss, the server determines whether the loss function of the initial task processing unit has reached the convergence condition. If it has not reached the convergence condition, the server adjusts the task processing parameters and learnable query parameters of the initial task processing unit using the model loss. When the loss function of the initial task processing unit reaches the convergence condition, the server obtains the task processing unit for object detection based on the saved task processing parameters and learnable query parameters after the last adjustment. The specific method for determining whether the convergence condition has been reached is similar to the previous embodiment and will not be repeated here.

[0165] As described above, the object annotation results can specifically indicate the labeled detection boxes and the annotation categories to which the labeled detection boxes belong in the interactive data samples, and can also obtain the predicted detection boxes and the predicted categories to which the predicted detection boxes belong in the interactive data samples. Therefore, the parameter adjustment method specifically considers the aforementioned detection boxes and categories: In a specific embodiment, for the initial task processing unit, based on the object annotation results and object detection results, the task processing parameters and learnable query parameters of the initial task processing unit are adjusted to obtain a task processing unit for object detection, including: for the initial task processing unit, based on the labeled detection boxes and predicted detection boxes, and the labeled categories and predicted categories, the task processing parameters and learnable query parameters of the initial task processing unit are adjusted to obtain a task processing unit for object detection.

[0166] Specifically, for the initial task processing unit, the server calculates the detection box loss based on the labeled detection box and the predicted detection box, and calculates the category loss based on the labeled category and the predicted category. Thus, the model loss can be obtained through the detection box loss and the category loss. Then, the feature extraction parameters of the feature extraction network are fixed, and the model loss, the task processing parameters of the initial task processing unit, and the learnable query parameters are adjusted to obtain the task processing unit for object detection.

[0167] For ease of understanding, such as Figure 11 As shown, the image encoder in CLIP extracts sample features 1102 (i.e., image features) from interactive data sample 1101. Then, sample features 1102 and learnable query parameters 1103 are used as inputs to the initial task processing unit. The initial task processing unit is specifically an encoder-decoder architecture for object detection. Sample features 1102 are specifically the input of the encoder in the initial task processing unit, while learnable query parameters 1103 are the input of the decoder in the initial task processing unit. Finally, the initial task processing unit outputs the object detection result 1104 for the interactive data sample, i.e., the predicted detection box of the query object in the interactive data sample and the predicted category to which the predicted detection box belongs.

[0168] It should be understood that the corresponding examples in the embodiments of this application are used to understand this solution, but should not be construed as specific limitations on this solution.

[0169] In this embodiment, feature extraction and target detection units are separated, and target detection is performed through learnable query parameters. Since the parameters corresponding to feature extraction do not need to be trained and can be reused directly, only the downstream task processing parameters and learnable query parameters need to be trained and deployed to improve training speed and inference efficiency, and also improve the efficiency of target retrieval.

[0170] Based on the detailed description of the foregoing embodiments, the complete flow of the multi-task processing method in the embodiments of this application will be described below. In one embodiment, such as... Figure 12 As shown, a multi-task processing method is provided, which is applied to... Figure 1 Taking server 104 as an example, it can be understood that this method can also be applied to terminal 102, and also to a system including terminal 102 and server 104, and is implemented through the interaction between terminal 102 and server 104. In this embodiment, the method includes the following steps:

[0171] Step 1201: In an interactive scenario that supports multimodal interaction, in response to a multi-task processing request, determine the target interactive data indicated by the multi-task processing request.

[0172] Step 1202: Based on the feature extraction network in the content processing model, extract the modal features of the target interaction data according to the target data modality corresponding to the target interaction data.

[0173] Step 1203: Extract the multiple processing tasks carried in the multi-task processing request. The multiple processing tasks include classification tasks, retrieval tasks, and object detection tasks.

[0174] Step 1204: For the classification task, determine the target task processing unit for data classification from the content processing model.

[0175] Step 1205: For the target task processing unit used for data classification, classify the modal features to obtain the classification results for the target interaction data.

[0176] Step 1206: For the retrieval task, determine the target task processing unit for data retrieval from the content processing model.

[0177] Step 1207: When the modal features include text features, for the target task processing unit used for data retrieval, determine the image-text feature similarity between the text features and the data features including image features; when the modal features include image features, for the target task processing unit used for data retrieval, determine the image-text feature similarity between the image features and the data features including text features; the feature similarity includes at least image-text feature similarity.

[0178] Step 1208: Select retrieval interaction data from multiple candidate interaction data that meets the retrieval criteria in terms of feature similarity with the target interaction data.

[0179] Step 1209: For the object detection task, determine the target task processing unit for object detection from the content processing model.

[0180] Step 1210: Determine the target parameters that match the target detection task; the target parameters are used to indicate the target object; for the target task processing unit used for target detection, the target object is detected based on modal features and target parameters to obtain the target detection result of the target interaction data.

[0181] It should be understood that the specific implementation methods of steps 1201 to 1210 are similar to those of the aforementioned embodiments, and will not be repeated here.

[0182] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0183] Based on the same inventive concept, this application also provides a multitasking apparatus for implementing the multitasking method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more multitasking apparatus embodiments provided below can be found in the limitations of the multitasking method described above, and will not be repeated here.

[0184] In one embodiment, such as Figure 13 As shown, a multi-task processing device is provided, including: a data determination module 1302, a feature extraction module 1304, a processing unit determination module 1306, and a task processing module 1308, wherein:

[0185] The data determination module 1302 is used to determine the target interactive data indicated by the multi-task processing request in response to the multi-task processing request in an interactive scenario that supports multi-modal interaction, wherein the target interactive data has at least one data modality.

[0186] Feature extraction module 1304 is used to extract modal features of target interaction data according to the target data modality corresponding to the target interaction data based on the feature extraction network in the content processing model;

[0187] The processing unit determination module 1306 is used to extract multiple processing tasks carried in the multi-task processing request and determine the target task processing unit that matches each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples.

[0188] The task processing module 1306 is used to perform task processing on modal features based on each target task processing unit to obtain the multi-task processing results of the target interaction data.

[0189] In one embodiment, the plurality of processing tasks includes at least a classification task;

[0190] The processing unit determination module is specifically used to determine the target task processing unit for data classification from the content processing model for classification tasks.

[0191] The task processing module is specifically used to classify modal features for the target task processing unit used for data classification, and obtain the classification results for the target interactive data; the multi-task processing results include at least the classification results.

[0192] In one embodiment, the plurality of processing tasks includes at least a retrieval task;

[0193] The processing unit determination module is specifically used to determine the target task processing unit for data retrieval from the content processing model for the retrieval task.

[0194] The task processing module is specifically used to determine the feature similarity between the modal features and the data features of each candidate interaction data for the target task processing unit used for data retrieval, and to select the retrieval interaction data from multiple candidate interaction data whose feature similarity with the target interaction data meets the retrieval conditions; the multi-task processing result includes at least the retrieval interaction data.

[0195] In one embodiment, the task processing module is specifically configured to determine the image-text feature similarity between text features and data features including image features when the modal features include text features; and to determine the image-text feature similarity between image features and data features including text features when the modal features include image features; wherein the feature similarity includes at least image-text feature similarity.

[0196] In one embodiment, the plurality of processing tasks includes at least an object detection task;

[0197] The processing unit determination module is specifically used to determine the target task processing unit for target detection from the content processing model for the target detection task.

[0198] The task processing module is specifically used to determine the target parameters that match the target detection task; the target parameters are used to indicate the target object; for the target task processing unit used for target detection, the target object is detected based on modal features and target parameters to obtain the target detection result of the target interaction data; the multi-task processing result includes at least the target detection result.

[0199] In one embodiment, the multitasking device further includes a task processing unit acquisition module;

[0200] The task processing unit acquisition module is used to acquire interactive data samples and extract sample features of the interactive data samples based on the feature extraction network. The feature extraction parameters of the feature extraction network are fixed parameters. Based on the sample features, classification prediction is performed through the initial task processing unit to obtain the classification prediction result of the interactive data samples. Based on the sample labels of the interactive data samples and the classification prediction result, the task processing parameters of the initial task processing unit are adjusted to obtain the task processing unit used for data classification.

[0201] In one embodiment, the sample label of the interactive data sample is used to characterize the target category;

[0202] The task processing unit acquisition module is specifically used to perform classification prediction based on sample features through the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category; the classification prediction result is the classification probability result; based on the target category and the classification probability result, the task processing parameters of the initial task processing unit are adjusted to obtain the task processing unit used for data classification.

[0203] In one embodiment, the interactive data sample is graphic-text interactive data; the sample features include image features and text features.

[0204] The task processing unit acquisition module is specifically used to perform random masking processing on image features and text features to obtain masked sample features. The masked sample features are either text masked sample features or image masked sample features. Based on the masked sample features, classification prediction is performed through the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category.

[0205] In one embodiment, the interaction data sample is image interaction data or text interaction data; the sample features are the image features corresponding to the image interaction data, or the text features corresponding to the text interaction data.

[0206] The task processing unit acquisition module is specifically used to determine the sample features as text-masked sample features when the sample features are image features corresponding to the image interaction data; and to determine the sample features as image-masked sample features when the sample features are text features corresponding to the text interaction data. Based on the masked sample features, the initial task processing unit performs classification prediction to obtain the classification probability result of the interaction data sample belonging to the target category. The masked sample features are: text-masked sample features, or image-masked sample features.

[0207] In one embodiment, the task processing unit acquisition module is further configured to acquire interactive data samples and object annotation results of the interactive data samples; the interactive data samples are image interactive data; the object annotation results are used to characterize the labeled objects present in the interactive data samples; sample features of the interactive data samples are extracted based on a feature extraction network; the feature extraction parameters of the feature extraction network are fixed parameters; for the initial task processing unit, object detection is performed on the query object characterized by the learnable query parameters based on the sample features and learnable query parameters to obtain the object detection result of the query object in the interactive data samples; for the initial task processing unit, the task processing parameters and learnable query parameters of the initial task processing unit are adjusted based on the object annotation results and object detection results to obtain a task processing unit for target detection.

[0208] In one embodiment, the object annotation result specifically indicates the annotation detection box of the annotated object in the interactive data sample, as well as the annotation category to which the annotation detection box belongs;

[0209] The task processing unit acquisition module is specifically used to perform object detection on the initial task processing unit based on sample features and learnable query parameters, obtaining the query object represented by the learnable query parameters, and obtaining the predicted detection box and the predicted category to which the predicted detection box belongs in the interactive data sample. Based on the labeled detection box and the predicted detection box, as well as the labeled category and the predicted category, the task processing parameters and learnable query parameters of the initial task processing unit are adjusted to obtain the task processing unit used for object detection.

[0210] In one embodiment, a computer device is provided, which can be a server or a terminal. This embodiment uses a server as an example for description, and its internal structure diagram is as follows: Figure 14As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores content processing models and interactive data related to the embodiments of this application. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a multitasking method.

[0211] Those skilled in the art will understand that Figure 14 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0212] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0213] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0214] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0215] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the object or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0216] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based multitasking logic devices, etc., and are not limited to these.

[0217] The technical features in the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0218] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for multitasking, characterized in that, The method includes: In an interactive scenario that supports multimodal interaction, in response to a multi-task processing request, the target interactive data indicated by the multi-task processing request is determined, wherein the target interactive data has at least one data modality; Based on the feature extraction network in the content processing model, modal features of the target interaction data are extracted according to the target data modality corresponding to the target interaction data. Extract multiple processing tasks carried in the multi-task processing request, and determine target task processing units that match each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples. Based on each of the target task processing units, the modal features are processed to obtain the multi-task processing results of the target interaction data.

2. The method according to claim 1, characterized in that, The plurality of processing tasks includes at least a classification task; The step of determining the target task processing unit that matches each of the processing tasks from the content processing model includes: For the classification task, a target task processing unit for data classification is determined from the content processing model; The process of performing task processing on the modal features based on each of the target task processing units to obtain the multi-task processing result of the target interaction data includes: For the target task processing unit used for data classification, the modal features are classified to obtain a classification result for the target interactive data; the multi-task processing result includes at least the classification result.

3. The method according to claim 1, characterized in that, The plurality of processing tasks includes at least a retrieval task; The step of determining the target task processing unit that matches each of the processing tasks from the content processing model includes: For the retrieval task, a target task processing unit for data retrieval is determined from the content processing model; The process of performing task processing on the modal features based on each of the target task processing units to obtain the multi-task processing result of the target interaction data includes: For the target task processing unit used for data retrieval, the feature similarity between the modal features and the data features of each candidate interaction data is determined, and retrieval interaction data whose feature similarity with the target interaction data meets the retrieval conditions is selected from multiple candidate interaction data; the multi-task processing result includes at least the retrieval interaction data.

4. The method according to claim 3, characterized in that, The step of determining the feature similarity between the modal features and the data features of each candidate interaction data for the target task processing unit used for data retrieval includes: In the case where the modal features include text features, determine the image-text feature similarity between the text features and data features that include image features; When the modal features include image features, the image-text feature similarity between the image features and data features including text features is determined; the feature similarity includes at least the image-text feature similarity.

5. The method according to claim 1, characterized in that, The plurality of processing tasks includes at least a target detection task; The step of determining the target task processing unit that matches each of the processing tasks from the content processing model includes: For the target detection task, a target task processing unit for target detection is determined from the content processing model; The process of performing task processing on the modal features based on each of the target task processing units to obtain the multi-task processing result of the target interaction data includes: Determine the target parameters that match the target detection task; the target parameters are used to indicate the target object. For the target task processing unit used for target detection, the target object is detected based on the modal features and the target parameters to obtain the target detection result of the target interaction data; the multi-task processing result includes at least the target detection result.

6. The method according to claim 1, characterized in that, The methods for obtaining the task processing unit used for data classification in the content processing model include: Acquire interactive data samples, and extract sample features of the interactive data samples based on the feature extraction network; the feature extraction parameters of the feature extraction network are fixed parameters; Based on the sample features, classification prediction is performed by the initial task processing unit to obtain the classification prediction result of the interactive data sample; Based on the sample labels of the interactive data samples and the classification prediction results, the task processing parameters of the initial task processing unit are adjusted to obtain the task processing unit used for data classification.

7. The method according to claim 6, characterized in that, The sample labels of the interactive data samples are used to characterize the target category; The process of performing classification prediction on the interactive data samples based on the sample features through the initial task processing unit to obtain the classification prediction results includes: Based on the sample features, a classification prediction is performed by the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category; the classification prediction result is the classification probability result. The task processing parameters of the initial task processing unit are adjusted based on the sample labels of the interactive data samples and the classification prediction results to obtain the task processing unit for data classification, including: Based on the target category and the classification probability result, the task processing parameters of the initial task processing unit are adjusted to obtain the task processing unit used for data classification.

8. The method according to claim 7, characterized in that, The interactive data sample is text-image interactive data; the sample features include image features and text features. The step of performing classification prediction based on the sample features through the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category includes: The image features and the text features are randomly masked to obtain masked sample features; the masked sample features are either text masked sample features or image masked sample features. Based on the masked sample features, the initial task processing unit performs classification prediction to obtain the classification probability result of the interactive data sample belonging to the target category.

9. The method according to claim 7, characterized in that, The interactive data sample is either image interactive data or text interactive data; the sample features are image features corresponding to image interactive data, or text features corresponding to text interactive data. The step of performing classification prediction based on the sample features through the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category includes: When the sample feature is an image feature corresponding to the image interaction data, the sample feature is determined as the sample feature after text masking; When the sample feature is a text feature corresponding to the text interaction data, the sample feature is determined to be the sample feature after image masking. Based on the masked sample features, classification prediction is performed by the initial task processing unit to obtain the classification probability result of the interactive data sample belonging to the target category; the masked sample features are: the sample features after text masking, or the sample features after image masking.

10. The method according to claim 1, characterized in that, The methods for obtaining the task processing unit used for object detection in the content processing model include: Obtain interactive data samples and object annotation results for the interactive data samples; the interactive data samples are image interactive data; the object annotation results are used to characterize the annotated objects present in the interactive data samples; The feature extraction network extracts sample features from the interactive data samples; the feature extraction parameters of the feature extraction network are fixed parameters. For the initial task processing unit, based on the sample features and the learnable query parameters, object detection is performed on the query object represented by the learnable query parameters to obtain the object detection result of the query object in the interactive data sample; For the initial task processing unit, based on the object annotation results and the object detection results, the task processing parameters and the learnable query parameters of the initial task processing unit are adjusted to obtain the task processing unit for object detection.

11. The method according to claim 10, characterized in that, The object annotation result specifically indicates the annotation detection box of the annotated object in the interactive data sample, and the annotation category to which the annotation detection box belongs; The step of the initial task processing unit performing object detection on the query object represented by the learnable query parameters based on the sample features and learnable query parameters, and obtaining the object detection result of the query object in the interaction data sample, includes: For the initial task processing unit, based on the sample features and the learnable query parameters, the query object represented by the learnable query parameters is obtained and object detection is performed to obtain the predicted detection box in the interactive data sample and the predicted category to which the predicted detection box belongs; The step involves adjusting the task processing parameters and learnable query parameters of the initial task processing unit based on the object annotation results and the object detection results to obtain the task processing unit for object detection, including: For the initial task processing unit, based on the labeled detection box and the predicted detection box, as well as the labeled category and the predicted category, the task processing parameters and the learnable query parameters of the initial task processing unit are adjusted to obtain the task processing unit for target detection.

12. A multitasking processing device, characterized in that, The device includes: A data determination module is used to determine, in response to a multi-task processing request, the target interactive data indicated by the multi-task processing request in an interactive scenario that supports multimodal interaction, wherein the target interactive data has at least one data modality. The feature extraction module is used to extract modal features of the target interaction data according to the target data modality corresponding to the target interaction data, based on the feature extraction network in the content processing model. The processing unit determination module is used to extract multiple processing tasks carried in the multi-task processing request and determine target task processing units that match each processing task from the content processing model; the content processing model includes multiple task processing units trained using the same interactive data samples. The task processing module is used to perform task processing on the modal features based on each of the target task processing units to obtain the multi-task processing result of the target interaction data.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.