Sample data processing method and device, task processing method and device, equipment and medium

By generating loss information from a large multimodal model to filter multimodal datasets, the problem of high data cleaning costs in training large multimodal models is solved, and the training efficiency and performance of lightweight target models are improved.

CN120744489APending Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510826202.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, data cleaning is costly and difficult to ensure in training large multimodal models, which affects the learning effect and generalization ability of lightweight target models, especially when noisy and complex samples in multimodal datasets are difficult to handle.

Method used

A multimodal large model is used to process candidate sample data, indicate the target task according to the prompt information, generate loss information, filter the candidate sample data based on the loss information, and select target sample data suitable for lightweight target model training.

Benefits of technology

By directly filtering data through a large multimodal model with enhanced learning capabilities, it provides a clean and consistent sample data distribution, improves the training efficiency and performance of lightweight target models, and is suitable for resource-constrained or rapidly iterating multimodal data processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744489A_ABST
    Figure CN120744489A_ABST
Patent Text Reader

Abstract

The invention provides a sample data processing method and device, a task processing method and device, equipment and a medium, relates to the technical field of artificial intelligence and big data, in particular to the technical field of computer vision, deep learning, large models and the like, and can be applied to task scenes of text processing, image processing, audio processing, video processing and the like. According to the specific implementation scheme, processing is conducted according to a target task indicated by prompt information on the basis of multiple pieces of candidate sample data through a multi-modal large model, and multiple pieces of loss information used for the multiple pieces of candidate sample data are obtained; based on the multiple pieces of loss information used for the multiple pieces of candidate sample data, the multiple pieces of candidate sample data are filtered, and at least one piece of target sample data used for training a task processing model is obtained; wherein the parameter quantity of the multi-modal large model is greater than the parameter quantity of the task processing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of artificial intelligence and big data technology, particularly computer vision, deep learning, and large models, and can be applied to tasks such as text processing, image processing, audio processing, and video processing. More specifically, this disclosure provides a sample data processing method, a model training method, a task processing method, an apparatus, an electronic device, and a storage medium. Background Art

[0002] In the real world, data and information are not monomodal. For example, videos contain images and audio, and social media posts may contain text and images. To better understand and process this multimodal data and information, researchers have begun exploring models that can simultaneously process multimodal data. However, in the application of multimodal large language models (MLLMs), the quality of training data has a crucial impact on model performance. High-performing models are key to ensuring processing accuracy and efficiency in individual scenarios. Summary of the Invention

[0003] The present disclosure provides a sample data processing method, a model training method, a task processing method, an apparatus, an electronic device, and a storage medium.

[0004] According to one aspect of the present disclosure, a sample data processing method is provided, including: utilizing a multimodal large model to process a plurality of candidate sample data according to a target task indicated by prompt information, and obtaining a plurality of loss information for the plurality of candidate sample data; filtering the plurality of candidate sample data based on the plurality of loss information for the plurality of candidate sample data, and obtaining at least one target sample data for training a task processing model; wherein the number of parameters of the multimodal large model is greater than the number of parameters of the task processing model.

[0005] According to another aspect of the present disclosure, a task processing model training method is provided, comprising: training a task processing model to be trained using at least one target sample data to obtain a target task processing model; wherein the at least one target sample data is obtained according to the sample data processing method of the present disclosure.

[0006] According to another aspect of the present disclosure, a task processing method is provided, including: obtaining data to be processed; inputting the data to be processed into a target task processing model, and outputting a task processing result; wherein the target task processing model is trained based on the training method provided by the present disclosure.

[0007] According to another aspect of the present disclosure, a task processing model training device is provided, including: a processing module for using a multimodal large model to process a plurality of candidate sample data according to a target task indicated by prompt information, and obtain a plurality of loss information for the plurality of candidate sample data; a filtering module for filtering the plurality of candidate sample data based on the plurality of loss information for the plurality of candidate sample data, and obtain at least one target sample data for training the task processing model; wherein the parameter amount of the multimodal large model is greater than the parameter amount of the task processing model.

[0008] According to another aspect of the present disclosure, a task processing device is provided, comprising: a task processing model training device, comprising: a training module, for training a task processing model to be trained using at least one target sample data to obtain a target task processing model; at least one target sample data is obtained according to the sample data processing method of the present disclosure.

[0009] According to another aspect of the present disclosure, a task processing device is provided, including: an acquisition module for acquiring data to be processed; an input-output module for inputting the data to be processed into a target task processing model and outputting a task processing result; wherein the target task processing model is trained based on the training method of the present disclosure or trained based on the training device of the present disclosure.

[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.

[0011] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.

[0012] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method provided according to the present disclosure.

[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0015] Figure 1is a schematic diagram of an exemplary system architecture to which a sample data processing method, a model training method, a task processing method, and an apparatus can be applied according to an embodiment of the present disclosure;

[0016] Figure 2 is a flow chart of a sample data processing method according to one embodiment of the present disclosure;

[0017] Figure 3 is a flowchart of forward processing using a multimodal large model according to one embodiment of the present disclosure;

[0018] Figure 4 is a flow chart of a sample data processing method according to another embodiment of the present disclosure;

[0019] Figure 5 is a flowchart of a task processing model training method according to an embodiment of the present disclosure;

[0020] Figure 6 is a flowchart of a task processing method according to an embodiment of the present disclosure;

[0021] Figure 7 is a block diagram of a sample data processing apparatus according to an embodiment of the present disclosure;

[0022] Figure 8 is a block diagram of a task processing model training apparatus according to one embodiment of the present disclosure;

[0023] Figure 9 is a block diagram of a task processing device according to an embodiment of the present disclosure; and

[0024] Figure 10 A schematic block diagram of an example electronic device 1000 is shown, which may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] In the sample data processing scenario of multimodal data, due to the diversity of dataset modalities, the dataset may contain a large number of noise samples such as modality misalignment samples or semantically ambiguous samples, as well as difficult or complex samples that are difficult for the target model to learn. These problems seriously affect the learning effect and generalization ability of the lightweight target model.

[0027] To reduce difficult or complex samples, datasets can be cleaned or filtered. Data cleaning methods can rely on extensive manual annotation and computational resources, which are costly and difficult to guarantee effective cleaning results. Multimodal sample dataset filtering schemes can filter based on semantic similarity between modalities or the quality of a single modality, without considering the role of samples in model training. Therefore, a technical solution is urgently needed to reduce the cost of data cleaning and improve the training efficiency and performance of lightweight target models.

[0028] In view of this, an embodiment of the present disclosure provides a sample data processing method, comprising: utilizing a large multimodal model to process multiple candidate sample data according to a target task indicated by prompt information, thereby obtaining multiple loss information for the multiple candidate sample data. Based on the multiple loss information for the multiple candidate sample data, the multiple candidate sample data are filtered to obtain at least one target sample data for training a task processing model. The large multimodal model has a greater number of parameters than the task processing model. The at least one target sample data can then be used to train the task processing model to obtain a target task processing model. The large multimodal model has a greater learning capability (parameter count) than the task processing model. This method can provide a cleaner and more consistent data distribution for a lightweight target model by introducing the discriminative capabilities of a model with strong learning capabilities, thereby improving the training efficiency and performance of the lightweight target model. This method is applicable to scenarios where filtering of multimodal datasets is required, as well as to multimodal data processing scenarios where resources are limited or rapid iterative deployment is required.

[0029] Figure 1 This is a schematic diagram of an exemplary system architecture to which a sample data processing method, a model training method, a task processing method, and an apparatus can be applied according to an embodiment of the present disclosure. It should be noted that: Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.

[0030] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0031] Terminal devices 101, 102, and 103 may be various electronic devices having display screens and supporting input of prompt words representing target tasks, multimodal selection of candidate sample data types, triggering of sample data processing tasks, triggering of training tasks, triggering of task processing, and display of processing results, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. Users can input operations on the display screens of terminal devices 101, 102, and 103 and interact with server 105 via network 104 to receive or transmit data or information.

[0032] Server 105 can be a server that provides various services, such as a backend management server (for example only) that supports operations performed by users on terminal devices 101, 102, and 103. The backend management server can filter multiple candidate sample data based on prompt information and sample data processing task triggers entered by users on terminal devices 101, 102, and 103, and input at least one target sample data obtained after filtering into lightweight target data to train the lightweight target model. The backend management server can also perform task processing (e.g., audio segmentation, image recognition, audio recognition, text-and-image question answering, etc.) based on the data to be processed (e.g., text, images, audio, video, etc.) entered by clients 101, 102, and 103, or pre-stored data to be processed (e.g., text, images, audio, video, etc.) retrieved from a storage unit (e.g., a data pool), and transmit the processing results to the display screens of terminal devices 101, 102, and 103 for display.

[0033] It should be noted that the sample data processing method, model training method, and task processing method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the sample data processing device, model training device, and task processing device provided in the embodiments of the present disclosure can generally be set in the server 105. The sample data processing method, model training method, and task processing method provided in the embodiments of the present disclosure can also be executed by the terminal devices 101, 102, and 103. Accordingly, the sample data processing device, model training device, and task processing device provided in the embodiments of the present disclosure can also be set in the terminal devices 101, 102, and 103. The sample data processing method, model training method, and task processing method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, and 103 and / or the server 105. Accordingly, the sample data processing device, model training device, and task processing device provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, and 103 and / or the server 105.

[0034] I understand. Figure 1 The number and types of terminal devices, networks, and servers in the embodiment are merely illustrative. Any number of terminal devices, networks, and servers may be provided as required.

[0035] It is understood that the system architecture of the present disclosure is described above, and the method of the present disclosure will be described below. It is also understood that the sequence numbers of the various operations in the following method are merely used to indicate the operation for the purpose of description and should not be regarded as indicating the order in which the various operations must be performed. Unless explicitly stated, the method does not need to be executed in the exact order shown.

[0036] In the technical solutions disclosed herein, all information and data involved (including but not limited to data used for analysis, stored data, displayed data, etc.) are authorized information and data, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, adopt necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for the choice of authorization or rejection.

[0037] Figure 2 is a flowchart of a sample data processing method according to an embodiment of the present disclosure.

[0038] like Figure 2 As shown, the method 200 may include operations S210 to S220.

[0039] In operation S210 , a multimodal large model is used to process the plurality of candidate sample data according to the target task indicated by the prompt information, thereby obtaining a plurality of loss information for the plurality of candidate sample data.

[0040] According to an embodiment of the present disclosure, the prompt information may be a natural language instruction or description provided by the user, which is used to guide the multimodal large model to perform task processing. The prompt information may be associated with the task processing model and determined by the task that the task processing model needs to perform. For example, if the target task performed by the task processing model is an image classification task, the prompt information may be "Please determine whether the animal in this picture is a cat or a dog." For another example, if the target task performed by the task processing model is a picture-text question-answering task, the prompt information may be "Based on the picture and the question, answer: What is the weather in the picture?" That is, different target tasks performed by the task processing model may correspond to different prompt information input into the multimodal large model.

[0041] According to an embodiment of the present disclosure, the candidate sample data may be multimodal data, for example, it may include images, text, audio, video, and the like. For example, for a scenario of product images + product descriptions, the corresponding candidate sample data may be images + text descriptions. For another example, for a scenario of video clips + background music, the corresponding multiple candidate sample data may be videos + audios. When training a task processing model, the candidate sample data for different target tasks may be different, and the multiple candidate sample data input into the multimodal large model may be determined according to the tasks performed by the task processing model. For example, for a picture-text question-answering task, multiple candidate sample data may include images + text descriptions.

[0042] According to an embodiment of the present disclosure, a multimodal large model can be an artificial intelligence model that has been pre-trained based on large-scale data, which can simultaneously process and understand multimodal data, and realize the joint learning and generation of multimodal information through cross-modal alignment and fusion.

[0043] According to an embodiment of the present disclosure, since the prompt information is determined based on the target task performed by the task processing model, the multimodal large model can be processed based on multiple candidate sample data according to the target task indicated by the prompt information, so that the target sample data obtained by subsequent filtering is more in line with the training requirements of the task processing model.

[0044] According to an embodiment of the present disclosure, multiple candidate sample data may be input into a multimodal large model for processing to obtain the loss corresponding to each sample data.

[0045] In operation S220 , the plurality of candidate sample data are filtered based on the plurality of losses for the plurality of candidate sample data to obtain at least one target sample data for training the task processing model.

[0046] According to an embodiment of the present disclosure, filtering can be for eliminating invalid sample data from multiple candidate sample data. Invalid sample data can be sample data that is not suitable for task processing model training, and correspondingly, valid sample data can be sample data that is suitable for task processing model training. For the task processing model, not only erroneous sample data are invalid sample data, other data samples can be used as valid sample data. Some sample data have high complexity and can be used as high-quality sample data for models with strong learning capabilities. However, for lightweight task processing models, it is difficult to learn sample data with high complexity, so they cannot be used as valid sample data. That is, the present disclosure considers the effectiveness of sample data for the task processing model from multiple dimensions, and does not only measure the effectiveness of sample data for the task processing model from dimensions such as the complexity of the sample data itself and whether it is correct.

[0047] According to an embodiment of the present disclosure, the number of parameters of the multimodal large model is greater than the number of parameters of the task processing model. The larger the number of parameters of the model, the stronger the learning ability, that is, the learning ability of the multimodal large model is higher than the learning ability of the task processing model. The learning ability of a model can refer to the ability of the model to extract patterns, rules and relationships from input data in a data-driven manner, and use this knowledge to perform prediction, classification, generation or other tasks. The learning ability of a model is the core indicator for measuring model performance, which determines the model's generalization ability and actual application effect on unseen data.

[0048] By selecting a large multimodal model with stronger learning ability than the task target model, multiple candidate sample data that did not participate in the training are processed. If the loss of some sample data in the multiple candidate sample data is relatively high for the large multimodal model with stronger learning ability, it means that the large multimodal model also finds it difficult to learn the patterns, rules and relationships of this part of the sample data. In this case, the lightweight task processing model with relatively weak learning ability will also find it difficult to learn the patterns, rules and relationships of this part of the sample data.

[0049] It's important to note that a relatively weak learning ability of a task processing model doesn't necessarily mean poor performance. Rather, it means maintaining high data processing accuracy while maintaining low model complexity, computational complexity, and memory usage. This means it's not about more complex samples, but rather correct and appropriate samples. Correctness can mean that the semantics of each modality in the sample data are correct and aligned flawlessly. Appropriate sample data, on the other hand, means that the sample data presents a moderate learning difficulty for the task processing model, allowing the model to learn relevant information during training.

[0050] Through the sample data processing method of the embodiment of the present disclosure, a multimodal large model with stronger learning ability than the task processing model is selected to directly process multiple candidate sample data, which can fully utilize the strong learning ability of the large model. Based on the loss of multiple candidate sample data, the candidate sample data is directly filtered. Under the synergistic effect of the prompt information, the difficult sample data and the erroneous sample data that are invalid for the task processing model can be screened out, and the target sample data that is valid for the task processing model can be obtained, so that the target sample data can be better matched with the task processing model, which can effectively improve the accuracy of the task processing model, and thus enable the task processing model to perform task processing more accurately. In addition, this sample data processing method does not require additional training of a large model, and can directly use the trained multimodal large model, which improves the efficiency of sample data processing and saves resources.

[0051] Figure 3 It is a flowchart of forward processing using a multimodal large model according to an embodiment of the present disclosure.

[0052] like Figure 3As shown, in an embodiment of the present disclosure, a multimodal large model is used based on multiple candidate sample data to perform processing according to the target task indicated by the prompt information to obtain multiple loss information for the multiple candidate sample data, which may include operations S311 to S312.

[0053] In operation S311 , features of different modalities are extracted from the candidate sample data, and the features of different modalities are fused with the prompt information to obtain fused features.

[0054] In operation S312, forward processing is performed based on the fused features to obtain loss information for the candidate sample data.

[0055] According to an embodiment of the present disclosure, a multimodal large model can convert prompt information and multiple candidate sample data into a processable format, for example, converting prompt information into word vectors through word segmentation and embedding, and encoding modal data into feature vectors through an encoder.

[0056] A large multimodal model can include multiple modality-specific encoders to process data from different modalities. For example, a text encoder can encode prompt information and text into high-dimensional vectors, an image encoder can encode images into feature vectors, and an audio encoder can encode audio into feature vectors. A large multimodal model can understand the relationship between features from different modalities by aligning them into the same semantic space, for example, by aligning text features with image features through contrastive learning or cross-modal attention mechanisms.

[0057] Semantic features extracted from the prompt information can be fused with features from different modalities using methods such as concatenation, weighted summation, and attention mechanisms. Concatenation involves directly concatenating semantic features with features from different modalities. Weighted summation involves assigning weights to features from different modalities based on the task being performed by the task processing model and summing the weights. Attention mechanisms can dynamically fuse features through self-attention or cross-modal attention. Interaction layers can also be applied to the fused features to further extract cross-modal correlation information. For example, in image-text question answering tasks, the model needs to understand the relationship between the visual information in the image and the semantic information in the prompt word.

[0058] After feature fusion is complete, the fused features can be fed into subsequent modules of the multimodal model (such as fully connected layers and classifiers) for forward propagation. The model uses the learned weights and biases to perform nonlinear transformations on the input features, gradually extracting higher-level abstract features. Forward processing is then performed based on these abstract features, and the resulting loss is combined with the sample data to obtain the loss.

[0059] Through the sample data processing method of the embodiment of the present disclosure, since data filtering directly utilizes the forward processing function of the multimodal large model to process the multimodal candidate sample data, the data filtering requirements can be met based on the loss. It does not involve the training of the multimodal large model and does not require most of the back propagation process of the multimodal large model. It can save resources and improve data filtering efficiency.

[0060] In an embodiment of the present disclosure, using a multimodal large model based on multiple candidate sample data, processing according to the target task indicated by the prompt information, and obtaining multiple loss information for the multiple candidate sample data may also include:

[0061] The multimodal large model is used to process multiple candidate sample data based on the target task indicated by the prompt information, and multiple processing results are obtained for the multiple candidate sample data respectively.

[0062] Based on the multiple processing results and the true labels of the multiple candidate sample data, the loss information of each candidate sample data is determined.

[0063] According to an embodiment of the present disclosure, the processing result of the candidate sample data may be: the judgment result of the category or attribute of the candidate sample data output by the multimodal large model after processing the target task. For example, for classification tasks, the category label (such as cat or dog) is output; for regression tasks, a continuous value (such as the size of the target object in the image) is output; for generation tasks, a text description (such as the title of the picture or the answer to the question) is output. The true label may be the correct category or attribute annotated for each candidate sample data among multiple candidate sample data. Taking into account that multiple candidate sample data can contain multiple modal information such as text, images, layout structure, etc., the cross-entropy loss function can be used to determine the loss information based on the true label and the processing result of the candidate sample data. This loss information characterizes the semantic difference between the features of different modalities and the difficulty of understanding the target task for the multimodal large model.

[0064] Through the sample data processing method of the embodiment of the present invention, the loss used is cross-entropy loss, which can reflect the difficulty of the model in understanding the semantic differences and task objectives between different modalities, and thus serve as an effective indicator to measure the quality of the sample, and then screen out high-quality sample data for the task processing model, which can effectively improve the accuracy of the task processing model, and thus enable the task processing model to accurately perform task processing.

[0065] In an embodiment of the present disclosure, filtering the plurality of candidate sample data based on the plurality of loss information for the plurality of candidate sample data to obtain at least one target sample data for training the task processing model may include:

[0066] Based on the plurality of loss information, loss distribution information for the plurality of candidate sample data is obtained.

[0067] Multiple candidate sample data are filtered based on the loss distribution information to obtain at least one target sample data.

[0068] According to the embodiments of the present disclosure, the loss distribution can be modeled to obtain loss distribution information. The modeling methods may include statistical modeling, probabilistic modeling, and the like. The reason why the sample data processing method of the present disclosure chooses to eliminate invalid sample data after obtaining the loss distribution through modeling, rather than directly eliminating invalid sample data based on the loss value of each candidate sample data, may be that directly eliminating sample data based on the loss value lacks a global perspective and the threshold is difficult to determine. Directly setting a fixed threshold (such as 0.75) cannot adapt to the characteristics of different data sets or tasks: the loss values ​​of some tasks are naturally high (such as classification tasks of complex scenes), and the fixed threshold may mistakenly delete normal samples; the loss values ​​of some tasks are naturally low (such as regression tasks of simple tasks), and the fixed threshold may not be able to eliminate truly abnormal samples. For example, in an image classification task, if samples with a loss value greater than 0.7 are directly eliminated, but the loss values ​​of the model are generally high in the early stages of training (such as a mean of 0.8), a large number of normal samples may be mistakenly deleted.

[0069] Through the sample data processing method of the embodiment of the present disclosure, invalid sample data is eliminated after obtaining the loss distribution, rather than directly filtering abnormal samples based on the comparison of the loss value with a fixed threshold. This can take into account the global perspective and adapt to the characteristics of different data sets or tasks, making the target task processing model more accurate and robust, thereby improving the accuracy and efficiency of task processing.

[0070] In an embodiment of the present disclosure, the loss information may include a loss value. Filtering multiple candidate sample data based on the loss distribution information to obtain at least one target sample data for training the task processing model may include:

[0071] The loss threshold is determined based on the mean and standard deviation of the loss distribution information.

[0072] Candidate sample data with loss values ​​greater than a loss threshold are filtered out from the plurality of candidate sample data to obtain at least one target sample data.

[0073] Figure 4 is a flowchart of a sample data processing method according to another embodiment of the present disclosure.

[0074] like Figure 4As shown, multiple candidate sample data and prompt information in the first data pool 410 can be input into the multimodal large model 420 for forward processing, and the processing results 430 of high-quality candidate sample data corresponding to each sample data are output. The loss value is calculated based on the processing results of the candidate sample data and the true label, and modeling is performed to obtain loss distribution information. Based on the loss distribution information, the mean μ and variance σ of all losses are determined. Based on the mean μ and variance σ, the loss threshold is determined, and the candidate sample data with loss values ​​greater than the loss threshold in the multiple candidate sample data are filtered out to obtain at least one target sample data 440, and then the at least one target sample data 440 is stored in the second data pool 450. When the training task is triggered, at least one target sample data 440 can be obtained from the second data pool 450 and input into the task processing model to train the task processing model.

[0075] The loss threshold can be set to μ+2σ, filtering out sample data with loss values ​​greater than μ+2σ from multiple candidate sample data, and retaining sample data with loss values ​​less than or equal to μ+2σ.

[0076] For example, if μ = 0.5 and σ = 0.1, the loss threshold is 0.5 + 2 × 0.1 = 0.7. Sample data with a loss value less than or equal to 0.7 are retained, and sample data with a loss value greater than 0.7 are discarded.

[0077] Through the sample data processing method of the embodiment of the present disclosure, by reasonably setting the loss threshold, invalid sample data in multiple candidate sample data can be more accurately filtered out, further improving the training efficiency and performance of the model.

[0078] In an embodiment of the present application, the target task includes at least one of the following:

[0079] Text processing tasks, image processing tasks, audio processing tasks, and video processing tasks.

[0080] According to embodiments of the present disclosure, text processing tasks may include, for example, text classification, text generation, information extraction, text question-and-answer (Q&A), information recommendation, and the like. Text classification may include, for example, spam detection, sentiment analysis, and topic classification. Text generation may include, for example, machine translation, text summarization, dialogue generation, and creative writing. Information extraction may include, for example, named entity recognition, relationship extraction, and event extraction. Text question-and-answer (Q&A) may include, for example, retrieving and generating answers to user questions from text or a knowledge base. Information recommendation may include, for example, recommending similar text content based on the text.

[0081] According to embodiments of the present disclosure, image processing tasks may include, for example, image classification, target detection, image segmentation, image generation, face recognition, pose estimation, image denoising and enhancement, and so on. Image classification may include, for example, object recognition, scene classification, and disease diagnosis. Target detection may include, for example, detecting pedestrians, vehicles, and traffic signs on the road, detecting abnormal behavior or suspicious objects in surveillance footage, and detecting the location of defective products or parts on a production line. Image segmentation may include, for example, assigning each pixel in an image to a specific category to achieve refined image understanding, such as medical image segmentation and remote sensing image segmentation. Image generation may include, for example, generating new images based on input, such as style transfer, super-resolution reconstruction, and image inpainting. Face recognition may include, for example, identifying or verifying the identity of a face in an image, for use in security, payment, social networking, and other scenarios. Pose estimation may include, for example, detecting the locations of key points of a person or object in an image, for use in motion analysis and motion capture. Image denoising and enhancement may include, for example, removing noise from an image or enhancing image quality, for use in low-light and blurry scenes.

[0082] According to embodiments of the present disclosure, audio processing tasks may include, for example, speech recognition, speech synthesis, speech verification, speech enhancement and noise reduction, audio classification, audio detection, audio source separation, and emotion recognition. Speech recognition, for example, involves converting speech signals into text for applications such as speech transcription, voice assistants, and real-time subtitles. Speech synthesis, for example, involves converting text into natural and fluent speech for applications such as voice broadcasting, virtual anchors, and accessible reading. Speech verification, for example, involves identifying or verifying the speaker's identity for applications such as security, payment, and personalized services. Speech enhancement and noise reduction, for example, involves removing noise or interference from speech to improve speech quality. Audio classification, for example, involves categorizing audio into predefined categories, such as environmental sound classification or musical style classification. Audio detection, for example, involves detecting and locating specific events (such as breaking glass) in an audio stream. Audio source separation, for example, involves isolating individual sound sources from mixed audio, such as separating human voice from accompaniment or separating multiple speakers. Emotion recognition, for example, involves identifying emotional states (such as happiness, sadness, and anger) in speech for applications such as human-computer interaction and mental health monitoring.

[0083] According to embodiments of the present disclosure, video processing tasks may include, for example, target detection and tracking, video generation and editing, video enhancement and restoration, video search and recommendation, video interaction, augmented reality, and the like. Target detection and tracking may include, for example, real-time identification of objects (such as vehicles and pedestrians) in a video and tracking their motion trajectories. Video generation and editing may include, for example, generating new video content or synthesizing virtual scenes. Video enhancement and restoration may include, for example, removing noise or blur from a video to improve image quality. Video search and recommendation may include, for example, searching based on video content (such as objects, scenes, and actions). Video interaction may include, for example, controlling video playback or interaction through gestures. Augmented reality may include, for example, overlaying virtual objects or information on a video.

[0084] It should be noted that the specific application scenarios listed above are only exemplary. Different task processing models can be trained to achieve corresponding tasks according to different needs.

[0085] Based on the above sample data processing method, an embodiment of the present disclosure also provides a task processing model training method. Figure 5 4 is a flowchart of a task processing model training method according to an embodiment of the present disclosure.

[0086] like Figure 5 As shown, the method 500 may include operations S510 to S520.

[0087] In operation S510 , at least one target sample data is acquired.

[0088] According to an embodiment of the present disclosure, at least one target sample data is obtained based on the above-mentioned sample data processing method.

[0089] In operation S520 , the to-be-trained task processing model is trained using at least one target sample data to obtain a target task processing model.

[0090] It should be noted that the specific details of the task processing model training method are similar to the sample data processing method and will not be repeated here.

[0091] Through the task processing model training method of the disclosed embodiment, a multimodal large model with stronger learning ability than the task processing model is selected to directly process multiple candidate sample data, which can fully utilize the large model's stronger learning ability. Based on the loss of the multimodal large model, data filtering is directly performed. Under the synergistic effect of prompt information, difficult samples and erroneous samples that are invalid for the task processing model can be screened out, and training samples that are valid for the task processing model can be obtained, so that at least one target sample data is more capable of matching the task processing model, effectively improving the accuracy of the task processing model.

[0092] Based on the above-mentioned task processing model training method, an embodiment of the present disclosure also provides a task processing method. Figure 6 is a flowchart of a task processing method according to an embodiment of the present disclosure.

[0093] like Figure 6 As shown, the method 600 may include operations S610 to S620.

[0094] In operation S610 , data to be processed is acquired.

[0095] According to an embodiment of the present disclosure, the data to be processed may include at least one of text, image, audio, and video.

[0096] In operation S620 , the data to be processed is input into the target task processing model, and the task processing result is output.

[0097] According to an embodiment of the present disclosure, the target task processing model can be trained based on the above-mentioned task processing model training method. For specific implementation details, please refer to the training method embodiment section, which will not be repeated here.

[0098] Through the task processing method of the embodiment of the present disclosure, a multimodal large model with stronger learning ability than the task processing model is selected to directly process multiple candidate sample data, which can fully utilize the large model's stronger learning ability. Based on the loss of the multimodal large model, data filtering is directly performed. Under the synergistic effect of prompt information, difficult samples and erroneous samples that are invalid for the task processing model can be screened out, and training samples that are valid for the task processing model can be obtained, so that at least one target sample data is more able to match the task processing model, improving the accuracy of the task processing model, and thus enabling the task processing model to perform task processing more accurately.

[0099] The following will be combined Figure 7 A sample data processing device according to an embodiment of the present disclosure is schematically described. Figure 7 is a block diagram of a sample data processing apparatus according to an embodiment of the present disclosure.

[0100] like Figure 7 As shown, the sample data processing device 700 may include a processing module 710 and a filtering module 720 .

[0101] Processing module 710 is configured to utilize the multimodal large model to process the plurality of candidate sample data according to the target task indicated by the prompt information, thereby obtaining a plurality of loss information for the plurality of candidate sample data. In one embodiment, processing module 710 may be configured to perform operation S210 described above, which will not be further described here.

[0102] The filtering module 720 is configured to filter the plurality of candidate sample data based on the plurality of loss information for the plurality of candidate sample data to obtain at least one target sample data for training the task processing model. In one embodiment, the filtering module 720 may be configured to perform the operation S220 described above, which will not be described in detail here.

[0103] According to an embodiment of the present disclosure, the processing module 710 uses a multimodal large model to process multiple candidate sample data according to the target task indicated by the prompt information to obtain multiple loss information for the multiple candidate sample data, which may include:

[0104] Features of different modalities are extracted from candidate sample data, and the features of different modalities are fused with prompt information to obtain fused features.

[0105] Forward processing is performed based on the fused features to obtain loss information for candidate sample data.

[0106] According to an embodiment of the present disclosure, the processing module 710 uses a multimodal large model to process multiple candidate sample data according to the target task indicated by the prompt information to obtain multiple loss information for the multiple candidate sample data, and may also include:

[0107] The multimodal large model is used to process multiple candidate sample data based on the target task indicated by the prompt information, and multiple processing results are obtained for the multiple candidate sample data respectively.

[0108] Based on multiple processing results and the true labels of multiple candidate sample data, the loss information of each candidate sample data is determined. The loss information represents the semantic differences between the features of different modalities and the difficulty of understanding the target task of the multimodal large model.

[0109] According to an embodiment of the present disclosure, the filtering module 720 filters the plurality of candidate sample data based on the plurality of loss information for the plurality of candidate sample data to obtain at least one target sample data for training the task processing model, which may include:

[0110] Based on the multiple loss information, loss distribution information of the multiple candidate sample data is obtained.

[0111] Multiple candidate sample data are filtered based on the loss distribution information to obtain at least one target sample data.

[0112] According to an embodiment of the present disclosure, the loss information includes a loss value. The filtering module 720 filters the plurality of candidate sample data based on the loss distribution information, which may include:

[0113] The loss threshold is determined based on the mean and standard deviation of the loss distribution information.

[0114] Candidate sample data with loss values ​​greater than a loss threshold are filtered out from the plurality of candidate sample data to obtain at least one target sample data.

[0115] According to an embodiment of the present disclosure, the target task may include at least one of the following:

[0116] Text processing tasks, image processing tasks, audio processing tasks, and video processing tasks.

[0117] It should be noted that other embodiment details of the sample data processing device and the technical effects brought about are the same or similar to the embodiment details of the sample data processing method and the technical effects brought about, and will not be repeated here.

[0118] The following will be combined Figure 8 A task processing model training device according to an embodiment of the present disclosure is schematically described. Figure 8 4 is a block diagram of a task processing model training device according to an embodiment of the present disclosure.

[0119] like Figure 8 As shown, the task processing model training device 800 may include a first acquisition module 810 and a training module 820.

[0120] The first acquisition module 810 is configured to acquire at least one target sample data. In one embodiment, the first acquisition module 810 may be configured to execute the operation S510 described above, which will not be described in detail herein.

[0121] The training module 820 is configured to train the to-be-trained task processing model using at least one target sample data to obtain a target task processing model. In one embodiment, the training module 820 may be configured to perform the operation S520 described above, which will not be described in detail herein.

[0122] It should be noted that the details of other embodiments of the task processing model training device and the technical effects brought about are the same or similar to the details of the embodiments of the task processing model training method and the technical effects brought about, and will not be repeated here.

[0123] The following will be combined Figure 9 A task processing device according to one embodiment of the present disclosure is schematically described. Figure 9 is a block diagram of a task processing device according to an embodiment of the present disclosure.

[0124] like Figure 9 As shown, the task processing device 900 may include a second acquisition module 910 and an input / output module 920 .

[0125] The second acquisition module 910 is configured to acquire the data to be processed. In one embodiment, the second acquisition module 910 may be configured to execute the operation S610 described above, which will not be described in detail herein.

[0126] The input / output module 920 is configured to input the data to be processed into the target task processing model and output the task processing result. In one embodiment, the input / output module 920 may be configured to execute the operation S620 described above, which will not be described in detail here.

[0127] According to an embodiment of the present disclosure, the target task processing model can be obtained by training based on the training method of the present disclosure or by training using the training device of the present disclosure.

[0128] It should be noted that the details of other embodiments of the task processing device and the technical effects brought about are the same or similar to the details of the embodiments of the task processing method and the technical effects brought about, and will not be repeated here.

[0129] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0130] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0131] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. Computing unit 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0132] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0133] Computing unit 1001 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1001 performs the various methods and processes described above, such as the text processing method and / or the method for deploying a deep learning framework. For example, in some embodiments, the text processing method and / or the method for deploying a deep learning framework can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the text processing method and / or deep learning framework deployment method described above may be performed. Alternatively, in other embodiments, computing unit 1001 may be configured to execute the text processing method and / or deep learning framework deployment method in any other appropriate manner (e.g., via firmware).

[0134] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0135] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0136] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0138] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0139] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0140] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0141] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A sample data processing method, comprising: Using the multimodal large model to process the plurality of candidate sample data according to the target task indicated by the prompt information, thereby obtaining a plurality of loss information for the plurality of candidate sample data; as well as filtering the plurality of candidate sample data based on a plurality of loss information for the plurality of candidate sample data to obtain at least one target sample data for training the task processing model; Among them, the number of parameters of the multimodal large model is greater than the number of parameters of the task processing model.

2. The method according to claim 1, wherein The multimodal large model is used to process the plurality of candidate sample data according to the target task indicated by the prompt information to obtain a plurality of loss information for the plurality of candidate sample data, including: Extracting features of different modalities from the candidate sample data, and fusing the features of different modalities with the prompt information to obtain fused features; and Forward processing is performed based on the fused features to obtain loss information for the candidate sample data.

3. The method according to claim 2, wherein: The method of using the multimodal large model to process the plurality of candidate sample data according to the target task indicated by the prompt information to obtain a plurality of loss information for the plurality of candidate sample data further includes: Using the multimodal large model to process the plurality of candidate sample data according to the target task indicated by the prompt information, thereby obtaining a plurality of processing results respectively for the plurality of candidate sample data; and Based on the multiple processing results and the respective true labels of the multiple candidate sample data, the loss information of each candidate sample data is determined, and the loss information represents the semantic difference between the features of the different modalities and the difficulty of understanding the target task of the multimodal large model.

4. The method according to claim 1, wherein The filtering of the plurality of candidate sample data based on the plurality of loss information for the plurality of candidate sample data to obtain at least one target sample data for training the task processing model includes: Based on the plurality of loss information, obtaining loss distribution information for the plurality of candidate sample data; and The plurality of candidate sample data are filtered based on the loss distribution information to obtain at least one target sample data.

5. The method according to claim 4, wherein The loss information includes a loss value, The filtering of the plurality of candidate sample data based on the loss distribution information to obtain at least one target sample data for training the task processing model includes: determining a loss threshold based on the mean and standard deviation of the loss distribution information; and The candidate sample data whose loss value is greater than the loss threshold are filtered out from the plurality of candidate sample data to obtain at least one target sample data.

6. The method according to claim 1, wherein The target tasks include at least one of the following: Text processing tasks, image processing tasks, audio processing tasks, and video processing tasks.

7. A task processing model training method, comprising: Using at least one target sample data to train a task processing model to be trained, to obtain a target task processing model; The at least one target sample data is obtained according to the method according to any one of claims 1 to 6.

8. A task processing method, comprising: Input the data to be processed into the target task processing model and output the task processing results; Wherein, the target task processing model is obtained by training based on the training method described in claim 7.

9. A sample data processing device, comprising: a processing module, configured to use the multimodal large model to perform processing based on the plurality of candidate sample data according to the target task indicated by the prompt information, and obtain a plurality of loss information for the plurality of candidate sample data; as well as a filtering module, configured to filter the plurality of candidate sample data based on a plurality of loss information for the plurality of candidate sample data to obtain at least one target sample data for training the task processing model; Among them, the number of parameters of the multimodal large model is greater than the number of parameters of the task processing model.

10. A task processing model training device, comprising: A training module, configured to train a task processing model to be trained using at least one target sample data to obtain a target task processing model; The at least one target sample data is obtained according to the method according to any one of claims 1 to 6.

11. A task processing device comprising: An acquisition module is used to obtain data to be processed; as well as An input and output module, used to input the data to be processed into the target task processing model and output the task processing result; Wherein, the target task processing model is obtained by training based on the training method described in claim 7 or by training based on the training device described in claim 9.

12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.