Large model optimization method and device based on small samples, equipment and storage medium
By vertically optimizing and data enhancement processing of multimodal data in small sample scenarios, a hybrid data set is constructed to optimize the base model, which solves the problem of overfitting the model in small sample scenarios, and achieves better adaptation and generalization capabilities of the model to multimodal service tasks and improves the model.
Patent Information
- Application Number
- CN202510380927.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In small sample scenarios, due to insufficient training samples, the model is prone to overfitting and cannot cover a wider range of scenarios and situations, making it difficult for the model to adapt to multimodal service tasks.
By acquiring initial multimodal data, determining its corresponding vertical field, and optimizing the pre-acquisitioned third-party large language model. Then the initial text data and image data are extracted, and the enhancement processing is performed separately to generate enhanced text data and selected images. This data is constructed into a hybrid dataset and the base model is optimized using this dataset to obtain an optimized model adapted to multimodal service tasks.
By enhancing the coverage and situation of the data, the model can better adapt to multimodal service tasks and improve the adaptability and generalization capabilities of the model.
Smart Images

Figure CN120180133A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model optimization, and specifically relates to a method, device, equipment, and storage medium for optimizing large models based on small samples. Background Art
[0002] With the development of technology, digital services on the Internet are emerging in an endless stream, such as online car-hailing, Internet hospitals, etc. Facing increasingly complex data service tasks, it is becoming increasingly important to build a model or training framework that can quickly adapt to multi-modal service tasks. Traditional machine learning methods often rely on a large amount of labeled data to improve the accuracy and generalization ability of the model. However, in some actual application scenarios, it is costly and difficult to obtain a large-scale data set. Therefore, more and more attention is now shifted to small sample scenarios.
[0003] In small sample scenarios, due to insufficient training samples, the model is prone to overfitting and cannot cover a wider range of scenarios and situations, resulting in the model being difficult to adapt to multi-modal service tasks. Summary of the Invention
[0004] In view of this, this application provides a method, device, equipment, and storage medium for optimizing large models based on small samples, which is used to solve the problem that in small sample scenarios, due to insufficient training samples, the model is prone to overfitting and cannot cover a wider range of scenarios and situations, resulting in the model being difficult to adapt to multi-modal service tasks.
[0005] To achieve the above objectives, the following solutions are proposed:
[0006] In a first aspect, a method for optimizing a large model based on small samples includes:
[0007] Obtain initial multi-modal data;
[0008] Determine the vertical domain corresponding to the initial multi-modal data, and optimize a pre-obtained third-party large language model based on the vertical domain;
[0009] Extract the initial text data from the initial multi-modal data, and input the initial text data into the optimized third-party large language model to obtain the output enhanced text data;
[0010] Extract the initial image data from the initial multi-modal data, and perform data augmentation processing on the initial image data to obtain each selected image;
[0011] Construct a mixed data set from the initial multi-modal data, enhanced text data, and each selected image;
[0012] Optimize a pre-established base large model using the mixed data set to obtain an optimized multi-modal base large model.
[0013] Preferably, optimizing the pre-obtained third-party large language model based on the vertical domain includes:
[0014] Establishing an initial tag text that matches the vertical domain;
[0015] Expanding the initial tag text to obtain specific description information;
[0016] Constructing a loss function according to the specific description information;
[0017] Using the loss function to optimize the soft prompts of the pre-obtained third-party large language model.
[0018] Preferably, the process of the optimized third-party large language model processing the initial text data and outputting the enhanced text data includes:
[0019] Establishing a target question based on the vertical domain;
[0020] Decomposing the target question to obtain each step text;
[0021] Constructing a logical reasoning process from each step text;
[0022] Generating respective corresponding prompt texts for each step text;
[0023] Adding the prompt texts to the logical reasoning process according to their respective corresponding step texts to obtain a thought chain;
[0024] Reasoning on the input initial text data according to the thought chain, obtaining the enhanced text data and outputting it.
[0025] Preferably, reasoning on the input initial text data according to the thought chain, obtaining the enhanced text data and outputting it includes:
[0026] Obtaining domain-specific knowledge corresponding to the vertical domain;
[0027] Reasoning on the initial text data according to the thought chain;
[0028] While reasoning, real-time extracting the output answers corresponding to each step text;
[0029] For each step text, adjusting the prompt text corresponding to the next step text according to its corresponding output answer;
[0030] While adjusting, adding the domain-specific knowledge to the prompt text to infer the enhanced text data corresponding to the initial text data.
[0031] Preferably, the processing and enhancement of the initial image data to obtain each selected image includes:
[0032] Identifying the main object in the initial image data and determining the position, contour, and category of the main object;
[0033] Performing segmentation processing on the initial image data based on the position, contour, and category to obtain first image data;
[0034] Determining the key labels and text attributes corresponding to the vertical domain and generating a structured text description according to the key labels and text attributes;
[0035] Processing the structured text description using a pre-trained text-to-image model to obtain each task main image corresponding to the structured text description;
[0036] Screening each task main image corresponding to each initial image data to obtain each selected image.
[0037] Preferably, the screening of each task main image corresponding to each initial image data to obtain each selected image includes:
[0038] Performing a preliminary screening on each task main image according to a preset clarity threshold, integrity threshold, label matching threshold, and background complexity threshold to obtain each roughly selected image;
[0039] Inputting each roughly selected image and the initial image data into a pre-trained expert model, so that the expert model performs precise screening on each roughly selected image according to the initial image data, and outputs the screened images as each selected image.
[0040] Preferably, the construction of the hybrid dataset from the initial multimodal data, enhanced text data, and each selected image includes:
[0041] Obtaining the prompt text and enhanced text data corresponding to the initial text data;
[0042] For each selected image, associating the prompt text and enhanced text data with the selected image to generate each image-text data pair;
[0043] Performing authenticity scoring on each image-text data pair corresponding to each selected image to screen out each target data pair;
[0044] Constructing a data pair set from each target data pair and performing data balancing processing on the data pair set to obtain each balanced data pair;
[0045] Construct a hybrid dataset from the initial multimodal data, each of the balanced data pairs, and each selected image.
[0046] In a second aspect, an optimization device for large models based on small samples includes:
[0047] A multimodal data acquisition module for acquiring initial multimodal data;
[0048] A large language model optimization module for determining the vertical domain corresponding to the initial multimodal data and optimizing a pre-acquired third-party large language model based on the vertical domain;
[0049] A model processing module for extracting the initial text data from the initial multimodal data and inputting the initial text data into the optimized third-party large language model to obtain the output enhanced text data;
[0050] A data augmentation processing module for extracting the initial image data from the initial multimodal data and performing data augmentation processing on the initial image data to obtain each selected image;
[0051] A hybrid dataset construction module for constructing a hybrid dataset from the initial multimodal data, the enhanced text data, and each selected image;
[0052] A base large model optimization module for optimizing a pre-established base large model using the hybrid dataset to obtain an optimized multimodal base large model.
[0053] In a third aspect, an optimization device for large models based on small samples includes a memory and a processor;
[0054] The memory is used for storing programs;
[0055] The processor is used for executing the programs to implement each step of the optimization method for large models based on small samples as described in any item of the first aspect.
[0056] In a fourth aspect, a storage medium stores a computer program, and when the computer program is executed by a processor, each step of the optimization method for large models based on small samples as described in any item of the first aspect is implemented.
[0057] As can be seen from the above technical solution, the present application obtains initial multimodal data; determines the vertical domain corresponding to the initial multimodal data, and optimizes a pre-acquired third-party large language model based on the vertical domain; extracts the initial text data in the initial multimodal data, inputs the initial text data into the optimized third-party large language model to obtain the output enhanced text data; extracts the initial image data in the initial multimodal data, and performs data enhancement processing on the initial image data to obtain each selected image; constructs a mixed dataset from the initial multimodal data, the enhanced text data, and each selected image; and uses the mixed dataset to optimize a pre-established base large model to obtain an optimized multimodal base large model. The present application classifies the initial multimodal data and processes it according to the classified data. The third-party large language model can be optimized through the vertical domain of the initial multimodal data, making the third-party large language model more targeted. Then, the optimized third-party large language model can better process the initial text data in the multimodal data to obtain enhanced text data, and then perform data enhancement processing on the initial image data in the multimodal data to obtain selected images. Through the above processing process, the coverage range and situation of the data can be made wider, enabling the model to adapt to multimodal service tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0059] Figure 1 It is an optional flowchart of a large model optimization method based on small samples provided by an embodiment of the present application;
[0060] Figure 2 It is an optional structural block diagram of a large model optimization method based on small samples provided by an embodiment of the present application;
[0061] Figure 3 It is a schematic structural diagram of a large model optimization device based on small samples provided by an embodiment of the present application;
[0062] Figure 4 It is a schematic structural diagram of a large model optimization device based on small samples provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0064] With the development of technology, digital services on the Internet are emerging in an endless stream, such as online car-hailing, Internet hospitals, etc. In the face of increasingly complex data service tasks, it is becoming increasingly important to build a model or training framework that can quickly adapt to multi-modal service tasks. Traditional machine learning methods often rely on a large amount of labeled data to improve the accuracy and generalization ability of the model. However, in some actual application scenarios, it is costly and difficult to obtain a large-scale data set. Therefore, more and more attention is now shifted to small-sample scenarios. When faced with a small amount of manually labeled data, it is of great significance to conduct efficient training to ensure the prediction performance of the model in such scenarios. Traditional machine learning methods often rely on a large amount of labeled data to improve the accuracy and generalization ability of the model. However, in some actual application scenarios, it is costly and difficult to obtain a large-scale data set. The significance of small-sample training lies in that it can break through the data volume limit, extract effective information from limited samples through delicate model design and innovative algorithms, and improve the prediction ability and adaptability of the model. Especially in the field of digital services, many tasks in this field often involve privacy and security, and labeled data is usually scarce and expensive. The method of constructing small samples for training not only helps to improve the model's learning ability for scarce data, but also promotes the application of data services in small-sample scenarios, promotes the popularization and implementation of intelligent technologies in more industries, and exploring and developing small-sample training technologies has broad application prospects and profound industrial significance.
[0065] However, in small-sample scenarios, due to insufficient training samples, the model is prone to overfitting and cannot cover a wider range of scenarios and situations, resulting in the model being difficult to adapt to multi-modal service tasks.
[0066] To solve the above defects of the prior art, an embodiment of the present invention provides a large model optimization method based on small samples. This method can be applied to various computer terminals or intelligent terminals, and its execution subject can be a processor or server of a computer terminal or intelligent terminal. The method flow chart of the method is as Figure 1 shown, and specifically includes:
[0067] S1: Obtain initial multi-modal data.
[0068] The initial multi-modal data obtained in this application contains various types or structures of data, such as text data, form data, picture data, video frame data, etc., so as to cover a wider range.
[0069] After obtaining the initial multi-modal data, in order to make the subsequent process more efficient, the multi-modal data is preprocessed, including removing redundant data, which can ensure the alignment of different modal data in the time or space dimension. For example, the narration or descriptive text corresponding to video frame data at different times should be timestamp-aligned to make the multi-modal data more regular.
[0070] The initial multi-modal data obtained in this application is a small amount of manually annotated data based on the vertical domain task.
[0071] S2: Determine the vertical domain corresponding to the initial multi-modal data, and optimize the pre-obtained third-party large language model based on the vertical domain.
[0072] In this step, a third-party large language model is introduced, and the third-party large language model is optimized according to the vertical domain corresponding to the initial multi-modal data, so as to perform relevant processing on the initial multi-modal data by virtue of the context understanding ability and reason reasoning ability of the third-party large language model.
[0073] S3: Extract the initial text data from the initial multi-modal data, and input the initial text data into the optimized third-party large language model to obtain the output enhanced text data.
[0074] Since the initial multi-modal data contains various types of data, different types of data are separated and processed respectively. Among them, for the text data and form data in the initial multi-modal data, which are collectively referred to as the initial text data, inputting it into the optimized third-party large language model can generate high-quality enhanced text based on a small amount of annotated initial text data.
[0075] The third-party large language model adopts a large language model that can generate various instruction texts including question answering and reasoning for specific tasks.
[0076] S4: Extract the initial image data from the initial multi-modal data, and perform data augmentation processing on the initial image data to obtain each selected image.
[0077] Since the initial multi-modal data also contains picture data and video frame data, it is extracted and collectively referred to as the initial image data, and then data augmentation processing is performed on the initial image data to obtain each selected image.
[0078] In this step, the third-party large language model for images and video frames can also be called. Using this third-party large language model, the problem of scarce samples in some categories can be alleviated by the way of text-to-image, and the data distribution and data deviation can be balanced. This third-party large language model refers to the third-party diffusion model.
[0079] S5: Construct a hybrid dataset from the initial multimodal data, enhanced text data, and each selected image.
[0080] During the construction process, multi-metric data evaluation can be performed for screening and regeneration. For example, multimodal data containing duplicate content or irrelevant information can be deleted to ensure the quality and relevance of the dataset. At the same time, the initial multimodal data, enhanced text data, and each selected image can be screened according to diversity metrics (such as the variation range and distribution characteristics of samples) and authenticity metrics (such as the generation quality of data and the degree of compliance with actual tasks), and low-quality or non-compliant samples can be removed to construct a high-quality multimodal dataset with both rich diversity and high credibility.
[0081] S6: Optimize the pre-established base large model using the hybrid dataset to obtain an optimized multimodal base large model.
[0082] Instruction fine-tuning can be performed on the data service base large model to be optimized based on the established hybrid dataset to quickly align it with different data service task scenarios and generate standardized output results, such as Figure 2 shown.
[0083] From the above technical solutions, it can be seen that this application obtains initial multimodal data; determines the vertical domain corresponding to the initial multimodal data, and optimizes the pre-obtained third-party large language model based on the vertical domain; extracts the initial text data from the initial multimodal data, inputs the initial text data into the optimized third-party large language model to obtain the output enhanced text data; extracts the initial image data from the initial multimodal data, and performs data augmentation processing on the initial image data to obtain each selected image; constructs a hybrid dataset from the initial multimodal data, enhanced text data, and each selected image; and optimizes the pre-established base large model using the hybrid dataset to obtain an optimized multimodal base large model. This application classifies the initial multimodal data and processes it separately according to the classified data. The third-party large language model can be optimized through the vertical domain of the initial multimodal data, making the third-party large language model more targeted. Then, the optimized third-party large language model can better process the initial text data in the multimodal data to obtain enhanced text data, and perform data augmentation processing on the initial image data in the multimodal data to obtain selected images. Through the above processing process, the coverage range and situation of the data can be made more extensive, enabling the model to adapt to multimodal service tasks.
[0084] In the method provided by the embodiments of the present invention, the process of optimizing the pre-obtained third-party large language model based on the vertical domain is specifically described as follows:
[0085] Establish initial label text matching the vertical domain;
[0086] Expand the initial label text to obtain specific description information;
[0087] Construct a loss function according to the specific description information;
[0088] Optimize the soft prompt of the pre - obtained third - party large - language model using the loss function.
[0089] Specifically, when optimizing a third - party large - language model based on a vertical domain, it is necessary to establish initial label text matching the vertical domain, which can be a single keyword or an image category name, etc. This initial label text is closely related to the vertical domain, ensuring accuracy and consistency. Then, it needs to be expanded into a more semantically informative detailed image description. In one example, the vertical domain is the field of automobile driving, with abnormal vehicle driving behavior as the task. Taking "Name: XXX, Vehicle type: XXX, Accident type: Rear - end collision" as the initial label text, after expansion, it becomes: "This online car - hailing driver, named XXX, male, drives a vehicle of type XXX and is involved in a rear - end collision accident. According to risk assessment, there are potential risks in this accident, and it is necessary to further analyze the reasons and liability attribution to evaluate the safety of his service."
[0090] Then, in this process, it combines the background knowledge and task requirements of the vertical domain to generate structured or natural - language detailed descriptions, providing more rich - semantic information support for image generation, task modeling, and subsequent model fine - tuning.
[0091] Among them, the loss function constructed according to the specific description information is the function construction after freezing the third - party large - language model. Then, for this large - language model Ψ, a learnable tensor θ (also known as "soft prompt") is constructed, and the following loss function is constructed:
[0092] ;
[0093] Among them, is the specific description information. By using keywords or descriptions related to the vertical domain to optimize the soft prompt instead of modifying the parameters of the entire third - party large - language model, the problems of overfitting and parameter inflation can be avoided. The above soft - prompt learning is not only applicable to scenarios with scarce data but also supports cross - modal tasks, ensuring that the third - party large - language model generates text data focusing on specific scenarios or tasks in vertical - domain tasks.
[0094] The following will elaborate on the process of the optimized third - party large - language model in this application processing the initial text data and outputting the enhanced text data.
[0095] Establish a target problem based on the vertical domain;
[0096] Decompose the target problem to obtain each step text;
[0097] Construct a logical reasoning process from each step text;
[0098] Generate respective corresponding prompt texts for each step text;
[0099] Add the prompt texts to the logical reasoning process according to their respective corresponding step texts to obtain a thought chain;
[0100] Reason about the input initial text data according to the thought chain to obtain enhanced text data and output it.
[0101] Among them, for the step of reasoning about the input initial text data according to the thought chain to obtain enhanced text data and output it, it includes:
[0102] Obtain domain-specific knowledge corresponding to the vertical domain;
[0103] Reason about the initial text data according to the thought chain;
[0104] During the reasoning, real-time extract the output answers corresponding to each step text;
[0105] For each step text, adjust the prompt text corresponding to the next step text according to its corresponding output answer;
[0106] During the adjustment, add the domain-specific knowledge to the prompt text to infer the enhanced text data corresponding to the initial text data.
[0107] Specifically, the above process can be divided into three processes: soft prompt learning including specific vertical domain tasks, prompt text generation based on the thought chain, and prompt text generation based on vertical domain task labels and text attributes.
[0108] In the above process, according to the generated prompt texts, the model can be gradually guided to generate content according to the thought chain, thereby improving the answering ability of complex tasks. Among them, the prompt text for each step clarifies the partial reasoning or operations that the model needs to complete, making the generation process more organized. And during the generation process, adjusting the prompt text of the subsequent step text in real time according to the output of each step can ensure the coherence and accuracy of the thought chain, and at the same time enhance the control of the generated output answers. Adding domain-specific knowledge to the prompt text can improve the understanding and adaptation ability of the third-party large language model for vertical domain tasks.
[0109] Next, taking abnormal vehicle driving as an example, the process of generating the above-mentioned thinking chain to obtain enhanced text data will be specifically described as follows:
[0110] Input: "Suppose we are analyzing the driving behavior data of a vehicle, and the data includes timestamp, speed, steering wheel angle, and braking intensity. Please list the categories of driving behaviors that these data may reflect.
[0111] Model Output: "These data can reflect the following categories of driving behaviors: normal driving, speeding, hard braking, sudden lane change, fatigue driving, etc."
[0112] Input: "Suppose the following behaviors are detected: the speed exceeds 120 km / h, and hard braking occurs 5 times within 10 minutes. Please determine whether the driving behavior is abnormal based on these conditions."
[0113] Model Output: "Abnormal driving behavior: The speed exceeding 120 km / h belongs to speeding, and frequent hard braking belongs to a possible dangerous behavior."
[0114] Input: "Based on the above abnormal situations, generate a detailed abnormal description for the driver, including problems and suggestions."
[0115] Model Output: "Speeding (speed exceeding 120 km / h) and frequent hard braking (hard braking 5 times within 10 minutes) are detected. These behaviors may increase the risk of accidents. It is recommended that the driver comply with the speed limit regulations, maintain a steady driving, and avoid frequent hard braking."
[0116] In the above process, another effective prompting text can also be designed by combining the initial label text and text attributes of the vertical domain task to generate a high-quality data set, so as to support the training of the subsequent base large model. According to the requirements of the vertical domain task, key labels and text attributes are extracted, and the data distribution is analyzed to identify possible biases. Subsequently, the trained attribute classifier is used to check the attribute distribution in the generated data, identify significant biases and explicitly add attribute constraints to this prompting text to guide the generation model to output diverse data containing target attributes. By dynamically adjusting the prompting design, it is ensured that the generated data covers a wider range of attribute scopes.
[0117] In an example, taking abnormal vehicle driving as an example again, instead of simply constructing picture descriptions and abnormal recognition instructions, this application uses key labels and attribute distributions (such as "driver behavior", "driving environment", "degree of distraction", "potential consequences", and "intervention suggestions", etc.) to generate instruction text results based on different attribute values. For example, the attribute of "driver behavior" includes "electronic device use", "interaction behavior", and "other distraction behaviors". At this time, a third-party large language model can be used to construct instructions:
[0118] Usage of electronic devices: "Is the driver using a mobile phone? Please specifically describe whether it is making a call, sending a text message, or using other applications." "Is the driver operating a navigation device? For example, adjusting the destination, viewing the map, etc."
[0119] Interactive behaviors: "Is the driver talking to the passengers in the vehicle? Does this behavior cause the driver's attention to shift?" "Does the driver have the behavior of looking back? Please describe the possible reasons, such as interacting with the passengers in the back seat or checking items."
[0120] Other distracting behaviors: "Does the driver have the behaviors of eating, drinking, or other non-driving-related operations? Please list the specific behaviors."
[0121] In this way, it can ensure the balance of various attribute data, and then efficiently train the model to ensure that the model can judge whether there is a driving abnormal behavior and its details, and at the same time provide improvement suggestions.
[0122] The following embodiments will explain in detail the steps of processing and enhancing the initial image data to obtain each selected image in the present application.
[0123] Identify the main object in the initial image data, and determine the position, contour, and category of the main object;
[0124] Perform segmentation processing on the initial image data based on the position, contour, and category to obtain the first image data;
[0125] Determine the key labels and text attributes corresponding to the vertical field, and generate a structured text description according to the key labels and text attributes;
[0126] Use the pre-trained text-to-image model to process the structured text description to obtain each task main image corresponding to the structured text description;
[0127] Screen each task main image corresponding to each initial image data to obtain each selected image.
[0128] Among them, for the step of screening each task main image corresponding to each initial image data to obtain each selected image, it includes:
[0129] Preliminarily screen each task main image according to the preset clarity threshold, integrity threshold, label matching threshold, and background complexity threshold to obtain each rough-selected image;
[0130] Input each of the roughly selected images and the initial image data into a pre-trained expert model, so that the expert model can accurately screen each roughly selected image according to the initial image data, and output the screened images as each selected image.
[0131] In the above solution, the initial image data can be preprocessed first, such as denoising, duplicate removal, removing blurred, unevenly exposed or low-quality frames, and then normalized, such as adjusting the image size, resolution, and color space, etc., so as to unify the input format. Next, segmentation is achieved by identifying the main object in the initial image data, which can be carried out by pre-establishing a segmentation scheme, so that the obtained first image data is an image data with the theme centered and clearly visible.
[0132] For the key labels and text attributes corresponding to the vertical domain, they often describe the visual concepts and semantic contexts of the image data. Therefore, structured text descriptions can be generated based on them to clarify the requirements of the image data. Then, using a pre-trained text-to-image model (diffusion model or generative adversarial network model) as the basic generation framework, loading domain-related data or fine-tuning the model to adapt to specific tasks. Visual label corresponding conditional prompts, such as color, shape, scene, etc., can also be added to the input data of the text-to-image model. This can guide the text-to-image model to generate synthetic images that only contain the main content or the main body of the task, that is, the task main body images, so as to ensure the consistency between the images and the task requirements.
[0133] However, considering that the task main body images are not yet fine and accurate enough, conduct preliminary screening and accurate screening on each task main body image. For the preliminary screening, task main body images that are lower than any one of the clarity threshold, integrity threshold, label matching threshold, or background complexity threshold can be deleted; subsequently, input the generated images and real images into a pre-trained expert model, which is a Transformer-based model. The expert model sets a distance metric function and a screening threshold to eliminate abnormal images that are irrelevant to the initial image data, and then accurately screens to obtain synthetic image data that is similar to the initial image data, which is used to make up for the common multi-modal data imbalance problem.
[0134] Furthermore, for the process of constructing a mixed dataset from the initial multi-modal data, enhanced text data, and each selected image, it can include the following steps:
[0135] Obtain the prompt text and enhanced text data corresponding to the initial text data;
[0136] For each selected image, associate the prompt text and enhanced text data with this selected image to generate each image-text data pair;
[0137] Authenticity scores are calculated for each image-text data pair corresponding to each selected image to filter out each target data pair;
[0138] A data pair set is constructed from each of the target data pairs, and the data pair set is processed for data balancing to obtain each balanced data pair;
[0139] A mixed data set is constructed from the initial multimodal data, each of the balanced data pairs, and each selected image.
[0140] Specifically, for the prompt text and the enhanced text data, noisy text such as misspelled words and text with irrelevant descriptions can be removed first, so as to retain high-quality descriptive text that meets semantic requirements. Then, the CLIP model (Contrastive Language–Image Pretraining) is used to calculate the similarity between each image-text data pair of the selected image and the prompt text, i.e., the enhanced text data, filter out image-text data pairs with high matching degrees, eliminate image-text data pairs with large semantic deviations, and align the selected image with the prompt text and the enhanced text data in terms of time or scenario.
[0141] Next, the pre-obtained scoring model can be called to perform authenticity scoring and inference on each image-text data pair, calculate authenticity indicators in each image-text data pair, including semantic consistency, label accuracy, and correlation between the image and the text, and set reasonable thresholds to filter out image-text data pairs higher than the thresholds as target data pairs, so as to ensure that the data used in the training and inference processes has high quality.
[0142] In addition to filtering by authenticity scoring, it is also necessary to calculate the difference degree between each target data pair and the entire data pair set or a specific category, measure the diversity of the data pair set, calculate the vector distance between target data pairs (such as using cosine similarity in the embedding space) to evaluate the diversity, measure the distribution of each category in the data pair set, ensure the balance of the number of target data pairs for each category or label, avoid the model being overly biased towards a certain category, and through multiple rounds of screening and enhancement, ensure that the data pair set covers sufficient categories and attributes, and avoid duplicate samples and overfitting.
[0143] Finally, in the process of constructing a mixed dataset from the initial multimodal data, each balanced data pair, and each selected image, the proportions of the initial multimodal data, balanced data pairs, and selected images can be reasonably designed to control the impact on the training of the base large model. The proportion of the initial multimodal data can be set to 30%, and the proportion of the balanced data pairs and selected images can be set to 70% according to the task requirements. Or the proportion can be adjusted according to the specific distribution of the real data. After construction, quality inspection is carried out to filter out noisy and low-quality data, ensuring the overall quality of the mixed dataset, and adjusting the labels or re-calibrating the samples to ensure the effectiveness of the mixed dataset for model training.
[0144] Therefore, this application enables the base large model to adapt to the task requirements in different vertical fields. Through text enhancement and image enhancement processing, the initial text data is automatically converted into instruction text adapted to the vertical field, and the common class imbalance problem is alleviated through text-to-image generation, improving the fast adaptation ability of the base large model in small-sample task scenarios and being applicable to various data structures and data service tasks. Compared with the existing method that only processes a single type of data, this application can extract valuable features from different data sources (such as images, texts, forms, video frames, etc.), perform effective fusion and enhancement, and thus generate synthetic data with higher information richness and stronger generalization ability. This multimodal synthetic data can not only cover a wider range of scenarios and situations but also demonstrate more excellent performance and accuracy in multiple tasks, significantly improving the application effect of the model. In addition, this application generates text data for specific tasks in the vertical field, which includes not only traditional descriptive content but also richer and more complex types such as reasoning and decision support, enabling the model to better understand the task requirements and consider more context and details during execution, thus effectively solving complex logical reasoning scenarios and task requirements that are difficult to handle in the existing technology and improving the efficiency and accuracy of data processing.
[0145] Corresponding to Figure 1 the method described above, the embodiment of the present invention also provides an optimization device for a large model based on small samples, used for Figure 1 the specific implementation of the method in Figure 3 . For the introduction of the optimization device for a large model based on small samples, as Figure 3 shown, the device may include:
[0146] A multimodal data acquisition module 10, configured to acquire initial multimodal data;
[0147] The large language model optimization module 20 is used to determine the vertical domain corresponding to the initial multimodal data, and optimize the pre-obtained third-party large language model based on the vertical domain;
[0148] The model processing module 30 is used to extract the initial text data in the initial multimodal data, input the initial text data into the optimized third-party large language model, and obtain the output enhanced text data;
[0149] The data enhancement processing module 40 is used to extract the initial image data in the initial multimodal data, and perform data enhancement processing on the initial image data to obtain each selected image;
[0150] The mixed dataset construction module 50 is used to construct a mixed dataset from the initial multimodal data, enhanced text data, and each selected image;
[0151] The base large model optimization module 60 is used to optimize the pre-established base large model using the mixed dataset to obtain an optimized multimodal base large model.
[0152] As can be seen from the above technical solution, this application obtains initial multimodal data; determines the vertical domain corresponding to the initial multimodal data, and optimizes the pre-obtained third-party large language model based on the vertical domain; extracts the initial text data in the initial multimodal data, inputs the initial text data into the optimized third-party large language model, and obtains the output enhanced text data; extracts the initial image data in the initial multimodal data, and performs data enhancement processing on the initial image data to obtain each selected image; constructs a mixed dataset from the initial multimodal data, enhanced text data, and each selected image; uses the mixed dataset to optimize the pre-established base large model to obtain an optimized multimodal base large model. This application classifies the initial multimodal data and processes them separately according to the classified data. The third-party large language model can be optimized through the vertical domain of the initial multimodal data, making the third-party large language model more targeted. Then, the optimized third-party large language model can better process the initial text data in the multimodal data to obtain enhanced text data, and then perform data enhancement processing on the initial image data in the multimodal data to obtain selected images. Through the above processing process, the coverage range and situation of the data can be wider, enabling the model to adapt to multimodal service tasks.
[0153] Furthermore, the embodiments of this application provide a large model optimization device based on few-shot learning. Optionally, Figure 4 shows the hardware structure block diagram of the large model optimization device based on few-shot learning. Refer to Figure 4, the hardware structure of the model optimization device may include: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.
[0154] In the embodiments of the present application, the number of the processor 01, the communication interface 02, the memory 03, and the communication bus 04 is at least one, and the processor 01, the communication interface 02, and the memory 03 complete mutual communication through the communication bus 04.
[0155] The processor 01 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.
[0156] The memory 03 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0157] Among them, the memory stores a program, and the processor can call the program stored in the memory. The program is used to execute the following large model optimization method based on small samples, including:
[0158] Obtain initial multimodal data;
[0159] Determine the vertical field corresponding to the initial multimodal data, and optimize the pre-obtained third-party large language model based on the vertical field;
[0160] Extract the initial text data in the initial multimodal data, and input the initial text data into the optimized third-party large language model to obtain the output enhanced text data;
[0161] Extract the initial image data in the initial multimodal data, and perform data enhancement processing on the initial image data to obtain each selected image;
[0162] Construct a mixed dataset from the initial multimodal data, the enhanced text data, and each selected image;
[0163] Use the mixed dataset to optimize the pre-established base large model to obtain an optimized multimodal base large model.
[0164] Optionally, the refined functions and extended functions of the program can refer to the description of the large model optimization method based on small samples in the method embodiments.
[0165] The embodiments of the present application further provide a storage medium, which can store a program suitable for a processor to execute. When the program runs, it controls the device where the storage medium is located to execute the following large model optimization method based on small samples, including:
[0166] Obtain initial multimodal data;
[0167] Determine the vertical domain corresponding to the initial multimodal data, and optimize the pre-obtained third-party large language model based on the vertical domain;
[0168] Extract the initial text data from the initial multimodal data, and input the initial text data into the optimized third-party large language model to obtain the output enhanced text data;
[0169] Extract the initial image data from the initial multimodal data, and perform data enhancement processing on the initial image data to obtain each selected image;
[0170] Construct a mixed dataset from the initial multimodal data, enhanced text data, and each selected image;
[0171] Use the mixed dataset to optimize the pre-established base large model to obtain an optimized multimodal base large model.
[0172] Specifically, the storage medium can be a computer-readable storage medium, and the computer-readable storage medium can be an electronic memory such as a flash memory, EEPROM (electrically erasable programmable read-only memory), EPROM, hard disk, or ROM.
[0173] Optionally, the refined functions and extended functions of the program can refer to the description of the large model optimization method based on small samples in the method embodiments.
[0174] In addition, in each embodiment of the present disclosure, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a live broadcast device, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present disclosure.
[0175] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0176] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0177] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A large model optimization method based on small samples, characterized in that: include: Acquire initial multimodal data; Determine a vertical field corresponding to the initial multimodal data, and optimize a pre-acquired third-party large language model based on the vertical field; Extracting initial text data from the initial multimodal data, and inputting the initial text data into an optimized third-party large language model to obtain output enhanced text data; Extracting initial image data from the initial multimodal data, and performing data enhancement processing on the initial image data to obtain selected images; Constructing a hybrid dataset from the initial multimodal data, the enhanced text data, and each selected image; The pre-established large base model is optimized using the mixed data set to obtain an optimized multi-modal large base model.
2. The method according to claim 1, characterized in that The optimizing the pre-acquired third-party large language model based on the vertical field includes: Establish initial label text that matches the vertical field; Expanding the initial label text to obtain specific description information; Constructing a loss function according to the specific description information; The loss function is used to optimize the soft prompts of the pre-acquired third-party large language model.
3. The method according to claim 1, characterized in that The process of processing the initial text data with the optimized third-party large language model and outputting enhanced text data includes: Establish target questions based on the verticals; Decompose the target problem to obtain the text of each step; Constructing a logical reasoning process from the texts of each of the steps; Generate corresponding prompt text for each step text; Add the prompt texts to the logical reasoning process according to the corresponding step texts to obtain a thinking chain; The input initial text data is inferred according to the thought chain to obtain enhanced text data and output it.
4. The method according to claim 3, characterized in that The input initial text data is inferred according to the thought chain to obtain enhanced text data and output it, including: Acquire domain-specific knowledge corresponding to the vertical field; Reasoning the initial text data according to the thought chain; While reasoning, the output answer corresponding to each step text is extracted in real time; For each step text, adjust the prompt text corresponding to the next step text according to its corresponding output answer; While adjusting, the domain specific knowledge is added to the prompt text to infer enhanced text data corresponding to the initial text data.
5. The method according to claim 1, characterized in that The processing and enhancement of the initial image data to obtain each selected image includes: Identify the main object in the initial image data, and determine the position, outline and category of the main object; Perform segmentation processing on the initial image data based on the position, contour and category to obtain first image data; Determine key tags and text attributes corresponding to the vertical field, and generate structured text descriptions according to the key tags and text attributes; Processing the structured text description using a pre-trained text graph model to obtain task subject images corresponding to the structured text description; The task subject images corresponding to the initial image data are screened to obtain selected images.
6. The method according to claim 5, characterized in that The step of screening each task subject image corresponding to each initial image data to obtain each selected image includes: Preliminarily screening each of the task subject images according to a preset clarity threshold, integrity threshold, label matching threshold, and background complexity threshold to obtain each roughly selected image; Each of the rough selected images and the initial image data is input into a pre-trained expert model, so that the expert model can accurately screen each of the rough selected images according to the initial image data, and output the screened images as each of the selected images.
7. The method according to any one of claims 1 to 6, characterized in that: The hybrid data set is constructed by using the initial multimodal data, the enhanced text data and each selected image, including: Acquire prompt text and enhanced text data corresponding to the initial text data; For each selected image, the prompt text and enhanced text data are associated with the selected image to generate each image-text data pair; Perform authenticity scoring on each image-text data pair corresponding to each selected image to screen out each target data pair; Constructing a data pair set from each of the target data pairs, and performing data balancing processing on the data pair set to obtain each balanced data pair; A mixed dataset is constructed from the initial multimodal data, each of the balanced data pairs, and each of the selected images.
8. A large model optimization device based on small samples, characterized in that: include: A multimodal data acquisition module, used to acquire initial multimodal data; A large language model optimization module, used to determine a vertical field corresponding to the initial multimodal data, and optimize a pre-acquired third-party large language model based on the vertical field; A model processing module, used for extracting initial text data from the initial multimodal data, inputting the initial text data into an optimized third-party large language model, and obtaining output enhanced text data; A data enhancement processing module, used for extracting initial image data from the initial multimodal data, and performing data enhancement processing on the initial image data to obtain each selected image; A hybrid dataset construction module, used to construct a hybrid dataset from the initial multimodal data, the enhanced text data and each selected image; The base large model optimization module is used to optimize the pre-established base large model using the mixed data set to obtain an optimized multi-modal base large model.
9. A large model optimization device based on small samples, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the large model optimization method based on small samples as described in any one of claims 1-7.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the large model optimization method based on small samples as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Construction and training method of electric power vision multi-granularity pre-training large model
CN115240075A
Prearranged plan text extraction method and device and storage medium
CN116340532A
Image generation model training method and device, equipment and storage medium
CN116721334A
Image generation method and device, electronic equipment and storage medium
CN117132456A
Systems and methods for analyzing text extracted from images and performing appropriate transformations on the extracted text
US12033620B1