Visual language model fine tuning method and device, computer equipment and storage medium

By generating and verifying multiple candidate answers, the visual language model is updated, and the weak performance of large visual language models in learning with few samples is solved, and the application effect in the financial and medical fields is improved.

CN120508788APending Publication Date: 2025-08-19PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510602846.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing large-scale visual language models have weak performance and insufficient generalization capabilities in learning in small samples, making it difficult to effectively utilize limited data in the financial and medical fields, resulting in problems such as misjudgment of credit risk, investment decision-making errors, and clinical diagnosis deviations.

Method used

By obtaining training data from the target field, initializing the pre-trained visual language model, generating multiple candidate answers, and updating the model based on the verification results to mine cross-modal feature associations and improve model performance and generalization capabilities.

Benefits of technology

It improves the performance and generalization ability of the model under the conditions of few samples to meet the needs of high-quality models in the financial and medical fields, and reduces the probability of risk misjudgment and diagnostic bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508788A_ABST
    Figure CN120508788A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, financial science and technology and digital medical treatment, discloses a visual language model fine tuning method and device, computer equipment and a storage medium, and can optimize market trend prediction and medical record analysis application. The method comprises the steps of obtaining training data of a target field; initializing the pre-training visual language model as a basic model; generating a plurality of candidate answers of the training data through the basic model; and verifying the plurality of candidate answers and updating the basic model according to a verification result to obtain a target model. According to the method, the plurality of candidate answers of the training data are output through the basic model, and then the basic model is updated based on the verification result, so that potential cross-modal feature association in the data can be mined by utilizing the plurality of candidate answers, limited data can be more fully utilized, the performance and generalization ability of the model in few-sample learning are improved, and the learning efficiency is improved. And the application requirements of the financial and medical fields on the high-quality model are practically met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence, financial technology, and digital medicine, and in particular to a method, apparatus, computer equipment, and storage medium for fine-tuning a visual language model. Background Art

[0002] Visual language models are a type of multimodal artificial intelligence model that can jointly process visual data such as images, videos, medical images, charts, etc., and language data such as text, instructions, annotations, natural language descriptions, etc. They are widely used in intelligent scenarios such as image-text translation and visual question answering.

[0003] Existing large-scale vision-language models (LVLMs) primarily rely on supervised fine-tuning (SFT) for model fine-tuning. This method primarily uses a single annotation result from a single sample for verification and update. This makes it difficult to exploit cross-modal feature associations in limited data, such as the deep semantic mapping between visual images and text reports. This results in weak performance and insufficient generalization in few-shot learning. In practical applications, particularly in finance and healthcare, high-quality annotated data is often extremely scarce, making it even more difficult for vision-language models in these fields to achieve the required high performance and generalization.

[0004] In the financial sector, weak model performance and insufficient generalization capabilities can lead to misjudgments of credit risk. For example, when faced with chart data or unstructured text reports for new financial products, traditional fine-tuned models may fail to learn the trend fluctuation patterns implicit in historical data, leading to investment decision-making errors or systemic risks. In the medical field, weak model performance and insufficient generalization capabilities can lead to clinical diagnostic bias, such as missed detection or misdiagnosis of tiny lesions in CT images. The root cause is that traditional methods are unable to learn the association logic between fine-grained features such as the texture and boundaries of the lesions and the pathology text guidelines through limited labeled data, resulting in incomplete extraction of key medical features. In severe cases, this can delay treatment and threaten the patient's life.

[0005] Therefore, there is an urgent need for a new visual language model fine-tuning method that can efficiently utilize small-sample data and improve generalization capabilities, so as to solve the practical application difficulties of existing fine-tuning methods in the financial and medical fields. Summary of the Invention

[0006] The present invention provides a visual language model fine-tuning method, apparatus, computer equipment, and storage medium to address the technical problems of weak performance and insufficient generalization ability of existing large-scale visual language models in few-sample learning.

[0007] First, a method for fine-tuning a visual language model is provided, including:

[0008] Obtain training data for the target domain;

[0009] Initialize the pre-trained visual language model as the base model;

[0010] generating a plurality of candidate answers for the training data using the basic model;

[0011] Verify multiple candidate answers and update the base model according to the verification results to obtain a target model.

[0012] Secondly, a visual language model fine-tuning device is provided, including:

[0013] Acquisition module, used to obtain training data in the target domain;

[0014] Initialization module, used to initialize the pre-trained visual language model as the base model;

[0015] a processing module, configured to generate a plurality of candidate answers for the training data using the basic model;

[0016] An updating module is used to verify multiple candidate answers and update the basic model according to the verification results to obtain a target model.

[0017] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned visual language model fine-tuning method when executing the computer program.

[0018] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned visual language model fine-tuning method are implemented.

[0019] The beneficial effects of the present invention compared with the existing technology are: the present invention outputs multiple candidate answers for training data through a basic model, and then updates the basic model based on the verification results, which helps to use multiple candidate answers to explore potential cross-modal feature associations in the data. Compared with supervised fine-tuning that only relies on a single sample and a single annotation result, it can make more effective use of limited data, break the limitations of traditional methods on data, improve the performance and generalization ability of the model in few-sample learning, and effectively meet the application needs of the financial and medical fields for high-quality models.

[0020] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In addition, in order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 2 is a schematic diagram of an application environment of a visual language model fine-tuning method according to an embodiment of the present invention;

[0022] Figure 2 1 is a flow chart of a method for fine-tuning a visual language model according to an embodiment of the present invention;

[0023] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S10;

[0024] Figure 4 yes Figure 3 A schematic flow chart of a specific implementation of step S30;

[0025] Figure 5 yes Figure 4 A schematic flow chart of a specific implementation of step S40;

[0026] Figure 6 yes Figure 5 A schematic flow chart of a specific implementation of step S41;

[0027] Figure 7 yes Figure 5 A schematic flow chart of a specific implementation of step S45;

[0028] Figure 8 1 is a schematic structural diagram of a visual language model fine-tuning device according to an embodiment of the present invention;

[0029] Figure 9 is a structural diagram of a computer device in one embodiment of the present invention;

[0030] Figure 10 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0032] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0033] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0034] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0035] See also Figure 1 and Figure 2 , Figure 1 Schematic diagram of the application scenario of the visual language model fine-tuning method provided in an embodiment of the present invention. The client can communicate with the server through a network; the server can obtain training data of the target domain stored from the client or the server; and initialize the pre-trained visual language model as a basic model; generate multiple candidate answers for the training data through the basic model; verify the multiple candidate answers and update the basic model according to the verification results to obtain a target model. The client provides an interactive platform for the user, so that the user can process the visual language through the visual language model or target model carried by the server. For example, the user sends data such as images and texts to the target model of the server through the client, and the target model returns the processing results to the client through the server. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented as an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0036] See also Figure 2 As shown, Figure 2 A schematic flow chart of a visual language model fine-tuning method provided in an embodiment of the present invention. The visual language model fine-tuning method includes the following steps:

[0037] S10: Obtain training data in the target domain.

[0038] Step S10 provides data support for subsequent model fine-tuning. Only by obtaining data that meets the target domain can the visual language model be learned and optimized in a targeted manner.

[0039] It is understandable that the training data can be obtained from the client or from the server storage module. The visual language model fine-tuning method of this embodiment is used to fine-tune and strengthen the visual language model. In the financial field, the training data can be visual data and language data such as financial market trend graphics and customer transaction data icons; financial institutions can use the fine-tuned or strengthened visual language model to analyze multimodal data such as market trends and customer behavior, providing a scientific basis for risk assessment and investment decisions; they can also use the visual language model to process text reports, numerical data and charts and other information to predict market fluctuations and credit risks and assist in formulating investment strategies; they can also detect fraud through the visual language model, and identify abnormal patterns and potential fraudulent activities by analyzing transaction data, user behavior and related text information, thereby improving the security of financial transactions; they can also use the visual language model to analyze text data such as customer inquiries and complaints and multimodal information such as customer operation videos to accurately understand customer needs, improve service quality and efficiency, and enhance customer satisfaction. In the medical field, training data can be visual data and language data such as CT, MRI imaging data, and medical records; fine-tuned or enhanced visual language models can assist doctors in analyzing medical images such as CT and MRI, accurately locating lesions and identifying abnormal tissues, and improving diagnostic accuracy and efficiency. Especially in the scenario of few-sample learning, it can effectively alleviate the problem of scarce labeled data; it can also analyze patients' multimodal data, predict disease risks and progression, provide support for early intervention and prevention, and help promote the development of personalized medicine and disease management.

[0040] In some embodiments of the present invention, Figure 3 As shown, a specific training data acquisition solution is provided, in S10, that is, obtaining training data of the target field, which specifically includes the following steps S11-S13.

[0041] S11: Acquire target data of a target domain, where the target data includes visual data and / or language data.

[0042] For step S11, the target data includes visual data and language data. The acquisition of these multimodal data broadens the breadth and depth of model learning. The extensive collection of data of different modalities in the target field enables the model to access multi-dimensional information, which helps to learn cross-modal features.

[0043] More specifically, visual data refers to images, videos, graphical structured data, and the like generated by sensors such as cameras, scanners, or visualization tools; text data refers to unstructured or semi-structured data composed of natural language symbols, which can be manually entered or generated by natural language processing (NLP) models. Most target data includes both visual data and language data associated with visual data, such as videos and subtitles within them. In the financial sector, visual data can include stock candlestick charts, heat maps, capital flow diagrams, check / contract scans, ATM surveillance videos, and facial / fingerprint biometric images; text data can include listed company financial statements, financial news text, customer risk assessment questionnaires, credit approval opinion records, and compliance review reports. Target data that is a mixture of visual and textual data can include bank statements, robo-advisory reports, and digital RMB wallet interfaces. Bank statements include scanned images and table text; robo-advisory reports include charts and analytical text; and digital RMB wallet interfaces include the user interface and transaction record text. In the medical field, visual data may include X-rays, CT, MRI medical images, endoscopic surgery videos, pathological section microscopic images, skin lesion photos, drug packaging barcodes, and QR codes; text data may include electronic health records, doctors' handwritten prescriptions, drug instructions, scientific research papers, and patient consultation records; target data mixed with visual data and text data may include annotated medical imaging reports, screenshots of smart medical record systems, and medical bills, where annotated medical imaging reports include images and diagnostic conclusions, screenshots of smart medical record systems include structured charts and descriptions of the condition, and medical bills include images of examination sheets and detailed expense text.

[0044] S12: Cleaning and labeling the target data.

[0045] S13: Generate training data based on the cleaning and labeling results.

[0046] For steps S12-S13, the target data is cleaned to remove noise and erroneous information in the data, and the data is labeled so that the model can understand the meaning of the data and learn the correct pattern. The cleaning and labeling process ensures the quality of the training data, and the cleaned and labeled target data is organized into a format suitable for model training so that it can be effectively used by the model.

[0047] It is understandable that financial data cleaning can eliminate abnormal transaction data, mark the direction of market trends, and allow the model to accurately learn the rules. The reasonable generation of financial training data can help the model accurately predict market fluctuations and credit risks; medical data cleaning can remove image artifacts, mark the location and nature of lesions, and improve the accuracy of model diagnosis. The standardized generation of medical training data can enable the model to accurately locate lesions and identify abnormal tissues, which can improve the accuracy of medical data analysis and lay a good foundation for subsequent model training.

[0048] S20: Initialize the pre-trained visual language model as the base model.

[0049] The basic model of this embodiment uses a pre-trained visual language model as the initial model and performs initialization operations on the pre-trained visual language model. By leveraging the general knowledge and features already learned by the pre-trained visual language model, the cost of learning the model from scratch is reduced, which helps speed up training. In the financial and medical fields, pre-trained visual language models can quickly adapt to data in specific fields. For example, financial models use pre-trained language understanding capabilities to analyze text reports, and medical models use pre-trained image feature recognition capabilities to process medical images, significantly improving training efficiency.

[0050] S30: Generate multiple candidate answers for the training data using the basic model.

[0051] For step S30, the visual language model of this embodiment fully utilizes the model's generative capabilities to obtain multiple possible results, increasing the diversity of model learning and facilitating the discovery of potential correct answers. Among them, the candidate answers of this embodiment are candidate answers that retain the reasoning process, effectively ensuring the availability of candidate answers. It is understandable that when a financial model analyzes market data, multiple candidate answers can cover different investment strategy recommendations; when a medical model diagnoses, it can provide multiple possible lesion judgments, providing more references and improving the comprehensiveness and accuracy of the visual language model.

[0052] In some embodiments of the present invention, Figure 4 As shown, a specific candidate answer generation scheme is provided, in S30, that is, multiple candidate answers for the training data are generated by the basic model, which specifically includes the following steps S31-S34.

[0053] S31: Obtain a prompt word set and prompt word additional rules.

[0054] S32: Add prompt words to the training data according to the prompt word adding rule.

[0055] For steps S31-S32, the prompt word set includes several prompt words. Prompt words are instructions or questions input by the user into the model, guiding the model to generate answers with specific goals, formats, or fields. This helps align user intent with model output, significantly improving the accuracy and practicality of model results. Prompt words provide guidance for the model to generate candidate answers, making the generated answers more consistent with task requirements and domain characteristics. Combining prompt words with training data allows the model to think based on the prompt words when generating answers, improving the relevance and accuracy of candidate answers.

[0056] S33: Inputting the training data and the prompt words attached to the training data into the basic model.

[0057] S34: Loading the basic model, wherein the basic model outputs a plurality of candidate answers for the training data according to the prompt word.

[0058] For steps S33-S34, the basic model receives complete input information and outputs multiple possible results based on the training data and prompt words for subsequent verification and selection.

[0059] For example, in the application of financial investment decision-making, the prompt word may be "Based on the Q2 2024 financial report, list the net profit growth rate comparison of companies A, B and C", so as to provide multiple inferred candidate answers for risk assessment and investment decision-making in an intelligent and efficient manner; for example, in the application of digital medicine, the prompt word may be "Combined with the patient's skin lesion photos and symptom descriptions, infer the possible type of skin disease", or "Based on MRI images and electronic medical records, generate a preliminary diagnosis report for brain tumors", so as to provide medical staff with multiple inferred candidate answers in an intelligent and efficient manner.

[0060] S40: Verify multiple candidate answers and update the basic model according to the verification results to obtain a target model.

[0061] Verify the quality of candidate answers and optimize the model based on the verification results, moving it toward high performance and generalization. In the financial and medical fields, verification and updating can improve the accuracy of visual language model predictions and diagnoses, effectively reducing the probability of misjudgment and diagnostic bias.

[0062] In some embodiments of the present invention, Figure 5 As shown, a specific model verification and update solution is provided, in S40, that is, verifying multiple candidate answers and updating the basic model according to the verification results, which specifically includes the following steps S41-S45.

[0063] S41: Construct reward function.

[0064] The reward function is a mathematical function used to quantitatively evaluate the degree of match between the candidate answers generated by the model and the expected results. It plays a key guiding role in the fine-tuning process of the visual language model of this embodiment. For the multiple candidate answers generated, the reward function can assign a numerical value, namely the reward value, to each answer based on the task type and domain characteristics to measure the quality of the answer. This value can intuitively reflect the usefulness and accuracy of the candidate answer in completing the current visual task. By constructing a suitable reward function, a clear optimization direction is provided for the model, so that the model understands what kind of answer can obtain higher rewards, and then continuously adjusts its own strategy in this direction in subsequent training to optimize the performance of the model.

[0065] In some embodiments of the present invention, Figure 6 As shown, a specific reward function construction scheme is provided, in S41, that is, constructing the reward function, which specifically includes the following steps S411-S412.

[0066] S411: Obtain the visual task type of the training data.

[0067] S412: Constructing a reward function according to the visual task type.

[0068] In the fields of finance and medicine, different visual tasks, such as financial chart analysis and medical image diagnosis, have corresponding reward functions that differ. Targeted reward functions can better optimize the model's performance on specific tasks.

[0069] For example, in the target detection task, the IoU value of the bounding box is used as a reward to construct a reward function. The specific reward function can be R base =IoU(B pred ,B gt ), prediction box B pred is the bounding box coordinate predicted by the model, B gt The coordinates of the bounding box are the real ones. By comparing their intersection over union (IoU), we can obtain multiple data with high correlation and use them as candidate answers. In the image classification task, the classification accuracy is used as a reward, and the data with an accuracy greater than the first threshold can be used as a candidate answer.

[0070] S42: Quantitatively score the candidate answers in sequence based on the reward function to obtain multiple initial reward values.

[0071] S43: Normalize all initial reward values to obtain a target reward value.

[0072] For steps S42-S43, normalization is a data preprocessing technique that maps the initial reward value to a specific interval, such as [0,1] or [-1,1]. In the fine-tuning process of the visual language model of this embodiment, the initial reward values calculated by the reward function for different candidate answers may have different scales and distributions. If these initial reward values are directly fed back to the model for learning, the model may be overly sensitive to larger reward values during learning due to the large difference in numerical values, and ignore the information contained in smaller reward values. Through normalization, all initial reward values are unified within a standard range, so that the model can treat the reward feedback of each candidate answer fairly, avoid learning bias due to differences in the scale of the reward values, and more effectively learn the value of different candidate answers, optimize its own strategy to improve performance.

[0073] In finance, when models analyze various financial data to generate different investment strategy recommendations and obtain initial reward values, these reward values may vary significantly depending on the measurement indicators, such as some based on the amount of return and others on risk control indicators. After normalization, the model can learn from these reward values more evenly, comprehensively consider multiple factors to optimize investment strategies, avoid excessive focus on a single type of reward, and improve the rationality and stability of financial decision-making. In the medical field, normalizing the initial reward values of different candidate answers in diagnostic tasks, such as reward values based on the diagnostic accuracy of different diseases and the accuracy of lesion localization, allows the model to more objectively evaluate the value of each diagnostic result, avoid neglecting other important factors due to excessively high reward values for a single indicator, and thus improve the comprehensiveness and accuracy of medical diagnoses and reduce the risk of misdiagnosis.

[0074] S44: Outputting the target reward value and the candidate answer corresponding to the target reward value to the reinforcement learning algorithm.

[0075] S45: Updating the strategy parameters of the basic model through the reinforcement learning algorithm to obtain a target model.

[0076] In steps S44-S45, the evaluation results are fed back to the reinforcement learning algorithm to provide data support for model strategy updates. In the financial and medical fields, the model can adjust its strategy based on the reward value, making it more inclined to generate high-reward answers in the future, thereby improving the model's adaptability and accuracy.

[0077] For example, in an application of visual language models to assist financial investment decision-making, investment recommendations are provided based on market chart data and text reports. The following three candidate answers and their corresponding initial reward values are obtained: Candidate A has an expected annualized return of 15% and a volatility of 20%, with an initial reward of 80 points; Candidate B has an expected annualized return of 10% and a volatility of 15%, with an initial reward of 60 points; and Candidate C has an expected annualized return of 20% and a volatility of 30%, with an initial reward of 70 points. The sum of all initial rewards is 210 points. Candidates A, B, and C, and their initial rewards, are normalized to give a target reward of 80÷210≈0.38 for Candidate A, 60÷210≈0.29 for Candidate B, and 70÷210≈0.33 for Candidate C. These target rewards and the corresponding candidate answers are fed into a reinforcement learning algorithm, which adjusts the model's policy parameters based on the rewards, making it more likely that the model will generate similarly high-reward investment recommendations in the future.

[0078] For example, in an application where a visual language model assists in lung CT image diagnosis, three diagnostic results are generated as candidate answers. Candidate D accurately identifies pneumonia and provides a detailed description, earning an initial reward of 90 points; Candidate E correctly identifies the disease but doesn't specify its type, earning an initial reward of 30 points; and Candidate F incorrectly identifies the disease and earning an initial reward of 10 points. The total initial reward is 130 points, with the target reward for Candidate D being 90÷130≈0.69, Candidate E being 30÷130≈0.23, and Candidate F being 10÷130≈0.08. The target reward values and corresponding results are fed into a reinforcement learning algorithm, which adjusts the model's policy parameters to make it more likely to generate accurate and detailed diagnostic results.

[0079] In some embodiments of the present invention, Figure 7 As shown, a specific basic model updating solution is provided. In S45, the strategy parameters of the basic model are updated by the reinforcement learning algorithm to obtain the target model, which specifically includes the following steps S451-S452.

[0080] S451: Adopting a policy gradient algorithm as the reinforcement learning algorithm, and calculating the policy gradient of the basic model based on the target reward value.

[0081] Policy gradient algorithms are an important class of algorithms in reinforcement learning, used to optimize an agent's policy. A policy can be understood as the way an agent chooses actions in different states. In this example, the agent is a visual language model. In reinforcement learning, an agent interacts with its environment, takes actions, and receives rewards in return. The goal is to learn an optimal policy that maximizes long-term cumulative rewards. Policy gradient algorithms achieve this by directly optimizing the policy. They do not rely on an estimated value function, but instead directly adjust the policy parameters. Specifically, the policy is typically represented by a parameterized function, such as a neural network, whose parameters determine the probability distribution of each action in different states. The core idea of the policy gradient algorithm is to calculate the gradient of the policy parameters. This gradient reflects the direction and extent of the impact of small changes in the policy parameters on the cumulative reward. By updating the policy parameters along this gradient, the performance of the policy can be gradually improved, allowing the agent to obtain higher rewards in future interactions.

[0082] In step S451, the policy gradient algorithm analyzes the target reward value for each candidate answer and determines how the current model policy, i.e., the method used to generate the candidate answer, affects the reward. If a candidate answer receives a high target reward value, the algorithm calculates the policy parameter gradient that increases the probability of generating that answer. Conversely, if a candidate answer receives a low target reward value, the algorithm calculates the gradient that decreases the probability of generating that answer.

[0083] For example, in finance, if an investment recommendation receives a high reward, the algorithm will calculate how to adjust the model parameters, making it more likely to generate similar high-quality recommendations when encountering similar financial data in the future. For example, in medical imaging diagnosis, when processing a lung CT image, the model generates several diagnostic results, each with a corresponding target reward. Based on these rewards, the policy gradient algorithm calculates how to adjust the model parameters, making it more likely to generate a diagnosis with a high reward the next time a similar CT image is processed.

[0084] S452: Adjust the policy parameters of the basic model according to the policy gradient.

[0085] The policy parameters of the base model are adjusted based on the calculated policy gradient. Specifically, the policy gradient indicates the direction of increasing reward values. By adjusting the parameters along this direction, the model can continuously optimize its own strategy.

[0086] After step S4, the method further includes: iteratively training the target model to obtain a fine-tuned visual language model.

[0087] Repeat steps S30-S40, using the latest target model as the base model for training after each iteration. Each round of iterative training updates the model parameters based on the results of the previous round of training. The latest model has already learned the features and patterns in the previous training data and accumulated a certain amount of experience. Using it as the base model for continued training can further optimize based on previous learning results and gradually improve model performance. Repeating the training process multiple times continuously optimizes the model, gradually improving and stabilizing its performance. After repeated adjustments, the model can learn better strategies and generate more consistent and rewarding answers when faced with new visual and language data, thereby improving application performance in fields such as finance and healthcare.

[0088] In some embodiments of the present invention, before step S10, the following step may be included: constructing a multi-task learning framework to input training data of multiple related visual tasks into the basic model.

[0089] In specific implementation, independent reward functions are designed for multiple related visual tasks, and the expected rewards of multiple tasks are simultaneously optimized through the reinforcement learning algorithm.

[0090] In this embodiment, the multi-task learning framework enables the model to simultaneously learn multiple related visual tasks, meaning that different tasks can share underlying feature representations. For example, in finance and healthcare, tasks such as analyzing market trends and detecting fraud may both require extracting underlying features from the data, such as time series changes and abnormal patterns. Through multi-task learning, the model can share these common underlying feature extraction capabilities when learning these different tasks, reducing duplication and improving learning efficiency. In healthcare, tasks such as analyzing CT images and predicting disease risk may both involve extracting underlying features such as tissue texture and structure from images. The multi-task learning framework allows the model to share the learning results of these common features, reducing training costs. Simultaneously learning multiple tasks helps the model capture a wider range of patterns and regularities, thereby enhancing its generalization ability. Models trained on a single task may only learn patterns specific to that specific task, resulting in poor performance when faced with new, unseen data or tasks. In a multi-task learning framework, the model is trained on multiple tasks, exposing it to a more diverse range of samples and patterns. In practical applications, especially in finance and healthcare, high-quality annotated data is often extremely scarce. Multi-task learning frameworks can fully utilize limited data resources and improve data utilization efficiency by sharing data across multiple tasks. For example, some financial data may contain information related to both market trend analysis and customer credit assessment. Multi-task learning frameworks allow models to simultaneously learn knowledge for different tasks from this data, making limited data more effective. In the healthcare field, a patient's medical record data may contain both imaging information for disease diagnosis and basic patient information and medical history for disease risk prediction. Multi-task learning frameworks enable models to simultaneously learn relevant knowledge for multiple tasks from this data, alleviating the problem of data scarcity.

[0091] It can be seen that in the above scheme, the method of outputting multiple candidate answers for training data through the basic model and then updating the basic model based on the verification results can help to use multiple candidate answers to explore potential cross-modal feature associations in the data. Compared with supervised fine-tuning that only relies on a single sample and a single annotation result, it can make more full use of limited data, break the limitations of traditional methods on data, improve the performance and generalization ability of the model in few-sample learning, and effectively meet the application needs of finance and medical fields for high-quality models.

[0092] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0093] In one embodiment, the present invention provides a visual language model fine-tuning device, which corresponds to the visual language model fine-tuning method in the above embodiment. Figure 8 As shown, the visual language model fine-tuning device includes an acquisition module 101, an initialization module 102, a processing module 103 and an update module 104. The functional modules are described in detail as follows:

[0094] An acquisition module 101 is used to acquire training data of a target domain;

[0095] Initialization module 102, used to initialize the pre-trained visual language model as a base model;

[0096] A processing module 103 is configured to generate a plurality of candidate answers for the training data using the basic model;

[0097] The updating module 104 is configured to verify multiple candidate answers and update the basic model according to the verification results to obtain a target model.

[0098] In one embodiment, the acquisition module 101 is specifically configured to:

[0099] Acquiring target data of a target domain, wherein the target data includes visual data and / or language data;

[0100] Cleaning and labeling the target data;

[0101] Generate training data based on the cleaning and labeling results.

[0102] In one embodiment, the processing module 103 is specifically configured to:

[0103] Obtaining a set of prompt words and additional rules for prompt words;

[0104] Adding prompt words to the training data according to the prompt word addition rule;

[0105] Inputting the training data and the prompt words attached to the training data into the basic model;

[0106] The basic model is loaded, and the basic model outputs multiple candidate answers for the training data according to the prompt word.

[0107] In one embodiment, the update module 104 is specifically configured to:

[0108] Constructing a reward function;

[0109] Quantitatively scoring the candidate answers in sequence based on the reward function to obtain multiple initial reward values;

[0110] Normalize all initial reward values to get the target reward value;

[0111] Outputting the target reward value and the candidate answer corresponding to the target reward value to a reinforcement learning algorithm;

[0112] The strategy parameters of the basic model are updated through the reinforcement learning algorithm to obtain the target model.

[0113] For the specific definition of the visual language model fine-tuning device, please refer to the definition of the visual language model fine-tuning method above, which will not be repeated here. The various modules in the above-mentioned visual language model fine-tuning device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0114] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side method for fine-tuning a visual language model.

[0115] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a visual language model fine-tuning method.

[0116] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0117] Obtain training data for the target domain;

[0118] Initialize the pre-trained visual language model as the base model;

[0119] generating a plurality of candidate answers for the training data using the basic model;

[0120] Verify multiple candidate answers and update the base model according to the verification results to obtain a target model.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0122] Obtain training data for the target domain;

[0123] Initialize the pre-trained visual language model as the base model;

[0124] generating a plurality of candidate answers for the training data using the basic model;

[0125] Verify multiple candidate answers and update the base model according to the verification results to obtain a target model.

[0126] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant description in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0127] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0128] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0129] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A visual language model fine-tuning method, characterized in that: include: Obtain training data for the target domain; Initialize the pre-trained visual language model as the base model; generating a plurality of candidate answers for the training data using the basic model; Verify multiple candidate answers and update the base model according to the verification results to obtain a target model.

2. The visual language model fine-tuning method according to claim 1, characterized in that: The step of obtaining a training dataset for a target domain includes: Acquiring target data of a target domain, wherein the target data includes visual data and / or language data; Cleaning and labeling the target data; Generate training data based on the cleaning and labeling results.

3. The visual language model fine-tuning method according to claim 2, characterized in that: Generating a plurality of candidate answers for the training data using the basic model includes: Obtaining a set of prompt words and additional rules for prompt words; Adding prompt words to the training data according to the prompt word addition rule; Inputting the training data and the prompt words attached to the training data into the basic model; The basic model is loaded, and the basic model outputs multiple candidate answers for the training data according to the prompt word.

4. The visual language model fine-tuning method according to claim 3, characterized in that: The verifying the plurality of candidate answers and updating the basic model according to the verification results includes: Constructing a reward function; Quantitatively scoring the candidate answers in sequence based on the reward function to obtain multiple initial reward values; Normalize all initial reward values to get the target reward value; Outputting the target reward value and the candidate answer corresponding to the target reward value to a reinforcement learning algorithm; The strategy parameters of the basic model are updated through the reinforcement learning algorithm to obtain the target model.

5. The visual language model fine-tuning method according to claim 4, characterized in that: The constructing of the reward function includes: Obtaining a visual task type of the training data; Construct a reward function based on the type of vision task.

6. The visual language model fine-tuning method according to claim 5, characterized in that: Updating the strategy parameters of the base model by the reinforcement learning algorithm to obtain a target model includes: Adopting a policy gradient algorithm as the reinforcement learning algorithm, and calculating the policy gradient of the basic model based on the target reward value; Adjust the policy parameters of the base model according to the policy gradient.

7. The visual language model fine-tuning method according to claim 1, characterized in that: After verifying the plurality of candidate answers and updating the base model according to the verification results to obtain the target model, the method further includes: The target model is iteratively trained to obtain a fine-tuned visual language model.

8. A visual language model fine-tuning device, characterized in that: include: Acquisition module, used to obtain training data in the target domain; Initialization module, used to initialize the pre-trained visual language model as the base model; a processing module, configured to generate a plurality of candidate answers for the training data using the basic model; An updating module is used to verify multiple candidate answers and update the basic model according to the verification results to obtain a target model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the visual language model fine-tuning method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the visual language model fine-tuning method according to any one of claims 1 to 7 are implemented.