Multi-modal large model supervision-based cue word optimization training method and system

Through collaborative training and dual supervision optimization of prompt words of multimodal large models, the problem of inefficient prompt words optimization in port navigation monitoring is solved, automated and high-precision prompt word generation is realized, and the adaptability and accuracy of multimodal large models are improved.

CN120297423AActive Publication Date: 2025-07-11ZHEJIANG WHYIS TECH CO LTD

Patent Information

Application Number
CN202510776516.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

In the existing port navigation monitoring, traditional target detection models have problems such as category limitations and insufficient information fusion capabilities. The existing prompt word optimization methods rely on manual design, are inefficient and lack automated and efficient optimization methods in complex multimodal scenarios.

Method used

The prompt word optimization training method based on multimodal large models is adopted. By constructing a training set, generating initial prompt words, filtering and optimization processes, combining the collaborative training of complex and simple large models, dual supervision and feature fusion are introduced to optimize prompt words to improve accuracy and generalization capabilities.

Benefits of technology

The level of prompt word generation and optimization automation in port navigation monitoring tasks has been significantly improved, the degree of manual participation has been reduced, and the comprehensive reasoning ability of multimodal large models has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297423A_ABST
    Figure CN120297423A_ABST
Patent Text Reader

Abstract

The invention discloses a cue word optimization training method and system based on multi-modal large model supervision. The method comprises the following steps: screening and optimizing a plurality of initial cue words by adopting a simple multi-modal large model, a complex multi-modal large model and a first text large model to obtain a simple optimal cue word and a complex optimal cue word; inputting the simple optimal cue word, the training set and the manual calibration problem into a simple multi-modal large model for training to obtain a simple total loss value; judging whether the total simple loss value is within a preset threshold range, and if yes, ending training; and otherwise, repeating the optimization of the process until the simple total loss value is within the preset threshold range, stopping training, and taking the simple optimal cue word obtained by the last update as the target optimal cue word. According to the method, manual intervention is reduced through cooperative training of a complex large model and a simple large model; manual calibration data and the reasoning process of a complex large model are combined, and the cue word accuracy is improved through double supervision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and multimodal large models. Specifically, it relates to a method and system for optimizing and training prompts based on multimodal large model supervision. Background Art

[0002] Existing port navigation monitoring mainly relies on traditional object detection models for object recognition, which has problems such as category limitations and insufficient information fusion capabilities. Multimodal large models have the ability to fuse vision and semantics, can perform open-set recognition, and adapt to the complex environment of ports. However, due to the fact that the inference performance of multimodal large models highly depends on the quality of prompts, existing prompt optimization usually relies on manual design, which has problems such as understanding deviations and low efficiency. Although attempts to automatically optimize prompts have emerged in recent years, they are single-modal optimizations based on reinforcement learning. In complex multimodal scenarios, especially in port monitoring applications, there is no mature method to achieve efficient automation of prompt optimization. Therefore, there is an urgent need for an automated and highly accurate multimodal prompt optimization method. Summary of the Invention

[0003] In an embodiment of the present invention, a method and system for optimizing and training prompts based on multimodal large model supervision are provided to solve the problems of understanding deviations and low efficiency existing in the prior art when relying on manual design of prompts.

[0004] To achieve the above object, on the one hand, the present invention provides a method for optimizing and training prompts based on multimodal large model supervision, the method comprising: S1, constructing a training set including artificially calibrated problems, artificial analysis processes, and artificial answers; S2, using the training set with a complex multimodal large model to generate multiple initial prompts; S3, screening the multiple initial prompts with a simple multimodal large model to obtain simple initial prompts; screening the multiple initial prompts with a complex multimodal large model to obtain complex initial prompts; S4, optimizing the simple initial prompts through a simple multimodal large model, a complex multimodal large model, and a first text large model to obtain simple optimal prompts; optimizing the complex initial prompts through a simple multimodal large model, a complex multimodal large model, and a first text large model to obtain complex optimal prompts; S5, inputting the simple optimal prompts, the training set, and the artificially calibrated problems into the simple multimodal large model for training to obtain a simple total loss value, simple answers for each picture, and simple analysis processes for each picture; determining whether the simple total loss value is within a preset threshold range, if so, ending the training; otherwise, using the simple optimal prompts as simple initial prompts, using the complex optimal prompts as complex initial prompts, repeating S4 - S5 until the simple total loss value is within the preset threshold range, stopping the training, and using the simple optimal prompts updated for the last time as the target optimal prompts.

[0005] Optionally, after S5, it includes: inputting each picture corresponding to the simple answer error obtained from the last training into the encoding layer of the visual feature extraction model for feature extraction to obtain the visual feature vector of each picture; inputting the manual analysis process of each picture corresponding to the simple answer error obtained from the last training into the embedding encoder for feature extraction to obtain the text feature vector of each picture; fusing the visual feature vector and the text feature vector of the current picture to obtain the fused feature vector of the current picture; performing a clustering operation on the fused feature vectors of all pictures to obtain multiple clusters; adding the simple analysis process of the picture corresponding to the cluster center of each cluster to the target optimal prompt word.

[0006] Optionally, S2 includes: inputting the training set and the manually calibrated questions of each picture in the training set into the complex multi-modal large model to obtain the keywords of the questions; retrieving a reference prompt word template based on the keywords; inputting the reference prompt word template and the manually calibrated questions of all pictures into the complex multi-modal large model to obtain multiple initial prompt words.

[0007] Optionally, the screening of the multiple initial prompt words using a simple multi-modal large model to obtain simple initial prompt words includes: sequentially taking each initial prompt word as the current initial prompt word; inputting the current initial prompt word, the training set, and the manually calibrated questions into the simple multi-modal large model for training to obtain the simple inference answers and simple inference analysis processes of each picture; calculating the simple analysis process similarity loss value of each picture based on the manual analysis process and simple inference analysis process of each picture; calculating the score corresponding to the current initial prompt word based on the simple analysis process similarity loss value of each picture, the simple inference answer of each picture, and the manual answer of each picture; taking the initial prompt word corresponding to the highest score as the simple initial prompt word.

[0008] Optionally, the optimization of the simple initial prompt through the simple multi-modal large model, the complex multi-modal large model, and the first text large model to obtain the simple optimal prompt includes: inputting the simple initial prompt, the training set, and the manually calibrated problems into the simple multi-modal large model for model training to obtain the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture; inputting the complex initial prompt, the training set, and the manually calibrated problems into the complex multi-modal large model for model training to obtain the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture; according to the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture, and the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture, training the simple initial prompt through the first text large model to obtain the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, and the second simple wrong prompt; inputting the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, and the second simple wrong prompt into the first text large model for training to obtain the simple common prompt; inputting the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, the second simple wrong prompt, the simple common prompt, and the training set into the simple multi-modal large model to obtain the answers inferred from each prompt for each picture and the analysis process; according to the answers inferred from each prompt for each picture and the analysis process, optimizing the simple common prompt using the first text large model to obtain the simple optimal prompt.

[0009] Optionally, based on the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture, as well as the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture, training the simple initial prompt through a first large language model to obtain a first simple correct prompt, a second simple correct prompt, a first simple incorrect prompt, and a second simple incorrect prompt includes: inputting the simple inference analysis processes and human analysis processes of pictures where the simple inference answers are correct and the simple analysis process similarity loss values are greater than or equal to a first threshold into the first large language model for training to obtain a first comprehensive analysis process; inputting the simple initial prompt and the first comprehensive analysis process into the first large language model for training to obtain a first simple correct prompt; for pictures where the simple inference answers are correct and the simple analysis process similarity loss values are less than the first threshold, inputting the complex inference analysis processes of pictures where the complex inference answers are correct and the complex analysis process similarity loss values are greater than or equal to a second threshold, as well as the simple inference analysis processes and human analysis processes of the remaining pictures into the first large language model for training to obtain a second comprehensive analysis process; inputting the simple initial prompt and the second comprehensive analysis process into the first large language model for training to obtain a second simple correct prompt; performing a difference analysis on the simple inference analysis processes of pictures where the simple inference answers are incorrect and the complex inference answers are correct, respectively, with the human analysis process and the complex inference analysis process; inputting the difference analysis into the first large language model for training to obtain a third comprehensive analysis process; inputting the simple initial prompt and the third comprehensive analysis process into the first large language model for training to obtain a first simple incorrect prompt; performing a difference analysis on the human analysis processes of pictures where the simple inference answers are incorrect and the complex inference answers are incorrect, respectively, with the simple inference analysis process and the complex inference analysis process; inputting the difference analysis into the first large language model for training to obtain a fourth comprehensive analysis process; inputting the simple initial prompt and the fourth comprehensive analysis process into the first large language model for training to obtain a second simple incorrect prompt.

[0010] Optionally, the simple analysis process similarity loss value for each picture is calculated according to the following formula:

[0011]

[0012] where, is the simple analysis process similarity loss value of the current picture, is the simple inference analysis process of the current picture, is the human analysis process of the current picture, is the embedding encoder.

[0013] Optionally, the score corresponding to the current initial prompt is calculated according to the following formula:

[0014]

[0015] Among them, is the score corresponding to the current initial prompt word, is the number of pictures in the training set, is the similarity loss value of the simple analysis process of the i-th picture, is whether the simple inference answer of the i-th picture is the same as the manual answer. If it is the same, it is 1; if it is different, it is 0.

[0016] On the other hand, the present invention provides a prompt word optimization training system based on multi-modal large model supervision. The system includes: an artificial calibration unit for constructing a training set including artificial calibration questions, artificial analysis processes, and artificial answers; a generation unit for generating multiple initial prompt words from the training set using a complex multi-modal large model; a screening unit for screening the multiple initial prompt words using a simple multi-modal large model to obtain simple initial prompt words; screening the multiple initial prompt words using a complex multi-modal large model to obtain complex initial prompt words; an optimization unit for optimizing the simple initial prompt words through a simple multi-modal large model, a complex multi-modal large model, and a first text large model to obtain simple optimal prompt words; optimizing the complex initial prompt words through a simple multi-modal large model, a complex multi-modal large model, and a first text large model to obtain complex optimal prompt words; a judgment unit for inputting the simple optimal prompt words, the training set, and the artificial calibration questions into the simple multi-modal large model for training to obtain a simple total loss value, the simple answer for each picture, and the simple analysis process for each picture; judging whether the simple total loss value is within a preset threshold range. If so, the training ends; otherwise, taking the simple optimal prompt words as simple initial prompt words and the complex optimal prompt words as complex initial prompt words, repeating S4 - S5 until the simple total loss value is within the preset threshold range, stopping the training, and taking the simple optimal prompt words obtained by the last update as the target optimal prompt words.

[0017] Optionally, it further includes: an adding unit for: inputting each picture corresponding to the incorrect simple answer obtained from the last training into the encoding layer of the visual feature extraction model for feature extraction to obtain the visual feature vector of each picture; inputting the artificial analysis process of each picture corresponding to the incorrect simple answer obtained from the last training into the embedding encoder for feature extraction to obtain the text feature vector of each picture; fusing the visual feature vector and the text feature vector of the current picture to obtain the fused feature vector of the current picture; performing a clustering operation on the fused feature vectors of all pictures to obtain multiple clusters; adding the simple analysis process of the picture corresponding to the cluster center of each cluster to the target optimal prompt words.

[0018] Advantages of the present invention:

[0019] The present invention provides a method and system for optimizing and training prompting words based on the supervision of multi-modal large models. Among them, the method reduces manual intervention and the cost of optimizing prompting words through the collaborative training of complex large models and simple large models; combines manually calibrated data with the inference process of complex large models, and double supervision improves the accuracy and generalization ability of prompting words; introduces double supervision of question-answer loss and inference process loss, taking into account both accuracy and inference rationality; synchronously optimizes prompting words in two dimensions of logical analysis and visual recognition to solve the adaptation problem of complex scenarios; analyzes error cases through the fusion of visual and text features, discovers typical difficult cases based on clustering, and reversely optimizes prompting words; the method of the present invention greatly reduces the degree of manual participation and improves the automation level of prompting word generation and optimization; significantly improves the comprehensive inference ability of simple multi-modal large models in port navigation monitoring tasks. Brief Description of the Drawings

[0020] Figure 1 is a flowchart of a method for optimizing and training prompting words based on the supervision of multi-modal large models provided by an embodiment of the present invention;

[0021] Figure 2 is a schematic structural diagram of a system for optimizing and training prompting words based on the supervision of multi-modal large models provided by an embodiment of the present invention. Detailed Embodiments

[0022] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0023] Figure 1 is a flowchart of a method for optimizing and training prompting words based on the supervision of multi-modal large models provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0024] S1. Construct a training set including manually calibrated questions, manual analysis processes, and manual answers;

[0025] In an optional implementation manner, the training set is a ship picture training set; assuming that the ship picture training set includes 10 pictures, manually calibrate questions, make human judgment analysis processes, and answers to the questions for each picture.

[0026] The manual calibration questions for each image are usually in the form of multiple-choice questions, and the answers are given manually based on the questions. For example: How many cargo ships are parked at the port terminal in the video? A: 5, B: 6, C: 7.

[0027] Manual analysis process: There are 11 ships in the picture, 5 of which are on the sea and 6 are staying at the dock.

[0028] Human answer: B.

[0029] The problems of manual calibration for each image in the ship image training set are the same or similar.

[0030] S2, using a complex multimodal large model to generate multiple initial prompts for the training set;

[0031] In an optional embodiment, the S2 includes:

[0032] S21, inputting the training set and the manual calibration problem of each image in the training set into the complex multimodal large model to obtain keywords of the problem;

[0033] Input all the manual calibration questions for each image in the training set into a complex multimodal large model (such as the qwen2.5 vl 72B model). Instruct the large model to extract keywords for all questions. For example: Keywords: [pier], [docking], [cargo ship], [quantity statistics].

[0034] S22, obtaining a reference prompt word template based on a keyword online search;

[0035] The keywords extracted in the previous step are searched online (for example, by accessing a search engine or an industry knowledge base); existing prompt words (prompt words) related to the keywords are searched.

[0036] S23, inputting the reference prompt word template and the manual calibration problems of all pictures into the complex multimodal large model to obtain multiple initial prompt words.

[0037] Specifically, the retrieved reference prompt word template and the manual calibration problems of the above 10 pictures are input into the complex multimodal large model together; the complex large model is required to rewrite and expand the prompt words based on the reference prompt word template for all manual calibration problems, and generate at least multiple versions of initial prompt words (for example, 3 to 5).

[0038] S3, screening multiple initial prompt words using a simple multimodal large model to obtain simple initial prompt words; screening multiple initial prompt words using a complex multimodal large model to obtain complex initial prompt words;

[0039] In an alternative embodiment, the process of screening multiple initial prompting words using a simple multi-modal large model to obtain simple initial prompting words includes:

[0040] S311. Successively use each initial prompting word as the current initial prompting word;

[0041] S312. Input the current initial prompting word, the training set, and the manually calibrated questions into the simple multi-modal large model for training to obtain the simple inference answers for each picture and the simple inference analysis process;

[0042] In this application, the simple multi-modal large model is the qwen2.5 vl 7B model.

[0043] S313. Calculate the similarity loss value of the simple analysis process for each picture based on the manual analysis process and the simple inference analysis process for each picture;

[0044] The similarity loss value of the simple analysis process for each picture is calculated according to the following formula:

[0045]

[0046] where, is the similarity loss value of the simple analysis process for the current picture, is the simple inference analysis process for the current picture, is the manual analysis process for the current picture, is the embedding encoder.

[0047] S314. Calculate the score corresponding to the current initial prompting word based on the similarity loss value of the simple analysis process for each picture, the simple inference answers for each picture, and the manual answers for each picture;

[0048] The score corresponding to the current initial prompting word is calculated according to the following formula:

[0049]

[0050] where, is the score corresponding to the current initial prompting word, is the number of pictures in the training set (in the above embodiment, the number of pictures is 10), is the similarity loss value of the simple analysis process for the i-th picture, is whether the simple inference answer for the i-th picture is the same as the manual answer. If it is the same, it is 1 (i.e., the answer is correct); if it is different, it is 0 (i.e., the answer is wrong).

[0051] S315. Use the initial prompting word corresponding to the highest score as the simple initial prompting word.

[0052] The screening of multiple initial prompts using a complex multi-modal large model to obtain complex initial prompts includes:

[0053] S321. Sequentially take each initial prompt as the current initial prompt;

[0054] S322. Input the current initial prompt, the training set, and the manually calibrated questions into the complex multi-modal large model for training to obtain the complex inference answers for each picture and the complex inference analysis process;

[0055] S323. Calculate the complex analysis process similarity loss value for each picture based on the manual analysis process and the complex inference analysis process of each picture;

[0056] The complex analysis process similarity loss value of each picture is calculated according to the following formula:

[0057]

[0058] Where, is the complex analysis process similarity loss value of the current picture, is the complex inference analysis process of the current picture, is the manual analysis process of the current picture, is the embedding encoder.

[0059] S324. Calculate the score corresponding to the current initial prompt based on the complex analysis process similarity loss value of each picture, the complex inference answer of each picture, and the manual answer of each picture;

[0060] The score corresponding to the current initial prompt is calculated according to the following formula:

[0061]

[0062] Where, is the score corresponding to the current initial prompt, is the number of pictures in the training set (in the above embodiment, the number of pictures is 10), is the simple analysis process similarity loss value of the i-th picture, is whether the simple inference answer of the i-th picture is the same as the manual answer. If it is the same, it is 1 (i.e., the answer is correct); if it is different, it is 0 (i.e., the answer is wrong).

[0063] S325. Take the initial prompt corresponding to the highest score as the complex initial prompt.

[0064] S4. Optimize the simple initial prompt through the simple multimodal large model, the complex multimodal large model, and the first text large model to obtain the simple optimal prompt; optimize the complex initial prompt through the simple multimodal large model, the complex multimodal large model, and the first text large model to obtain the complex optimal prompt;

[0065] In this application, the first text large model may be the deepseek R1 text large model.

[0066] In an optional embodiment, the optimizing the simple initial prompt through the simple multimodal large model, the complex multimodal large model, and the first text large model to obtain the simple optimal prompt includes:

[0067] S411. Input the simple initial prompt, the training set, and the manually calibrated questions into the simple multimodal large model for model training to obtain the simple inference answers, the simple analysis processes, and the simple analysis process similarity loss values for each picture; input the complex initial prompt, the training set, and the manually calibrated questions into the complex multimodal large model for model training to obtain the complex inference answers, the complex analysis processes, and the complex analysis process similarity loss values for each picture;

[0068] Compare the simple inference answer of each picture with the manual answer. If it is correct (i.e., the same), it is 1; if it is wrong (i.e., different), it is 0.

[0069] Compare the complex inference answer of each picture with the manual answer. If it is correct (i.e., the same), it is 1; if it is wrong (i.e., different), it is 0.

[0070] The calculation methods of the simple analysis process similarity loss value and the complex analysis process similarity loss value are the same as those in S3 above.

[0071] S412. According to the simple inference answers, the simple analysis processes, and the simple analysis process similarity loss values of each picture, and the complex inference answers, the complex analysis processes, and the complex analysis process similarity loss values of each picture, train the simple initial prompt through the first text large model to obtain the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, and the second simple wrong prompt;

[0072] In an optional embodiment, S412 includes:

[0073] Input the simple inference analysis processes and the manual analysis processes of the pictures corresponding to the correct simple inference answers and the simple analysis process similarity loss values greater than or equal to the first threshold into the first text large model for training to obtain the first comprehensive analysis process; input the simple initial prompt and the first comprehensive analysis process into the first text large model for training to obtain the first simple correct prompt;

[0074] For example, among 10 pictures, 5 pictures have correct simple reasoning answers (i.e., the simple reasoning answers are the same as the manual answers). Among these 5 pictures, 2 pictures have a similarity loss value of the simple analysis process greater than or equal to the first threshold (set to 0.6 in the present invention). Then, the simple reasoning analysis processes and manual analysis processes of these 2 pictures are input into the first large text model for training to obtain the first comprehensive analysis process (i.e., integrating the advantages and disadvantages of all analysis processes from both visual and logical aspects); the simple initial prompt is combined with the advantages and disadvantages of vision and logic to regenerate the first simple correct prompt.

[0075] For the pictures with correct simple reasoning answers and a similarity loss value of the simple analysis process less than the first threshold, the complex reasoning analysis processes of the pictures with correct complex reasoning answers and a similarity loss value of the complex analysis process greater than or equal to the second threshold, as well as the simple reasoning analysis processes and manual analysis processes of the remaining pictures, are input into the first large text model for training to obtain the second comprehensive analysis process; the simple initial prompt and the second comprehensive analysis process are input into the first large text model for training to obtain the second simple correct prompt;

[0076] For example, among 10 pictures, 5 pictures have correct simple reasoning answers (i.e., the simple reasoning answers are the same as the manual answers). Among these 5 pictures, 3 pictures have a similarity loss value of the simple analysis process less than the first threshold (set to 0.6 in the present invention). Among these 3 pictures, 2 pictures have correct complex reasoning answers (i.e., the complex reasoning answers are the same as the manual answers) and a similarity loss value of the complex analysis process greater than or equal to the second threshold. Then, the complex reasoning analysis processes of these 2 pictures, as well as the simple reasoning analysis processes and manual analysis processes of 1 picture (i.e., the picture other than the 2 pictures with correct complex reasoning answers among the 3 pictures) are input into the first large text model for training to obtain the second comprehensive analysis process (i.e., integrating the disadvantages of all analysis processes from both visual and logical aspects); the simple initial prompt is combined with the disadvantages of vision and logic to regenerate the second simple correct prompt.

[0077] The simple reasoning analysis processes of the pictures corresponding to incorrect simple reasoning answers and correct complex reasoning answers are respectively analyzed for differences with the manual analysis process and the complex reasoning analysis process; the difference analysis is input into the first large text model for training to obtain the third comprehensive analysis process; the simple initial prompt and the third comprehensive analysis process are input into the first large text model for training to obtain the first simple incorrect prompt;

[0078] For example: Among 10 pictures, there are 5 pictures with incorrect simple reasoning answers (i.e., the simple reasoning answers are different from the manual answers). Among these 5 pictures, there are 2 pictures with correct complex reasoning answers (i.e., the complex reasoning answers are the same as the manual answers). Then, perform a differential analysis on the simple reasoning analysis process and the manual analysis process of these 2 pictures, and perform a differential analysis on the simple reasoning analysis process and the complex reasoning analysis process of these 2 pictures. Input all the differential analyses into the first text large model for training to obtain the third comprehensive analysis process (i.e., comprehensively summarize the commonalities and differences in terms of both vision and logic from all analysis processes); regenerate the first simple error prompt word by combining the simple initial prompt word with the analysis reasons of vision and logic.

[0079] Perform a differential analysis on the manual analysis process of the pictures corresponding to incorrect simple reasoning answers and incorrect complex reasoning answers, respectively, with the simple reasoning analysis process and the complex reasoning analysis process; input the differential analysis into the first text large model for training to obtain the fourth comprehensive analysis process; input the simple initial prompt word and the fourth comprehensive analysis process into the first text large model for training to obtain the second simple error prompt word.

[0080] For example: Among 10 pictures, there are 5 pictures with incorrect simple reasoning answers (i.e., the simple reasoning answers are different from the manual answers). Among these 5 pictures, there are 3 pictures with incorrect complex reasoning answers (i.e., the complex reasoning answers are different from the manual answers). Then, perform a differential analysis on the manual analysis process and the simple reasoning analysis process of these 3 pictures, and perform a differential analysis on the manual analysis process and the complex reasoning analysis process of these 3 pictures. Input all the differential analyses into the first text large model for training to obtain the fourth comprehensive analysis process (i.e., comprehensively summarize the commonalities and differences in terms of both vision and logic from all analysis processes); regenerate the second simple error prompt word by combining the simple initial prompt word with the analysis reasons of vision and logic.

[0081] S413. Input the first simple correct prompt word, the second simple correct prompt word, the first simple error prompt word, and the second simple error prompt word into the first text large model for training to obtain the simple common prompt word;

[0082] That is, the first text large model analyzes the commonalities of the first simple correct prompt word, the second simple correct prompt word, the first simple error prompt word, and the second simple error prompt word to obtain the simple common prompt word.

[0083] S414. Input the first simple correct prompt, the second simple correct prompt, the first simple incorrect prompt, the second simple incorrect prompt, the simple common prompt, and the training set into the simple multi-modal large model to obtain the answers and analysis processes inferred for each picture through each prompt; according to the answers and analysis processes inferred for each picture through each prompt, optimize the simple common prompt using the first text large model to obtain the simple optimal prompt.

[0084] The optimizing the simple common prompt using the first text large model according to the answers and analysis processes inferred for each picture through each prompt to obtain the simple optimal prompt includes:

[0085] Successively take each picture as the current picture. If the answer inferred for the current picture through the simple common prompt is incorrect, but the answer inferred for the current picture through the first simple correct prompt, or the second simple correct prompt, or the first simple incorrect prompt, or the second simple incorrect prompt is correct, then input the analysis process inferred for the current picture through the simple common prompt and the analysis process inferred for the current picture through the prompt corresponding to the correct inferred answer into the first text large model for training to obtain the fifth comprehensive analysis process; input the simple common prompt and the fifth comprehensive analysis process into the first text large model for training to obtain the simple optimal prompt.

[0086] For example: among 10 pictures, 5 pictures have incorrect answers inferred through the simple common prompt. Among these 5 pictures, the first picture has a correct answer inferred through the first simple correct prompt, the second picture has a correct answer inferred through the second simple correct prompt, the third picture has a correct answer inferred through the first simple incorrect prompt, the fourth picture has a correct answer inferred through the second simple incorrect prompt, and the fifth picture has incorrect answers inferred through the first simple correct prompt, the second simple correct prompt, the first simple incorrect prompt, and the second simple incorrect prompt.

[0087] Then, the analysis process obtained by inferring the first picture with a simple common prompt, the analysis process obtained by inferring the first picture with the first simple correct prompt; the analysis process obtained by inferring the second picture with a simple common prompt, the analysis process obtained by inferring the second picture with the second simple correct prompt; the analysis process obtained by inferring the third picture with a simple common prompt, the analysis process obtained by inferring the third picture with the first simple incorrect prompt; the analysis process obtained by inferring the fourth picture with a simple common prompt, the analysis process obtained by inferring the fourth picture with the second simple incorrect prompt are input into the first text large model for training to obtain the fifth comprehensive analysis process (i.e., integrating the defects of all analysis processes from both visual and logical aspects); the simple common prompt combines the visual and logical defects to regenerate the simple optimal prompt.

[0088] The optimization of the complex initial prompt through the simple multi-modal large model, the complex multi-modal large model, and the first text large model to obtain the complex optimal prompt includes:

[0089] S421. Input the simple initial prompt, the training set, and the manually calibrated problems into the simple multi-modal large model for model training to obtain the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture; input the complex initial prompt, the training set, and the manually calibrated problems into the complex multi-modal large model for model training to obtain the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture;

[0090] S422. According to the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture, and the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture, train the complex initial prompt through the first text large model to obtain the first complex correct prompt, the second complex correct prompt, the first complex incorrect prompt, and the second complex incorrect prompt;

[0091] The above S422 includes:

[0092] Input the complex inference analysis processes and manual analysis processes of the pictures where the complex inference answers are correct and the complex analysis process similarity loss values are greater than or equal to the first threshold into the first text large model for training to obtain the sixth comprehensive analysis process; input the complex initial prompt and the sixth comprehensive analysis process into the first text large model for training to obtain the first complex correct prompt;

[0093] Input the simple reasoning analysis processes of the pictures corresponding to the correct complex reasoning answers and the similarity loss values of the complex analysis processes less than the first threshold, the simple reasoning analysis processes of the pictures corresponding to the correct simple reasoning answers and the similarity loss values of the simple analysis processes greater than or equal to the second threshold, and the complex reasoning analysis processes and manual analysis processes of the remaining pictures into the first large text model for training to obtain the seventh comprehensive analysis process; input the complex initial prompt word and the seventh comprehensive analysis process into the first large text model for training to obtain the second complex correct prompt word;

[0094] Perform differential analysis on the complex reasoning analysis processes of the pictures corresponding to the incorrect complex reasoning answers and the correct simple reasoning answers, respectively, with the manual analysis process and the simple reasoning analysis process; input the differential analysis into the first large text model for training to obtain the eighth comprehensive analysis process; input the complex initial prompt word and the eighth comprehensive analysis process into the first large text model for training to obtain the first complex incorrect prompt word;

[0095] Perform differential analysis on the manual analysis processes of the pictures corresponding to the incorrect complex reasoning answers and the incorrect simple reasoning answers, respectively, with the complex reasoning analysis process and the simple reasoning analysis process; input the differential analysis into the first large text model for training to obtain the ninth comprehensive analysis process; input the simple initial prompt word and the ninth comprehensive analysis process into the first large text model for training to obtain the second complex incorrect prompt word.

[0096] S423. Input the first complex correct prompt word, the second complex correct prompt word, the first complex incorrect prompt word, and the second complex incorrect prompt word into the first large text model for training to obtain the complex common prompt word;

[0097] S424. Input the first complex correct prompt word, the second complex correct prompt word, the first complex incorrect prompt word, the second complex incorrect prompt word, the complex common prompt word, and the training set into the complex multi-modal large model to obtain the answers and analysis processes inferred by each prompt word for each picture; according to the answers and analysis processes inferred by each prompt word for each picture, optimize the complex common prompt word using the first large text model to obtain the complex optimal prompt word.

[0098] S5. Input the simple optimal prompt word, the training set, and the manually calibrated questions into the simple multi-modal large model for training to obtain the simple total loss value, the simple answers for each picture, and the simple analysis processes for each picture; determine whether the simple total loss value is within the preset threshold range. If so, end the training; otherwise, use the simple optimal prompt word as the simple initial prompt word, use the complex optimal prompt word as the complex initial prompt word, and repeat S4 - S5 until the simple total loss value is within the preset threshold range, then stop the training, and use the last updated simple optimal prompt word as the target optimal prompt word.

[0099] In an optional embodiment, the simple total loss value is calculated according to the following formula:

[0100]

[0101] where is the simple total loss value, is the number of pictures in the training set, is the similarity loss value of the simple analysis process of the i-th picture (calculated according to the simple analysis process and the manual analysis process of the i-th picture, which is the same as the calculation method in S3 above), is whether the simple answer of the i-th picture is the same as the manual answer. If it is the same, it is 1; if it is different, it is 0.

[0102] Judge whether the simple total loss value is within the preset threshold range. If so, end the training; otherwise, use the simple optimal prompt word as the simple initial prompt word, use the complex optimal prompt word as the complex initial prompt word, and repeat S4~S5 until the simple total loss value is within the preset threshold range, stop the training, and use the simple optimal prompt word obtained by the last update as the target optimal prompt word.

[0103] After S5, it includes:

[0104] Input each picture corresponding to the simple answer error obtained from the last training (that is, the simple answer is different from the manual answer) (that is, all error cases) into the encoding layer of the visual feature extraction model for feature extraction to obtain the visual features of each picture, and convert the visual features of each picture into 1*2048-dimensional feature values to obtain the visual feature vector of each picture;

[0105] Input the manual analysis process of each picture corresponding to the simple answer error obtained from the last training (that is, the simple answer is different from the manual answer) into the embedding encoder (that is, the beg_embeding encoder) for feature extraction to obtain the text features of each picture, and convert the text features of each picture into 1*2048-dimensional feature values to obtain the text feature vector of each picture;

[0106] Fuse the visual feature vector and the text feature vector of the current picture to obtain the fused feature vector of the current picture;

[0107] The fused feature vector of the current picture is calculated according to the following formula:

[0108]

[0109] where is the fused feature vector of the current picture, is the visual feature vector of the current image, is the text feature vector of the current image.

[0110] Perform clustering operations on the fusion feature vectors of all images to obtain multiple clusters; add the simple analysis process of the image corresponding to the cluster center of each cluster to the target optimal prompt.

[0111] In an optional embodiment, the fusion feature vectors of all images (i.e., the fusion feature vectors based on all error cases) are divided into 5 categories by k-mean clustering operation to obtain 5 clusters; take the cluster centers of the 5 clusters as typical error cases and add them to the target optimal prompt.

[0112] Adding typical error cases to the target optimal prompt can help us reverse-enhance the prompt, enabling the prompt to "consider" such situations in advance, thereby guiding the model to avoid errors.

[0113] After obtaining the target optimal prompt, input the ship image to be analyzed, the manual calibration problem of the ship image to be analyzed (this problem is similar to the problems in the above training set), and the target optimal prompt into a simple multi-modal large model for training to obtain the inference answer and inference analysis process of the ship image to be analyzed. Through the guidance of the target optimal prompt, the accuracy of the inference answer of the ship image to be analyzed can be greatly improved.

[0114] Figure 2 is a schematic structural diagram of a prompt optimization training system based on multi-modal large model supervision provided by an embodiment of the present invention; as Figure 2 shown, the system includes:

[0115] The manual calibration unit 201 is used to construct a training set including manual calibration problems, manual analysis processes, and manual answers;

[0116] The generation unit 202 is used to generate multiple initial prompts from the training set using a complex multi-modal large model;

[0117] The screening unit 203 is used to screen multiple initial prompts using a simple multi-modal large model to obtain simple initial prompts; screen multiple initial prompts using a complex multi-modal large model to obtain complex initial prompts;

[0118] The optimization unit 204 is used to optimize the simple initial prompts through a simple multi-modal large model, a complex multi-modal large model, and a first text large model to obtain simple optimal prompts; optimize the complex initial prompts through a simple multi-modal large model, a complex multi-modal large model, and a first text large model to obtain complex optimal prompts;

[0119] A judgment unit 205 is configured to input the simple optimal prompt words, the training set, and the manually calibrated questions into the simple multi-modal large model for training to obtain a simple total loss value, a simple answer for each picture, and a simple analysis process for each picture; determine whether the simple total loss value is within a preset threshold range. If so, end the training; otherwise, use the simple optimal prompt words as the simple initial prompt words, use the complex optimal prompt words as the complex initial prompt words, and repeat S4 - S5 until the simple total loss value is within the preset threshold range, stop the training, and use the simple optimal prompt words obtained by the last update as the target optimal prompt words.

[0120] In an optional embodiment, the system further includes: an adding unit, configured to:

[0121] Input each picture corresponding to the wrong simple answer obtained from the last training into the encoding layer of the visual feature extraction model for feature extraction to obtain a visual feature vector for each picture;

[0122] Input the manual analysis process of each picture corresponding to the wrong simple answer obtained from the last training into the embedding encoder for feature extraction to obtain a text feature vector for each picture;

[0123] Fuse the visual feature vector and the text feature vector of the current picture to obtain a fused feature vector for the current picture;

[0124] Perform a clustering operation on the fused feature vectors of all pictures to obtain multiple clusters; add the simple analysis process of the picture corresponding to the cluster center of each cluster to the target optimal prompt words.

[0125] The system of this application corresponds to the above method, and the specific implementation manners of the system will not be repeated here.

[0126] The method of the present invention reduces manual intervention and lowers the cost of prompt word optimization through the collaborative training of a complex large model and a simple large model; combines manually calibrated data with the inference process of the complex large model, and uses dual supervision to improve the accuracy and generalization ability of prompt words; introduces dual supervision of question answer loss and inference process loss, taking into account both accuracy and inference rationality; synchronously optimizes prompt words in two dimensions of logical analysis and visual recognition to solve the adaptation problem of complex scenarios; analyzes error cases through the fusion of visual and text features, discovers typical difficult cases based on clustering, and reversely optimizes prompt words; the method of the present invention greatly reduces the degree of manual participation, improves the automation level of prompt word generation and optimization; and significantly improves the comprehensive inference ability of the simple multi-modal large model in the port navigation monitoring task.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A prompting word optimization training method based on the supervision of a multimodal large model, characterized in that, Including: S1. Construct a training set including artificial calibration problems, artificial analysis processes, and artificial answers; S2. Use a complex multi-modal large model to generate multiple initial prompt words from the training set; S3. Screen the multiple initial prompt words using a simple multi-modal large model to obtain simple initial prompt words; Screen the multiple initial prompt words using a complex multi-modal large model to obtain complex initial prompt words; S4. Optimize the simple initial prompt words through a simple multi-modal large model, a complex multi-modal large model, and a first text large model to obtain simple optimal prompt words; Optimize the complex initial prompt words through a simple multi-modal large model, a complex multi-modal large model, and a first text large model to obtain complex optimal prompt words; S5. Input the simple optimal prompt words, the training set, and the artificial calibration problems into the simple multi-modal large model for training to obtain a simple total loss value, simple answers for each picture, and a simple analysis process for each picture; Determine whether the simple total loss value is within a preset threshold range. If so, end the training; Otherwise, use the simple optimal prompt words as simple initial prompt words, use the complex optimal prompt words as complex initial prompt words, and repeat S4 - S5 until the simple total loss value is within the preset threshold range. Stop the training and use the simple optimal prompt words obtained from the last update as the target optimal prompt words.

2. The method according to claim 1, wherein After S5, it includes: Input each picture corresponding to the incorrect simple answer obtained from the last training into the encoding layer of the visual feature extraction model for feature extraction to obtain the visual feature vector of each picture; Input the artificial analysis process of each picture corresponding to the incorrect simple answer obtained from the last training into the embedding encoder for feature extraction to obtain the text feature vector of each picture; Fuse the visual feature vector and the text feature vector of the current picture to obtain the fused feature vector of the current picture; Perform clustering operations on the fused feature vectors of all pictures to obtain multiple clusters; Add the simple analysis process of the picture corresponding to the cluster center of each cluster to the target optimal prompt words.

3. The method according to claim 1, wherein S2 includes: Input the training set and the artificial calibration problems of each picture in the training set into the complex multi-modal large model to obtain the keywords of the problems; Retrieve the reference prompt word template based on the keywords through network search; Input the reference prompt word template and the artificial calibration problems of all pictures into the complex multi-modal large model to obtain multiple initial prompt words.

4. The method according to claim 1, characterized in that, The screening of the multiple initial prompt words using a simple multi-modal large model to obtain simple initial prompt words includes: Successively use each initial prompt word as the current initial prompt word; Input the current initial prompt word, the training set, and the artificial calibration problems into the simple multi-modal large model for training to obtain the simple inference answers for each picture and the simple inference analysis process; Calculate the simple analysis process similarity loss value for each picture based on the artificial analysis process and the simple inference analysis process of each picture; Calculate the score corresponding to the current initial prompt word based on the simple analysis process similarity loss value of each picture, the simple inference answers of each picture, and the artificial answers of each picture; Take the initial prompt corresponding to the highest score as the simple initial prompt.

5. The method according to claim 1, wherein The optimization of the simple initial prompt through the simple multi-modal large model, the complex multi-modal large model, and the first text large model to obtain the simple optimal prompt includes: Input the simple initial prompt, the training set, and the manually calibrated questions into the simple multi-modal large model for model training to obtain the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture; input the complex initial prompt, the training set, and the manually calibrated questions into the complex multi-modal large model for model training to obtain the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture; Based on the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture, and the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture, train the simple initial prompt through the first text large model to obtain the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, and the second simple wrong prompt; Input the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, and the second simple wrong prompt into the first text large model for training to obtain the simple common prompt; Input the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, the second simple wrong prompt, the simple common prompt, and the training set into the simple multi-modal large model to obtain the answers and analysis processes inferred from each prompt for each picture; based on the answers and analysis processes inferred from each prompt for each picture, optimize the simple common prompt using the first text large model to obtain the simple optimal prompt.

6. The method according to claim 5, wherein The training of the simple initial prompt through the first text large model to obtain the first simple correct prompt, the second simple correct prompt, the first simple wrong prompt, and the second simple wrong prompt based on the simple inference answers, simple analysis processes, and simple analysis process similarity loss values for each picture, and the complex inference answers, complex analysis processes, and complex analysis process similarity loss values for each picture includes: Input the simple inference analysis processes and manual analysis processes of the pictures where the simple inference answers are correct and the simple analysis process similarity loss values are greater than or equal to the first threshold into the first text large model for training to obtain the first comprehensive analysis process; input the simple initial prompt and the first comprehensive analysis process into the first text large model for training to obtain the first simple correct prompt; Input the complex reasoning analysis process of the picture corresponding to the similarity loss value of the correct simple reasoning answer and the simple analysis process being less than the first threshold, the complex reasoning analysis process of the picture corresponding to the correct complex reasoning answer and the similarity loss value being greater than or equal to the second threshold, as well as the simple reasoning analysis process and the manual analysis process of the remaining pictures into the first large text model for training to obtain the second comprehensive analysis process; Input the simple initial prompt word and the second comprehensive analysis process into the first large text model for training to obtain the second simple correct prompt word; Conduct a difference analysis on the simple reasoning analysis process of the picture corresponding to the incorrect simple reasoning answer and the correct complex reasoning answer, respectively, with the manual analysis process and the complex reasoning analysis process; Input the difference analysis into the first large text model for training to obtain the third comprehensive analysis process; Input the simple initial prompt word and the third comprehensive analysis process into the first large text model for training to obtain the first simple incorrect prompt word; Conduct a difference analysis on the manual analysis process of the picture corresponding to the incorrect simple reasoning answer and the incorrect complex reasoning answer, respectively, with the simple reasoning analysis process and the complex reasoning analysis process; Input the difference analysis into the first large text model for training to obtain the fourth comprehensive analysis process; Input the simple initial prompt word and the fourth comprehensive analysis process into the first large text model for training to obtain the second simple incorrect prompt word.

7. The method according to claim 4, characterized in that The similarity loss value of the simple analysis process of each picture is calculated according to the following formula: ; Among them, is the similarity loss value of the simple analysis process of the current image, is the simple reasoning and analysis process of the current image, is the manual analysis process of the current image, is the embedding encoder.

8. The method according to claim 4, characterized in that, The score corresponding to the current initial prompt word is calculated according to the following formula: ; Among them, is the score corresponding to the current initial prompt word, is the number of pictures in the training set, is the similarity loss value of the simple analysis process of the i-th picture, indicates whether the simple inference answer of the i-th picture is the same as the manual answer. If it is the same, it is 1; if it is different, it is 0.

9. A prompt optimization training system based on the supervision of a multimodal large model, characterized in that Including: An artificial calibration unit for constructing a training set including artificial calibration questions, manual analysis processes, and artificial answers; A generation unit for generating multiple initial prompt words from the training set using a complex multi-modal large model; A screening unit for screening the multiple initial prompt words using a simple multi-modal large model to obtain a simple initial prompt word; Screen the multiple initial prompt words using a complex multi-modal large model to obtain a complex initial prompt word; An optimization unit for optimizing the simple initial prompt word through a simple multi-modal large model, a complex multi-modal large model, and the first large text model to obtain a simple optimal prompt word; Optimize the complex initial prompt word through a simple multi-modal large model, a complex multi-modal large model, and the first large text model to obtain a complex optimal prompt word; A judgment unit for inputting the simple optimal prompt word, the training set, and the artificial calibration questions into the simple multi-modal large model for training to obtain a simple total loss value, the simple answer of each picture, and the simple analysis process of each picture; Judge whether the simple total loss value is within the preset threshold range. If so, end the training; On the contrary, use the simple optimal prompt word as the simple initial prompt word, use the complex optimal prompt word as the complex initial prompt word, repeat S4 - S5 until the simple total loss value is within the preset threshold range, stop the training, and use the simple optimal prompt word updated last time as the target optimal prompt word.

10. The system according to claim 9, wherein It also includes: An addition unit for: For each image corresponding to the simple answer error obtained from the last training, input it into the encoding layer of the visual feature extraction model for feature extraction to obtain the visual feature vector of each image; For the manual analysis process of each image corresponding to the simple answer error obtained from the last training, input it into the embedding encoder for feature extraction to obtain the text feature vector of each image; Fuse the visual feature vector and the text feature vector of the current image to obtain the fused feature vector of the current image; Perform a clustering operation on the fused feature vectors of all images to obtain multiple clusters; add the simple analysis process of the image corresponding to the cluster center of each cluster to the target optimal prompt word.

Citation Information

Patent Citations

  • Large language model optimization generation method based on optimal cue word selection

    CN119476209A

Cited By

  • Multi-modal model training method and device based on thinking chain prompt pool

    CN121146070A

  • Multi-modal large model-based cue word automatic generation model training method and system

    CN121600511A

  • A prompt word automatic generation model training method and system based on a multi-modal large model

    CN121600511B

  • Multi-modal information-based catalyst literature prompt iterative optimization method and system

    CN122417216A