Prompt word optimization method and system based on approximate submodule function and continuous learning

Through the joint optimization method of approximate submodular functions and continuous learning, the shortcomings of the visual-language model in cue word selection in complex environments are solved, efficient and adaptive cue word updates are achieved, and the performance and robustness of multimodal image classification are improved.

CN120597895AActive Publication Date: 2025-09-05SHANDONG UNIV +1

Patent Information

Application Number
CN202511092853.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-05
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing prompt word selection methods based on vision-language models cannot effectively and automatically select the optimal prompt words under complex and open conditions, resulting in limited improvement in classification performance. In particular, they have poor generalization and lack of real-time performance in multimodal data processing, and rely on high-quality annotated data that is difficult to obtain.

Method used

A joint optimization method of approximate submodular function and continuous learning is adopted. By constructing a set of candidate prompt words, designing a combined objective function, combining a greedy selection algorithm and a multi-round iterative mechanism, the optimal prompt word is gradually selected and optimized in the discrete-continuous space to achieve adaptive updating of the prompt word.

Benefits of technology

It significantly improves the performance of the model in multimodal image classification tasks, enhances generalization ability and robustness, reduces dependence on labeled data, adapts to complex dynamic environments, and is suitable for scenarios with unstable data quality or scarce labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597895A_ABST
    Figure CN120597895A_ABST
Patent Text Reader

Abstract

The invention provides a cue word optimization method and system based on an approximate sub-module function and continuous learning, and relates to the technical field of artificial intelligence multi-mode perception.The method comprises the steps that a candidate cue word set is constructed, a combined objective function based on the property of the approximate sub-module function is designed, and a candidate cue word set is constructed; solving the combined objective function by adopting a greedy selection algorithm combining random disturbance, multi-round iteration and a task self-adaptive mechanism, realizing optimal selection of a candidate cue word set, gradually selecting a cue word with the maximum gain from the candidate cue word set, adding the cue word into an optimal subset, and when a new cue word is selected, selecting the cue word with the maximum gain into the optimal subset. And if the target cue words are selected, carrying out local optimization once, after all the target cue words are selected, carrying out joint optimization on all the cue words in a continuous space by adopting an alternate optimization strategy, and carrying out continuous iteration until the optimized cue words are obtained. According to the method and the device, efficient self-adaptive updating of the cue words of the language model is realized, so that the generalization performance and robustness of the model in zero-sample, few-sample and concept drift scenes are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence multimodal perception technology, and in particular to a prompt word optimization method and system based on approximate submodular functions and continuous learning. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the rapid development of artificial intelligence (AI), its application in complex environments is becoming increasingly widespread, particularly in the fields of multimodal perception and multi-task learning. Existing methods based on vision-language models (such as CLIP) have limitations in cue word selection and are unable to effectively and automatically select the optimal cue word. This severely hinders classification performance when faced with complex, open-ended recognition tasks. Therefore, a more effective cue word optimization method is urgently needed to enhance the model's recognition capabilities in diverse, dynamic, and uncertain environments, and improve its generalization and robustness.

[0004] When performing recognition tasks under complex and open conditions, data sources often have significant differences in modality, resolution, acquisition time, and spatial information. The fusion and processing of data from different modalities often face huge challenges. Especially in dynamic environments, the quality and stability of data are often affected by multiple factors, such as environmental changes, equipment performance, and uncertainty in acquisition conditions. In addition, the processing of these multimodal data usually requires a large amount of annotated data. However, the cost of obtaining high-quality annotated data is high, and the annotation process is often time-consuming and labor-intensive, making the acquisition of large-scale annotated data difficult.

[0005] Existing supervised methods rely heavily on high-quality and sufficient annotated data. However, in practical applications, especially when dealing with complex multimodal data, it is often difficult to obtain sufficient annotated data and achieve effective spatial and temporal alignment. Furthermore, existing multimodal data fusion methods suffer from poor generalization and real-time performance in the face of volatile environments and unstable data quality, making them difficult to meet the demands of complex tasks. Summary of the Invention

[0006] To address the above-mentioned issues, this paper proposes a prompt word optimization method and system based on approximate submodular functions and continuous learning. In a multimodal and multi-task environment, this method constructs a set of candidate prompt words and applies an approximate submodular utility function to evaluate their discriminability and diversity. The optimal subset is greedily selected, and then combined with continuous optimization to achieve efficient adaptive updating of the prompt words of the pre-trained language model, thereby improving the generalization performance and robustness of the model in zero-shot, few-shot, and concept drift scenarios.

[0007] According to some embodiments, the present disclosure adopts the following technical solutions: The prompt word optimization method based on approximate submodular functions and continuous learning includes: Obtain image data to be classified; The cue words corresponding to the image are optimized using an approximate submodular function and a continuous learning joint optimization strategy. The image data and the optimized cue words are input into a multimodal classification model for image-text matching, and the image classification result is output. Among them, the process of implementing the joint optimization strategy of approximate submodular function and continuous learning for prompt words includes: constructing a set of candidate prompt words, designing a combined objective function based on the properties of approximate submodular function, and solving the combined objective function using a greedy selection algorithm that combines random perturbation, multiple rounds of iteration and task adaptation mechanism to achieve optimal selection of the candidate prompt word set, gradually selecting the prompt word with the largest gain from the candidate prompt word set and adding it to the optimized subset. Each time a new prompt word is selected, a local optimization is performed. When all target prompt words are selected, all prompt words are jointly optimized in the continuous space using an alternating optimization strategy, and iterate continuously until the optimized prompt word is obtained.

[0008] According to some embodiments, the present disclosure adopts the following technical solutions: The prompt word optimization system based on approximate submodular functions and continuous learning includes: A data acquisition module, used to acquire image data to be classified; The image-text matching classification module is used to optimize the prompt words corresponding to the image using an approximate submodular function and a continuous learning joint optimization strategy; the image data and the optimized prompt words are input into the multimodal classification model for image-text matching, and the image classification result is output; Among them, the process of implementing the joint optimization strategy of approximate submodular function and continuous learning for prompt words includes: constructing a set of candidate prompt words, designing a combined objective function based on the properties of approximate submodular function, and solving the combined objective function using a greedy selection algorithm that combines random perturbation, multiple rounds of iteration and task adaptation mechanism to achieve optimal selection of the candidate prompt word set, gradually selecting the prompt word with the largest gain from the candidate prompt word set and adding it to the optimized subset. Each time a new prompt word is selected, a local optimization is performed. When all target prompt words are selected, all prompt words are jointly optimized in the continuous space using an alternating optimization strategy, and iterate continuously until the optimized prompt word is obtained.

[0009] According to some embodiments, the present disclosure adopts the following technical solutions: A computer program product includes a computer program. When the computer program is executed by a processor, the computer program implements the prompt word optimization method based on approximate submodular functions and continuous learning.

[0010] According to some embodiments, the present disclosure adopts the following technical solutions: A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the prompt word optimization method based on approximate submodular functions and continuous learning is implemented.

[0011] According to some embodiments, the present disclosure adopts the following technical solutions: An electronic device includes: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the prompt word optimization method based on approximate submodule functions and continuous learning.

[0012] Compared with the prior art, the present invention has the following beneficial effects: This proposed cue word optimization method, based on approximate submodular functions and continuous learning, significantly improves the performance of the CLIP model in multimodal image classification tasks through intelligent cue word selection and staged optimization training. By combining the optimization properties of approximate submodular functions, the efficiency of greedy selection algorithms, and the advantages of discrete and continuous learning methods, it achieves superior classification performance in complex open environments and effectively addresses typical challenges such as uneven data distribution, high label noise, and limited computing resources.

[0013] The disclosed method for optimizing prompt words based on approximate submodular functions and continuous learning proposes a prompt word optimization technique that combines approximate submodular functions and continuous learning. This method can effectively improve the optimization efficiency and quality of the prompt word set, thereby enhancing the generalization and robustness of the model in complex recognition tasks. By adaptively updating prompt words and performing real-time optimization, this technology can adapt to complex and dynamic environmental changes without relying on large amounts of labeled data. It is particularly suitable for scenarios with unstable data quality or where large amounts of labeled data are difficult to obtain. This method not only significantly improves the accuracy and efficiency of multimodal data processing, but also has broad application prospects and is suitable for recognition tasks under various complex and open conditions.

[0014] The disclosed prompt word optimization method based on approximate submodule function and continuous learning uses this function as an evaluation criterion for prompt word selection, ensuring that the selected prompt words are representative and have information coverage, thereby maximizing the model's expressiveness in multiple categories and multiple scenarios. In order to solve the problem of difficulty in identifying "long-tail categories" in open environment recognition, the present invention introduces a category adaptive weighting mechanism in the submodule function, dynamically adjusting the coverage evaluation value of the prompt word according to the category frequency, so that the prompt word selection process pays more attention to rare categories and edge scenarios, significantly improving the model's classification performance and overall generalization ability on the tail category. In addition, the approximate submodule function has significant advantages in theoretical computational efficiency, so that the prompt word screening process can be completed with limited computing resources even when faced with thousands of candidate words.

[0015] The proposed method for optimizing prompt words based on approximate submodular functions and continuous learning adopts a greedy selection strategy to efficiently implement the prompt word screening process, gradually selecting words with the largest marginal benefit from the candidate set to construct the optimal subset. The greedy algorithm has been proven to have theoretical performance guarantees when dealing with submodular function optimization problems, and can achieve the desired result with an approximate ratio of (approximately 63%) approaches the global optimal solution, making it particularly suitable for large-scale discrete choice tasks. Compared to traditional brute-force search or gradient-based global optimization methods, the greedy strategy is more scalable and computationally efficient, achieving high-quality solutions in an acceptable time.

[0016] This proposed method for optimizing prompt words, based on approximate submodular functions and continuous learning, further enhances the standard greedy framework by introducing multiple rounds of iteration and perturbation mechanisms. In each round, the algorithm constructs subsets from multiple initialization states, effectively increasing the diversity and coverage of the search space and reducing the risk of falling into local optima. Furthermore, the prompt word selection process comprehensively considers category distribution, task context, and historical selection paths, dynamically adjusting word priorities and guiding the model to focus on difficult categories and marginal examples within the task. This enhances the model's adaptability in complex environments and the global representativeness of the prompt word set.

[0017] This disclosed method for optimizing prompt words based on approximate submodule functions and continuous learning proposes a discrete-continuous fusion learning framework. High-quality prompt words are first selected using discrete submodule functions, and then end-to-end fine-tuned using learnable vectors in the CLIP framework. The training process employs a staged scheduling mechanism, with a fixed set of prompt words at each stage. Performance is evaluated after small-batch training, and the prompt word pool is dynamically updated. This method improves training efficiency and final classification accuracy while effectively maintaining the stability of the semantic structure of the prompt words. In actual deployment, this strategy can flexibly adapt to the training requirements of varying data volumes and scenarios, demonstrating strong versatility and portability.

[0018] The proposed method for optimizing prompt words based on approximate submodular functions and continuous learning proposes an efficient and intelligent prompt word learning algorithm by effectively integrating approximate submodular functions, an optimized greedy selection algorithm, and an improved discrete-continuous optimization learning method. Specific advantages include: (1) significantly reducing the reliance on labeled data, saving labor costs, and being particularly suitable for problems lacking precise labels in open scenarios; (2) improving the interpretability and generalization performance of prompt words, and enhancing the model's ability to handle new categories and unseen samples; and (3) being suitable for real-time target monitoring applications, supporting efficient deployment on edge computing devices, and meeting high-frequency, low-latency task requirements.

[0019] This disclosure is not only widely applicable to tasks such as ecological and environmental monitoring, marine pollution source tracing, and offshore activity supervision, but also has the potential to expand its application in multimodal recognition fields such as remote sensing image analysis, urban security, and agricultural monitoring. With the further development of large-scale model and small-sample learning and deployment requirements in the future, this invention has great industrial transformation value and research and promotion potential, and will play an important role in promoting the development of intelligent environmental protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which constitute a part of the present disclosure, are used to provide a further understanding of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure.

[0021] Figure 1 This is a flow chart of the joint optimization strategy of approximate submodular function and continuous learning according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram of a specific process of applying the method of an embodiment of the present disclosure to a method for classifying images of pollution in nearshore waters. DETAILED DESCRIPTION

[0022] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0023] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs.

[0024] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0025] Example 1 In one embodiment of the present disclosure, a prompt word optimization method based on approximate submodular functions and continuous learning is provided. The method is applied to image classification tasks. The prompt word optimization can be better applied to image classification and image-text matching, making the image classification results more accurate. The specific method steps include: Step 1: Obtain image data to be classified; Step 2: Optimize the prompt word corresponding to the image using an approximate submodular function and a continuous learning joint optimization strategy; input the image data and the optimized prompt word into a multimodal classification model for image-text matching, and output the image classification result; Among them, the process of implementing the joint optimization strategy of approximate submodular function and continuous learning for prompt words includes: constructing a set of candidate prompt words, designing a combined objective function based on the properties of approximate submodular function, and solving the combined objective function using a greedy selection algorithm that combines random perturbation, multiple rounds of iteration and task adaptation mechanism to achieve optimal selection of the candidate prompt word set, gradually selecting the prompt word with the largest gain from the candidate prompt word set and adding it to the optimized subset. Each time a new prompt word is selected, a local optimization is performed. When all target prompt words are selected, all prompt words are jointly optimized in the continuous space using an alternating optimization strategy, and iterate continuously until the optimized prompt word is obtained.

[0026] As an embodiment, the present disclosure proposes a prompt word optimization method based on approximate submodular functions and continuous learning. By effectively integrating approximate submodular functions, an optimized greedy selection algorithm, and an improved discrete-continuous optimization learning method, an efficient and intelligent prompt word learning optimization algorithm is proposed. The prompt word optimization process is described in detail below: Step 1: Construct a set of candidate prompt words; In this paper, to achieve high-quality, interpretable prompt word selection, we first construct a set of candidate prompt words with good semantic coverage from the English vocabulary. The candidate prompt word set is constructed based on the NLTK standard vocabulary and is screened using a variety of linguistic and model adaptability rules to ensure that the candidate words have semantic integrity, model compatibility, and applicable frequency. The specific steps are as follows: Step 1.1: Obtain candidate prompt words and conduct preliminary screening; This example uses the English corpus provided by NLTK (Natural Language Toolkit) as a basis, iterates through the words one by one, and only retains the words that meet the following conditions: (1) Alphabetic: exclude non-natural language words containing symbols, numbers, or special characters; (2) Semantic validity (WordNet coverage): By calling the WordNet interface, the legal words that can be found in its vocabulary are retained; (3) Frequency restriction (Zipf frequency ≥ 3.5): The Zipf frequency score defined in the wordfreq library is used to measure the commonness of words, remove extremely rare or domain-specific low-frequency words, and improve generalization ability.

[0027] Step 1.2: Model compatibility screening; Furthermore, to ensure the availability of candidate cue words in vision-language models such as CLIP, we further employ the BPE tokenizer used in CLIP to screen candidate words for encoding, retaining only those encoded as single tokens. This strategy avoids interference in the embedding space caused by syntactically incomplete or cross-token cue words, which helps improve the stability and consistency of subsequent cue word representations.

[0028] Finally, the remaining prompt words constitute the candidate prompt word set, which is recorded as:

[0029] Among them, each These are semantically valid atomic terms that have passed the initial multiple screening, have a moderate frequency (Zipf frequency ≥ 3.5), and can be directly processed by CLIP. This set of candidate cue words constitutes the input space for the cue word optimization process in this disclosure and is used for subsequent approximate submodular function evaluation and greedy selection operations.

[0030] Through the above processing, the candidate prompt word set constructed by the present disclosure has good semantic breadth, model compatibility and interpretability, and can provide a solid foundation for subsequent prompt word selection.

[0031] Step 2: Design a combined objective function based on the properties of the approximate submodular function; To efficiently select a set of semantically representative, discriminative, and complementary cue words from a large set of candidate cue words, this paper designs a combined objective function that combines task loss and semantic diversity. This function exhibits properties similar to submodular functions in practical optimization, providing a theoretical foundation and practical feasibility for subsequent greedy algorithms.

[0032] Step 2.1: Construct the combined objective function; In the CLIP model, prompt words directly affect the image-text matching score, so choosing the right prompt words is crucial to improving classification accuracy. However, there are currently the following problems and requirements: (1) Minimizing only the classification loss will cause the selected cue words to be overly concentrated in the semantic space, thereby reducing generalization ability; (2) If there is no regularization constraint, the selected prompt words may be highly redundant and unable to cover the diverse image content; (3) In open target monitoring tasks, the scenes are varied and the categories are complex, so the prompt words need to have both discriminative and semantic coverage.

[0033] Therefore, when constructing the objective function, the present disclosure takes into account both classification performance and semantic diversity, and constructs the following combined objective function:

[0034] in, CLIP-based contrast loss is used to measure the improvement of the selected cue words on the image discrimination ability; Indicates the semantic redundancy within the prompt word subset, using the average cosine similarity measurement; To control the strength of the diversity regularization term, adjust the trade-off between task-driven and semantic divergence; It is a set of selected category labels and candidate prompt words.

[0035] Furthermore, the combined objective function based on the properties of the approximate submodular function is composed of a classification performance term and a semantic diversity regularization term. It includes: (1) Classification performance items Calculation method: Given a category label and a set of candidate prompt words , construct the text prompt as: prompt="aphotoofa[CLASS]withemphasison: " "," ",…," Input the sentence into the CLIP text encoder and the image data into the image encoder to calculate the image-text matching similarity. , and then use cross entropy loss:

[0036] in, For samples i CLIP image features; For the real category Corresponding text features; is the learnable temperature parameter; is the total number of categories. This loss term encourages the subset of prompt words to maximize the category discrimination ability.

[0037] (2) Diversity regularization term Calculation method: To prevent the selected prompt words from clustering redundancy in the semantic space and reducing generalization performance, the present disclosure designs the following diversity regularization term:

[0038] in, represents the embedding vector of the candidate prompt word, This represents the embedding vector of the selected cue word. A larger similarity value indicates closer semantics. Minimizing this value promotes a more even distribution of selected cue words in the semantic space, thereby improving the model's ability to perceive diverse image content.

[0039] As an embodiment, the word vector of the prompt word is extracted by the tokenizer + embedding layer of CLIP to ensure consistency with the model reasoning process and high expressiveness.

[0040] Step 2.2 Mathematical analysis and application of properties of approximate submodular functions Although the combined loss function of the present disclosure is not a submodular function in the strict sense (i.e., it does not necessarily satisfy the diminishing marginal returns defined by all submodular functions), it exhibits approximate submodularity in practical applications, namely: (1) With the candidate prompt word set Extension, newly added prompt word pairs The improvement effect shows a marginal decreasing trend; (2) That is: there is a , for smaller sets Gain Higher than for larger sets Gain :

[0041] This property lays the foundation for the greedy algorithm employed in this paper: greedy solutions to approximate submodular functions are guaranteed to achieve theoretical near-optimality. Furthermore, this objective is highly stable in practical tasks, possesses clear physical interpretability, and is task-driven, significantly improving the classification performance of the CLIP model in complex scenarios.

[0042] Step 3: A greedy selection algorithm that combines random perturbations, multiple iterations, and task adaptation is used to solve the combined objective function. In order to efficiently solve the defined approximate submodular objective function, this paper designs a prompt word selection algorithm based on a greedy strategy, combining random perturbation, multiple rounds of iteration and task adaptation mechanism to achieve automatic construction and global optimization of prompt word subsets.

[0043] Step 3.1: Basic process of greedy selection algorithm; Suppose the candidate prompt word set is The optimization goal is to not exceed the set length Under the premise of , so that the joint loss function Minimum. The algorithm steps are as follows: 1. Initialize the optimized subset of prompt words to an empty set:

[0044] 2. In each iteration, traverse all candidate prompt words , calculate the loss reduction after adding it to the optimized subset:

[0045] 3. Select the candidate prompt word that brings the greatest loss reduction , add to the current collection:

[0046] 4. Repeat steps 2-3 until the maximum number of words is reached or the loss drops below the preset threshold, stop the iteration, and finally get the optimized subset under the current greedy iteration, and finally output the prompt word set This is the optimal subset under the current greedy iteration.

[0047] Since the constructed combined objective function has an approximate submodular property, the greedy strategy adopted in this disclosure can obtain the following theoretical guarantees: (1) In the standard submodular function minimization scenario, the greedy algorithm can obtain The approximate optimal solution of (2) Although It does not fully satisfy the submodule definition, but its marginal improvement shows a decreasing trend, so the greedy strategy still shows high reliability in practice; (3) The algorithm selection results disclosed in this paper are consistently superior to random sampling and continuous optimization in multiple tasks.

[0048] Step 3.2: Greedy selection algorithm enhancement mechanism; To prevent the greedy algorithm from falling into local optimality and improve the coverage of the search space, this paper designs the following two enhancement mechanisms: (1) Implementation process of multiple rounds of initialization combined with perturbation mechanism: This mechanism is mainly used to improve the diversity and robustness of greedy search in large-scale prompt word candidate sets and prevent it from falling into local optimality. The specific implementation is as follows: 1) Multiple rounds of parallel initialization: Set up several search channels (e.g. 3-5), each starting from a different initial prompt word; There are many strategies for selecting the initial cue words: such as random sampling, selecting high-frequency words, selecting words with high TF-IDF scores, or selecting words that are semantically close to the initial category.

[0049] 2) Injection of perturbation mechanism: Before each round of greedy selection begins, the current candidate word set is randomly shuffled; When selecting the word with the largest marginal benefit, if there are multiple candidate words with similar marginal benefits, the perturbed order will guide the search into different paths, thereby exploring more potential solution spaces; Stability and exploration can be balanced by setting random seeds and controlling the perturbation amplitude.

[0050] 3) Result fusion and selection: After multiple rounds of search, the set of prompt words obtained from each search path is evaluated separately (e.g., calculating the classification accuracy or target loss on the validation set); Finally, a set of prompt words with the best performance is selected as the output of the current stage.

[0051] (2) Implementation process of task adaptive weight adjustment mechanism: This mechanism addresses the problem of tip word selection being biased towards the head class under imbalanced category distribution (especially long-tail distribution), with the goal of improving the model's ability to recognize the tail class (low-frequency class). The specific implementation is as follows: 1) Category statistical modeling: Before starting prompt word selection, count each category in the training set The number of samples ; Calculate the weight of each class , for example, the inverse weight can be used ,in A small constant to prevent division by zero.

[0052] 2) Marginal gain adjustment: In the greedy algorithm, each candidate clue word is calculated Marginal gain When , the semantic relevance of the word to each category is considered (which can be measured by the similarity between CLIP text embedding and category name); Weight the marginal gain according to the weight of its associated category to obtain the adjusted marginal benefit:

[0053] in, Indicates prompt words and categories semantic relevance.

[0054] 3) Dynamic adjustment mechanism: As the training process progresses, the category distribution statistics can be dynamically updated and the weights can be adjusted in real time; In the multi-stage cue word selection, the initial stage can focus on global coverage, while the later stage can strengthen the tail class adaptability.

[0055] The computational complexity of the greedy selection algorithm disclosed in this paper is approximately , given a fixed set of candidate words and a limited number of categories, it exhibits good scalability. Compared to traditional gradient descent-based continuous optimization methods, this method does not rely on random initialization, produces stable and controllable results, and has clear semantic interpretability. The prompt words are readable natural language words. It also does not require an additional backpropagation stage, making it easy to deploy on resource-constrained platforms (such as drones). Therefore, this greedy selection algorithm not only has theoretical advantages but also meets the engineering requirements of high-performance, low-cost prompt word optimization.

[0056] Step 4: Perform local optimization and joint optimization in stages; In order to effectively integrate the discrete prompt word selection and continuous prompt word tuning process, this paper designs a staged alternating optimization mechanism to enhance the matching and generalization between prompt words and image semantics while improving the model's expressive ability.

[0057] Step 4.1: The entire cue word learning process is divided into two stages: Phase 1: Stepwise selection and local optimization During this phase, the model gradually selects semantically valid words from a set of candidate cue words. In each round, a cue word is selected and inserted into a fixed position in the predefined template. This word remains unchanged throughout subsequent training and does not participate in gradient updates. Simultaneously, only the continuous representation of the context token is optimized to adapt to the semantic information introduced by the newly added semantic words. Through this alternating process of "word selection-freezing-optimization," the model gradually adapts to the combination of semantic words, reducing misleading information caused by early selection biases and improving the final alignment between semantics and visual space.

[0058] Stage 2: Joint fine-tuning of fixed semantic words After the word selection phase is complete, the positions of all selected semantic cue words are fixed and no longer updated. The model then enters the joint fine-tuning phase, where it continues to optimize the parameters of the context tokens to improve the expressive power and classification discriminability of the overall embedding space. In this phase, all context tokens participate in the training, while the semantic cue words remain frozen.

[0059] Optimize the objective function During the entire training process, the classification loss consistent with the submodule function evaluation stage is used as the optimization target. Specifically, it is the standard contrast cross entropy loss of CLIP. (Same as step 2.1), used to measure image features and target category text features The degree of match between them.

[0060] This two-stage optimization strategy effectively balances semantic interpretability and model expressiveness. It achieves controllability and interpretability through fixed semantic terms, while leveraging continuously learnable tokens to improve model adaptability and robustness. This approach is particularly well-suited for monitoring tasks with complex semantics, scarce annotations, or blurred inter-class boundaries, offering significant advantages in stability and generalization.

[0061] Step 4.2: Training of cue word fine-tuning in continuous space; To more effectively integrate discretely selected cue words into the CLIP model's embedding space, this paper employs an iterative optimization strategy. During the cue word screening and training process, the model doesn't select all cue words at once and then optimizes them uniformly. Instead, it introduces cue words and conducts joint training in stages. Each iteration progressively optimizes the cue word representation to enhance semantic consistency and model adaptability.

[0062] Specifically, this disclosure sets a target size of 2–3 cue words for each dataset. During training, a new cue word is greedily selected and added to the current set. The model is trained for 5–10 epochs to fully adapt to the newly introduced semantic features. This process is repeated until all cue words have been selected. Afterward, the model uses the remaining training epochs to jointly fine-tune all cue words to further improve overall performance.

[0063] This alternating strategy has the following advantages: 1. Alleviate early selection bias: Gradual training enables the model to fully adjust its internal representation after each introduction of a new cue word, reducing the adverse effect of the initial cue word dominating model training.

[0064] 2. Improving semantic and representational alignment: Each newly introduced prompt word is semantically aligned through model training after being added, so that it forms a good synergy with existing prompt words and enhances the overall prompt expression ability.

[0065] 3. Improve convergence efficiency and generalization ability: The phased optimization strategy avoids the overfitting problem caused by one-time learning and achieves efficient convergence within a limited number of training rounds.

[0066] A unified classification loss is used throughout the training process As the sole optimization objective, this loss measures the degree of alignment between the image and its ground-truth category textual cue in the embedding space. By gradually strengthening this optimization objective, the cue word ultimately achieves a high degree of semantic alignment with the image category, significantly improving multimodal classification performance in complex open scenes.

[0067] Step 5: Deployment and application process; The approximate submodular function and continuous learning joint optimization strategy proposed in this disclosure can be integrated into existing multimodal visual recognition systems to improve image classification performance in environmental monitoring scenarios. It is particularly suitable for problem environments such as open environment monitoring where multimodal data is difficult to label, category distribution is uneven, and scene changes are drastic.

[0068] Example 2 In one embodiment of the present disclosure, the approximate submodular function and continuous learning joint optimization strategy proposed in the present disclosure are deployed in a modular manner in the "image semantic recognition engine" of the intelligent monitoring system. The main process includes: (1) Input: Receive video frames and high-resolution images from satellite remote sensing and drones; combine them with metadata such as acquisition time, location, and lighting information.

[0069] (2) A multimodal feature extraction module is implemented in the CLIP model, which uses pre-trained CLIP to extract image features and text features corresponding to prompt words; it supports floating-point precision and quantization deployment, and is compatible with edge computing devices.

[0070] (3) Prompt word generation and optimization module: In the offline stage, greedy sub-module selection and continuous fine-tuning process are run to generate task-specific prompt words; in the online stage, the prompt word list is loaded and the template is spliced ​​for classification reasoning.

[0071] (4) Image classification and decision module: Softmax normalize the image-text similarity results; select the most relevant category and output the pollution type, risk level or alarm signal.

[0072] (5) Back-end analysis and visualization module: Integrate the model recognition results with the GIS system to display the classification results and change trends in real time on the map or monitoring platform.

[0073] As an embodiment, the application scenario of the present disclosure is nearshore water pollution identification. The operation process of the above system modules is as follows: Step 1: Image acquisition: The monitoring terminal device periodically collects image data and uploads it to the central server; Step 2: Semantic recognition: The system loads the optimized prompt words and outputs the pollution type (such as red tide, oil pollution, floating plastic, etc.) based on the image and text comparison; Step 3: Event judgment and response. If the identification result is a high-risk category, the system triggers the early warning mechanism and notifies the operation and maintenance department; Step 4: Result recording and optimization feedback. The model recognition results are automatically archived and fine-tuned with feedback combined with manual corrections, which can be used for subsequent prompt word optimization and migration.

[0074] The application system design disclosed in this disclosure has the following practical advantages: (1) High deployment flexibility: compatible with GPU cloud deployment and lightweight edge deployment (such as Jetson, FPGA); (2) Strong data adaptability: It can process multiple sources of input, including high-resolution remote sensing images, nearshore video frames, and shipborne images; (3) Good human-computer interaction: The optimized prompt words are language interpretable, making it easier for experts to quickly understand the model behavior; (4) Rapid portability: The prompt word optimization process can be migrated and reused in other scenarios (such as inland water quality monitoring and meteorological disaster warning).

[0075] In summary, the present disclosure not only theoretically establishes a prompt word learning framework that combines approximate submodule optimization and continuous fine-tuning, but also realizes the complete implementation from method to system in the typical multimodal weak supervision scenario of environmental monitoring, with good engineering feasibility and social application value.

[0076] Example 3 In one embodiment of the present disclosure, a prompt word optimization system based on approximate submodular functions and continuous learning is provided, comprising: Obtain image data to be classified; The cue words corresponding to the image are optimized using an approximate submodular function and a continuous learning joint optimization strategy. The image data and the optimized cue words are input into a multimodal classification model for image-text matching, and the image classification result is output. Among them, the process of implementing the joint optimization strategy of approximate submodular function and continuous learning for prompt words includes: constructing a set of candidate prompt words, designing a combined objective function based on the properties of approximate submodular function, and solving the combined objective function using a greedy selection algorithm that combines random perturbation, multiple rounds of iteration and task adaptation mechanism to achieve optimal selection of the candidate prompt word set, gradually selecting the prompt word with the largest gain from the candidate prompt word set and adding it to the optimized subset. Each time a new prompt word is selected, a local optimization is performed. When all target prompt words are selected, all prompt words are jointly optimized in the continuous space using an alternating optimization strategy, and iterate continuously until the optimized prompt word is obtained.

[0077] Example 4 In one embodiment of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method for optimizing prompt words based on approximate submodular functions and continuous learning is implemented.

[0078] Example 5 In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the prompt word optimization method based on approximate submodular functions and continuous learning is implemented.

[0079] Example 6 In one embodiment of the present disclosure, an electronic device is provided, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device implements the prompt word optimization method based on approximate submodular functions and continuous learning.

[0080] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0082] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A prompt word optimization method based on approximate submodular functions and continuous learning, characterized in that: include: Obtain image data to be classified; Utilize the approximate submodular function and continuous learning joint optimization strategy to optimize the prompt words corresponding to the image; The image data and the optimized prompt words are input into the multimodal classification model for image-text matching, and the image classification result is output; Among them, the process of implementing the joint optimization strategy of approximate submodular function and continuous learning for prompt words includes: constructing a set of candidate prompt words, designing a combined objective function based on the properties of approximate submodular function, and solving the combined objective function using a greedy selection algorithm that combines random perturbation, multiple rounds of iteration and task adaptation mechanism to achieve optimal selection of the candidate prompt word set, gradually selecting the prompt word with the largest gain from the candidate prompt word set and adding it to the optimized subset. Each time a new prompt word is selected, a local optimization is performed. When all target prompt words are selected, all prompt words are jointly optimized in the continuous space using an alternating optimization strategy, and iterate continuously until the optimized prompt word is obtained.

2. The prompt word optimization method based on approximate submodular function and continuous learning according to claim 1, characterized in that: Extract candidate prompt words corresponding to the image, use the BPE word segmenter to encode and filter the candidate prompt words, retain the prompt words encoded as a single token, and the retained prompt words constitute the candidate prompt word set. Each prompt word is a semantically valid, moderately frequent atomic term that has passed multiple screenings and can be directly processed by CLIP.

3. The prompt word optimization method based on approximate submodular function and continuous learning according to claim 1, characterized in that: Combining task loss and semantic diversity, a combined objective function based on the properties of approximate submodular functions is designed. The combined objective function based on the properties of approximate submodular functions consists of a classification performance term and a semantic diversity regularization term, and the combined loss exhibits the properties of an approximate submodular function, including the fact that as the set of candidate prompt words expands, the improvement effect of newly added prompt words on the combined loss shows a marginal decreasing trend, and for a certain prompt word, the gain for a small set is higher than that for a large set.

4. The prompt word optimization method based on approximate submodular function and continuous learning according to claim 1, characterized in that: The goal of solving the combined objective function by using a greedy selection algorithm that combines random perturbations, multiple rounds of iterations, and a task adaptation mechanism is to select an optimized subset from the candidate prompt word set without exceeding a set length, so that the combined loss of the combined objective function is minimized. The greedy selection algorithm process includes: initializing the optimized subset of prompt words to an empty set. In each round of iteration, all candidate prompt words are traversed and the loss reduction brought about by adding them to the optimized subset is calculated. The candidate prompt word that brings the largest loss reduction is selected and added to the current optimized subset. The iteration is repeated until the maximum number of words is reached or the loss reduction is less than a preset threshold. The iteration is stopped to obtain the optimized subset under the current greedy iteration.

5. The prompt word optimization method based on approximate submodular function and continuous learning according to claim 4, characterized in that: To prevent the greedy algorithm from falling into local optimality and improve the coverage of the search space, two enhancement mechanisms are designed: adding random perturbations, starting the greedy process in parallel from different starting points, and randomly shuffling the candidate prompt words before each iteration to introduce perturbations to improve the diversity of search paths; In the marginal gain calculation, the category distribution information is integrated to weight the prompt words and perform adaptive weight adjustment.

6. The prompt word optimization method based on approximate submodular function and continuous learning according to claim 1, characterized in that: In the initial stage, a greedy algorithm is used to gradually select the cue word with the largest gain from the candidate cue word set. Whenever a new cue word is selected, the existing cue words are frozen and only local fine-tuning training is performed on the new cue word. After all target cue words are selected, the joint optimization stage is started, all cue word parameters are released, and joint fine-tuning is performed in the continuous embedding space. The joint fine-tuning process uses the same classification loss as the approximate submodular function evaluation process, namely the comparative cross-entropy loss. This loss is used to measure the similarity between image features and the true category cue word. The comparative cross-entropy loss is used as the only optimization target. Through the gradual strengthening of this optimization target, the cue word can eventually be highly semantically consistent with the image category.

7. A prompt word optimization system based on approximate submodular functions and continuous learning, characterized by: include: A data acquisition module, used to acquire image data to be classified; The image-text matching classification module is used to optimize the prompt words corresponding to the image using an approximate submodular function and a continuous learning joint optimization strategy; the image data and the optimized prompt words are input into the multimodal classification model for image-text matching, and the image classification result is output; Among them, the process of implementing the joint optimization strategy of approximate submodular function and continuous learning for prompt words includes: constructing a set of candidate prompt words, designing a combined objective function based on the properties of approximate submodular function, and solving the combined objective function using a greedy selection algorithm that combines random perturbation, multiple rounds of iteration and task adaptation mechanism to achieve optimal selection of the candidate prompt word set, gradually selecting the prompt word with the largest gain from the candidate prompt word set and adding it to the optimized subset. Each time a new prompt word is selected, a local optimization is performed. When all target prompt words are selected, all prompt words are jointly optimized in the continuous space using an alternating optimization strategy, and iterate continuously until the optimized prompt word is obtained.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the prompt word optimization method based on approximate submodular function and continuous learning according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the prompt word optimization method based on approximate submodular function and continuous learning is implemented as described in any one of claims 1 to 6.

10. An electronic device, characterized in that: include: A processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the prompt word optimization method based on approximate submodular function and continuous learning as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Visual language model training method, device, medium and computer program product

    CN118520933A

  • Large language model optimization generation method based on optimal cue word selection

    CN119476209A

  • Large language model discrete cue word searching method and device

    CN120296148A

  • Semantic understanding method and device

    WO2025000856A1

Cited By

  • Large language model fingerprint construction method and system based on maximum activation component coverage

    CN121706145A

  • A large language model fingerprint construction method and system based on maximum active component coverage

    CN121706145B

  • Crowd sensing task allocation method and system based on cue word and deduction decoupling

    CN122242772A

  • A crowd-sensing task allocation method and system based on decoupling of prompt words and inference

    CN122242772B