Product style classification method, system and equipment for optimizing multi-modal large model Prompt

By using the Qwen-VL multimodal large model and genetic algorithm to optimize the prompt, the problems of low efficiency and insufficient generalization ability in traditional product style classification are solved, and more accurate product style classification and multimodal information understanding are achieved.

CN120931985APending Publication Date: 2025-11-11SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510957451.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies rely on manual annotation in product style classification, which is inefficient and highly subjective. Convolutional neural networks cannot accurately classify different modal information and lack generalization ability.

Method used

The Qwen-VL multimodal large model is adopted, combined with the prompt optimization mechanism, thinking chain technology and genetic algorithm. Visual features are extracted through a large language model, visual encoder and position-aware visual language adapter, multimodal information is fused, and the instruction part of the prompt is optimized through genetic algorithm to generate the optimal instruction.

Benefits of technology

It achieves more accurate product style classification, improves the model's generalization ability and the interpretability of label generation, reduces classification errors, and enhances the ability to understand information from different modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931985A_ABST
    Figure CN120931985A_ABST
Patent Text Reader

Abstract

The invention discloses a product style classification method, system and device for optimizing a multi-modal large model Prompt, and relates to the technical field of product style classification. The method comprises the steps that a user product image is received and preprocessed; a classification model framework is constructed based on a Qwen-VL network architecture, Qwen-VL is adopted as a core model and used for extracting visual features of product images and fusing multi-modal information, and a Qwen-VL model comprises a large language model, a visual encoder and a position sensing visual language adapter. The multi-modal large model can process and understand information from different modals at the same time, the expression ability of the model is enhanced through feature fusion, complex information is more comprehensively understood and processed, the style features of products can be extracted and analyzed from the images for the current product design field, more accurate product style classification is achieved, and the method is suitable for the field of product design. And a thinking chain and a genetic algorithm are adopted to jointly optimize the prompt, so that the solvability and the stability of tag generation can be greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of product style classification technology, specifically to a product style classification method, system, and device for optimizing multimodal large model Prompt. Background Technology

[0002] With the advent of Industry 5.0, industrial production is increasingly transforming towards intelligence, personalization, and sustainable development, placing greater emphasis on collaboration between humans and artificial intelligence, as well as on personalization and user experience. Against this backdrop, consumer demand for personalized and unique product designs is constantly growing, making traditional, standardized production models increasingly unable to meet market changes. This is prompting a shift in design and production processes across various product sectors towards intelligence and a human-centered approach.

[0003] In the field of product design, current image processing and utilization methods mainly focus on extracting surface information such as color, shape, and pattern from images. They are significantly insufficient in extracting higher-level information such as product style. However, utilizing images to obtain product style information can not only improve the targeting of designs and enhance the consumer shopping experience, but also help optimize production processes, reduce waste, and produce products that better meet market demands, thereby reducing environmental impact. Image style classification technology, due to its ability to effectively connect design and consumer preferences, has become one of the key technologies for intelligent product design. It not only responds to the transformation needs of Industry 5.0, but is also an important technical means to achieve personalization and sustainable development.

[0004] However, in the traditional product style research stage, classification work mainly relies on manual annotation and expert opinions. Although it has a certain degree of professionalism, it has limitations such as low efficiency and strong subjectivity when processing large-scale data. The development of convolutional neural networks (CNN) has made significant breakthroughs in image classification. However, relying solely on visual features extracted by convolutional layers cannot classify completely accurately, and it is prone to poor adaptability and overfitting risks when dealing with different modal information. Therefore, this invention proposes a product style classification method, system, and device that optimizes the multimodal large model Prompt. Summary of the Invention

[0005] The purpose of this invention is to provide a product style classification method, system, and device for optimizing multimodal large model Prompt, which aims to extract and analyze product style features from images to achieve more accurate product style classification, and can simultaneously process and understand information from different modalities, thereby improving the generalization ability of the model.

[0006] According to a first aspect of the present invention, in order to achieve the above-mentioned objective, the present invention provides the following technical solution: an optimized product style classification method for multimodal large model Prompt, comprising the following steps:

[0007] Receive user product images and perform preprocessing;

[0008] A classification model framework is constructed using the Qwen-VL multimodal large model to extract visual features of product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter.

[0009] The Qwen-VL model is based on the introduction of a prompt optimization mechanism, which consists of two parts: instruction and demonstration.

[0010] Demonstration is constructed using the mind chain technology to inject expert knowledge in the product style domain into the Qwen-VL model and generate examples that include the reasoning process.

[0011] The instruction part of the prompt is optimized using a genetic algorithm to generate the optimal instruction for the product style classification task.

[0012] The preprocessed user product images and the optimized prompts are input into the Qwen-VL model to output product style classification results.

[0013] Furthermore, a classification model framework is constructed using the Qwen-VL multimodal large model to extract visual features from product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter, as detailed below:

[0014] (21) Large language model: Qwen-VL uses a large language model as its basic component. The large language model is initialized with pre-trained weights from Qwen-7B to process and analyze the input text information, providing a language foundation for subsequent visual language interaction.

[0015] (22) Visual encoder: The visual encoder of Qwen-VL uses the Vision Transformer (ViT) architecture and is initialized with the pre-trained weights of ViT-bigg from Openclip. During training and inference, the input image is adjusted to a specific resolution and then the image is segmented into small blocks with a stride of 14 to process the image and generate a set of image features, thereby transforming the visual information into a feature representation that the Qwen-VL model can process, enabling the Qwen-VL model to understand and process visual signals.

[0016] (23) Position-aware visual language adapter: The position-aware visual language adapter includes a randomly initialized single-layer cross-attention module that uses a set of trainable vectors as query vectors and uses the image features generated by the visual encoder as the key to the cross-attention operation. It compresses the visual feature sequence to 256 dimensions and introduces 2D absolute position encoding to mitigate the position details that may be lost during the compression process.

[0017] Furthermore, a demonstration is constructed using mind chain technology, as follows:

[0018] (31) Define the key dimension set D = {d1, d2, ..., d...} for product style classification. n}, where d i This indicates specific dimensions, including style, pattern, and color combination;

[0019] (32) For each dimension d i Establish style tags through expert knowledge j The associative mapping function: f(d) i →{s1,s2,…,s m};

[0020] (33) Select representative product examples. For each example product x, extract the feature vector V from the key dimension D. x =(v x,1 ,v x,2 ,…,v x,n ), where v x,i This indicates that the dress x is in dimension d. i Eigenvalues ​​on;

[0021] (33) Through the reasoning process R, the feature vector V x Mapping to style tags j :

[0022]

[0023] Where w i It is dimension d i The weights are denoted by sim, where sim represents the similarity measurement function.

[0024] (34) Based on the reasoning process R, generate style tags s for each example product x. j And reason about the basis of the record:

[0025]

[0026] Where θ is the confidence threshold;

[0027] (35) Input the example containing the reasoning process into Qwen-VL, guide the language big model to focus on the specific style features of the product, generate the product style reasoning process, and obtain the final result.

[0028] Furthermore, the instruction part of the prompt is optimized using a genetic algorithm, as follows:

[0029] (41) Initializing the population: The initial design prompt includes automatically generated model prompts and expert-built prompts;

[0030] (42) Selection: Imitating the natural selection process, the roulette wheel selection method is adopted. The accuracy of the initial group prompts on the development set is used as the fitness of the corresponding prompts. Based on the fitness of each prompt, a certain number of prompts are selected as the parents of the next generation.

[0031] (43) Crossover and mutation: In each iteration, the large model is treated as a genetic operator to generate a new population of prompts from the parent prompts in step (42);

[0032] Crossover: Cross over parts of the text or semantics of the parent prompts:

[0033] Mutation: Semantically delete or replace the statements that constitute the prompt in the offspring prompts population, with a certain probability to achieve mutation;

[0034] (43) New population: In each iteration, N new prompts are generated and merged with the existing N prompts population. They are evaluated on the development set, and the top N prompts are retained according to the scores to form a new generation population. This process is repeated until the number of iterations reaches a predetermined value, and the optimal instruction is obtained.

[0035] Furthermore, the preprocessed user product images and the optimized prompt are input into the Qwen-VL model to output product style classification results, as follows:

[0036] (51) The received user product images and the best prompt are used as input, and the Qwen-VL multimodal large model is used for image style classification.

[0037] (52) Qwen-VL first extracts the visual features of the image through Vision Transformer, segments the image into patches and extracts spatial features. The first four layers use a full attention mechanism to build a global semantic map, and the subsequent layers use a window attention mechanism to focus on local details. The text description is transformed into semantic features through Token embedding similar to BERT.

[0038] (53) The extracted visual features are projected onto the same dimension as the semantic features through a linear layer for alignment, and then deep interactive fusion is performed through MLP-based compression and fusion and a position-aware cross-attention module.

[0039] (54) Finally, the fused features are input into the classifier, and the classification result is output by the fully connected layer to realize product style classification.

[0040] According to a second aspect of the present invention, the present invention provides a product style classification system for a multimodal large model Prompt, used to implement the product style classification method for optimizing a multimodal large model Prompt described in the first aspect, comprising:

[0041] The receiving module is used to receive user product images and perform preprocessing.

[0042] The building module is used to construct a classification model framework using the Qwen-VL multimodal large model, which is used to extract visual features of product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter.

[0043] The Prompt optimization module is used to introduce a prompt optimization mechanism into the constructed Qwen-VL model. The Prompt consists of two parts: instruction and demonstration.

[0044] Demonstration is constructed using the mind chain technology to inject expert knowledge in the product style domain into the Qwen-VL model and generate examples that include the reasoning process.

[0045] The instruction part of the prompt is optimized using a genetic algorithm to generate the optimal instruction for the product style classification task.

[0046] The classification output module is used to input the preprocessed user product image and the optimized prompt into the Qwen-VL model and output the product style classification result.

[0047] According to a third aspect of the present invention, a terminal device is provided, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores the computer program capable of running on the processor, and when the processor loads and executes the computer program, it employs the product style classification method of the optimized multimodal large model Prompt described in the first aspect.

[0048] According to a fourth aspect of the present invention, the present invention provides a storage medium containing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to perform a product style classification method as described in the first aspect of the optimized multimodal large model Prompt.

[0049] This invention has at least the following beneficial effects:

[0050] 1. The classification method of this invention is dedicated to extracting and analyzing the style characteristics of products from images to achieve more accurate product style classification, which can provide a scientific basis for subsequent product development and market positioning;

[0051] 2. This invention addresses the problems of traditional product image classification algorithms, which are often based on a single modality, require large-scale data for training, and lack generalization ability. It proposes a multimodal large-scale model capable of simultaneously processing and understanding information from different modalities (such as images, text, and speech). Feature fusion enhances the model's expressive power, enabling a more comprehensive understanding and processing of complex information. Furthermore, large-scale data pre-training significantly improves the model's generalization ability, reducing classification errors caused by differences in data distribution.

[0052] 3. This invention addresses the problems of poor label interpretability and difficulty in prompt design for large models in specific domains. It employs a combination of CoT (Co-T) and genetic algorithm to optimize the prompt. CoT can greatly enhance the solvability and stability of label generation, while the genetic algorithm (GA) can utilize currently performing well-performing prompt combinations to generate new schemes through crossover and mutation operations.

[0053] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0054] Figure 1 This is a flowchart illustrating the classification method described in this invention;

[0055] Figure 2 This is a schematic diagram of the structural principle of the multimodal large model framework in this invention;

[0056] Figure 3 This is a schematic diagram of the genetic algorithm optimization suggestion generation process in this invention;

[0057] Figure 4 This is an example diagram illustrating product style reasoning in this invention. Detailed Implementation

[0058] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0059] Example 1:

[0060] The product style classification method for optimizing the multimodal large-scale model's prompt uses Qwen-VL as the core model and leverages Prompt Engineering to design and optimize the prompt to guide the large-scale model's output. Specifically, it utilizes the thought chain technique to construct demonstrations to optimize model interpretability; and employs a genetic algorithm to optimize the instruction part of the prompt to find the optimal prompt, such as... Figure 1 As shown.

[0061] Please see Figures 1-3 This invention provides a technical solution: an optimized product style classification method for multimodal large-model Prompt, comprising the following steps:

[0062] S1. Receive the user's product image and perform preprocessing, as follows:

[0063] First, select a clear image of the product's front. Then, adjust the image resolution to 448*448 pixels, ensuring the image is in RGB three-channel format. Next, normalize the pixel values ​​from [0, 255] to the model training distribution, using Qwen-VL specific parameters: mean =

[0064] [0.48145466, 0.4578275, 0.40821073], standard deviation std = [0.26862954,

[0065] [0.26130258, 0.27577711], (the normalization formula is: (pixel value / 255-mean) / std), and finally the processed tensor (format [C,H,W]) is expanded to [1,C,H,W] to fit the model input;

[0066] S2. A classification model framework is constructed using the Qwen-VL multimodal large model to extract visual features of product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter.

[0067] Figure 2 The overall network architecture of Qwen-VL consists of three parts: a large language model, a visual encoder, and a position-aware visual-language adapter.

[0068] Large Language Model (LLM): Qwen-VL uses a large language model as its basic component. This model is initialized using pre-trained weights from Qwen-7B. Therefore, this model has powerful language understanding and generation capabilities, and can process and analyze input text information, providing a language foundation for subsequent visual language interaction.

[0069] Visual encoder: Qwen-VL's visual encoder uses the Vision Transformer (ViT) architecture and is initialized with pre-trained weights from Openclip's ViT-bigg. During training and inference, the input image is adjusted to a specific resolution and then the image is segmented into small blocks with a stride of 14 to process the image and generate a set of image features, thereby transforming visual information into feature representations that the model can process, enabling the model to understand and process visual signals.

[0070] Position-Aware Visual-Language Adapter: To alleviate the efficiency issues caused by long image feature sequences, Qwen-VL introduces a visual-language adapter that compresses image features. This adapter includes a randomly initialized single-layer cross-attention module that uses a set of trainable vectors (embeddings) as query vectors and image features from the visual encoder as the key for the cross-attention operation. This mechanism compresses the visual feature sequence to 256 dimensions. Furthermore, considering the importance of positional information for fine-grained image understanding, to mitigate potential loss of positional details during compression, 2D absolute position encoding is introduced into the query key pair of the cross-attention mechanism. Subsequently, the 256-dimensional compressed image feature sequence is input into a large-scale language model, achieving the fusion of visual and linguistic features, enabling the model to better understand the relationship between visual content and text.

[0071] S3. Based on the constructed Qwen-VL model, a prompt optimization mechanism is introduced, which includes two parts: instruction and demonstration.

[0072] S31. Construct a demonstration using the mind chain technique to inject expert knowledge in the product style domain into the Qwen-VL model and generate examples that include the reasoning process;

[0073] The main idea of ​​the thought chain is to show a small number of examples to a large language model and explain the reasoning process in the examples, guiding the LLM to provide the reasoning process when answering and arrive at the final result. When the large model is performing classification tasks, the thought chain can be manually constructed, that is, the reasoning process constructed by domain experts is added to the examples by humans to optimize context learning. This method is also known as Few-shot CoT, as follows:

[0074] (S31.1) Define the key dimension set D = {d1, d2, ..., d...} for product style classification. n}, where d i This indicates specific dimensions, including style, pattern, and color combination;

[0075] (S31.2) For each dimension d i Establish style tags through expert knowledge j The associative mapping function: f(d) i →{s1,s2,…,s j};

[0076] (S31.3) Select representative product examples. For each example product x, extract the feature vector V from the key dimension D. x =(v x,1 ,v x,2 ,…,v x,n ), where v x,i This indicates that the dress x is in dimension d. i Eigenvalues ​​on;

[0077] (S31.3) Through the reasoning process R, the feature vector V is... x Mapping to style tags j :

[0078]

[0079] Where w i It is dimension d i The weights are denoted by sim, where sim represents the similarity measure function.

[0080] (S31.4) Based on the reasoning process R, generate style tags s for each example product x. j And reason about the basis of the record:

[0081]

[0082] Where θ is the confidence threshold;

[0083] (S31.5) Input the example containing the reasoning process into Qwen-VL, guide the language big model to focus on the specific style features of the product, generate the product style reasoning process, and obtain the final result;

[0084] Regarding the technical solution of this embodiment, taking the field of clothing products as an example, the specific process is as follows:

[0085] First, we need to clarify the key dimensions of dress style classification. These dimensions are the basis for reasoning and can help reveal the unique patterns and rules of different styles. We also need to collect expert knowledge from fashion experts and clothing designers about the relationship between each key dimension and different dress styles, and understand which common style types each attribute corresponds to, so as to provide an accurate basis for building examples later.

[0086] Let the key dimension for classifying dress styles be set D = {d1, d2, ..., d...} n}, where d i This represents a specific dimension (such as style, pattern, color combination, etc.), with each dimension d. i With style tags j The association can be modeled as a mapping function using expert knowledge: f(d i →{s1,s2,…,s j}, where s j These are style labels (such as "mass casual style" or "simple and mature style").

[0087] Secondly, representative dresses are selected as examples, ensuring that these dresses have distinct stylistic characteristics. For each example dress x, a feature vector V is extracted from the key dimension D. x

[0088] =(v x,1 ,v x,2 ,…,v x,n ), where v x,i This indicates that the dress x is in dimension d. i Eigenvalues ​​on;

[0089] Through the inference process R, the feature vectors are mapped to style labels:

[0090]

[0091] Where w i It is dimension d i The weights are denoted by sim, where sim represents the similarity measurement function.

[0092] Then, based on the reasoning process R, style tags s are generated for each example dress x. jAnd record the reasoning basis, such as:

[0093]

[0094] Where θ is the confidence threshold;

[0095] For example, A-line skirts, with their narrow top and wide bottom shape, give a youthful and lively feel. From a color perspective, classic black and white is more suitable for formal occasions while maintaining a sense of style, often associated with elegance and sophistication. Based on the above multi-dimensional analysis, each example skirt is labeled with an accurate style tag, such as "mass casual style" or "simple mature style," and the basis for choosing this style tag is clearly shown in the example. Finally, the constructed example containing the reasoning process is input into Qwen-VL, guiding the LLM to focus on the specific style characteristics of the image, perform style analysis on the clothing image within the domain, generate the skirt style reasoning process, and finally obtain the result: Output = Qwen-VL(R(v x ));

[0096] S32. Optimize the instruction part of the prompt using a genetic algorithm to generate the optimal instruction suitable for the product style classification task;

[0097] Genetic algorithms are optimization search algorithms that simulate natural selection and genetic mechanisms. They aim to optimize prompts for specific downstream tasks to maximize the potential of LLMs on that task. In the process of constructing prompts, Qwen-VL generates new candidate prompts based on genetic operators, while the genetic algorithm guides the optimization process to retain the best prompts. Genetic algorithms typically start with N initially designed prompts, and then use genetic operators (such as mutation and crossover) on the current population to iteratively generate new prompt word templates and update them according to the fitness function.

[0098] It mainly includes four steps: initial population - selection - crossover and mutation - new population.

[0099] 1. Initial Population: The initial design prompts include both automatically generated and expert-built models;

[0100] Most existing automatic suggestion methods ignore human prior knowledge. Therefore, this embodiment also uses existing human (expert) suggestions as the initial group to fully inject domain knowledge. In addition, EA usually starts from random solutions to generate a diverse group and avoid getting trapped in local optima. Therefore, this embodiment also introduces some suggestions generated by large models into the initial group.

[0101] 2. Selection: Mimicking the natural selection process, the Roulette Wheel Selection method is adopted. The accuracy of the initial group's prompts on the development set is used as the fitness of the corresponding prompt. Based on the fitness of each prompt, a certain number of prompts are selected as the parents of the next generation. i Let represent the performance score of the i-th prompt in a group containing N prompts. Selecting the i-th prompt as the parent prompt can be represented as... Individuals with high fitness are more likely to be selected, thus passing on their superior genes to the next generation;

[0102] 3. Crossover and mutation: In each iteration, the large model is treated as a genetic operator, generating a new population of prompts from the parent prompts in the second step;

[0103] Crossover is achieved by crossing parts of the text or semantics between two parent individuals (similar to chromosome crossover in genetics);

[0104] Mutation involves randomly altering the genes of an individual to semantically delete or replace the constituent statements of the prompts in the offspring's prompts population with a certain probability. To achieve this goal, this paper designs the steps and corresponding instructions of mutation and crossover operators to guide the large model in generating new prompts based on these steps. The entire process is as follows: Figure 3 As shown.

[0105] 4. New Population: In each iteration, Evo generates N new prompts and merges them with the existing population of N prompts. The prompts are evaluated on the development set, and the top N prompts are retained based on their scores to form a new generation of population. This process is repeated until the number of iterations reaches a predetermined value, at which point the algorithm stops.

[0106] S4. Input the preprocessed user product image and the optimized prompt into the Qwen-VL model to output the product style classification result;

[0107] Taking user-provided images or image sets and the best prompt as input, the Qwen-VL multimodal large model is used for image style classification. Qwen-VL first extracts visual features from the image using a Vision Transformer, segmenting the image into patches and extracting spatial features. The first four layers use a full attention mechanism to build a global semantic map, while subsequent layers use a window attention mechanism to focus on local details. Text descriptions are transformed into semantic features through BERT-like token embedding. Next, the visual features are projected onto the same dimension as the semantic features through linear layers for alignment. Then, deep interactive fusion is performed through MLP-based compression and fusion, as well as a position-aware cross-attention module. Finally, the fused features are input to the classifier, and the fully connected layer outputs the classification result, achieving product style classification.

[0108] The technical solution of the present invention will be further described below with reference to specific embodiments:

[0109] The functional requirements of the product style classification method based on genetic algorithms and thought chain optimization for multimodal large models (Prompt) can be summarized in the following three points:

[0110] The user uploads product images (image set). This method requires the user to upload product images and extract their visual features. This function allows the user to upload product images that need to be classified and preprocesses the uploaded images. Therefore, there are no special requirements for the image's size, proportion, and other attributes.

[0111] To obtain the optimal prompt, the prompt consists of two parts: instruction and demonstration.

[0112] The purpose of demonstration is to input specific product style knowledge into a large model and provide examples based on each style dimension, such as... Figure 4 As shown, the complete thought chain reasoning process is demonstrated and input into a multimodal large model to guide model learning and simulation. The purpose of the instruction is to explicitly tell the model that it needs to complete the task of style classification based on images. The initial instruction is used as the initial population of the genetic algorithm. Then, the fitness score of each prompt is evaluated based on the style classification accuracy output by the large model. The score is based on the confidence returned by the model and the multimodal context matching degree. Based on the fitness, the prompts with better performance are selected as parents. Subsequently, the information of the parent prompts is reorganized or some of the offspring are slightly semantically adjusted or keywords are replaced to generate new offspring prompts. The genetic algorithm stops after a set number of generations or when the fitness reaches a threshold, and finally the optimal instruction is obtained.

[0113] In summary, this invention addresses the problems of traditional product image classification algorithms, which are often based on a single modality, require large-scale data for training, and lack generalization ability. The multimodal large-scale model can simultaneously process and understand information from different modalities (such as images, text, and speech), and enhances the model's expressive power through feature fusion, enabling a more comprehensive understanding and processing of complex information. Simultaneously, large-scale data pre-training significantly improves the model's generalization ability, reducing classification errors caused by differences in data distribution. Furthermore, this invention addresses the issues of poor label interpretability and difficulty in prompt design in specific domains by employing a combination of CoT (Co-T) and genetic algorithms to jointly optimize the prompt. CoT can greatly enhance the solvability and stability of label generation, while the genetic algorithm (GA) can utilize combinations of currently performing prompts to generate new solutions through crossover and mutation operations.

[0114] Example 2:

[0115] This embodiment provides a product style classification system for a multimodal large model Prompt, used to implement the optimized multimodal large model Prompt product style classification method described in Embodiment 1, including:

[0116] The receiving module is used to receive user product images and perform preprocessing.

[0117] The building module is used to construct a classification model framework using the Qwen-VL multimodal large model, which is used to extract visual features of product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter.

[0118] The prompt optimization module is used to introduce a prompt optimization mechanism into the constructed Qwen-VL model. The prompt consists of two parts: instruction and demonstration.

[0119] Demonstration is constructed using the mind chain technology to inject expert knowledge in the product style domain into the Qwen-VL model and generate examples that include the reasoning process.

[0120] The instruction part of the prompt is optimized using a genetic algorithm to generate the optimal instruction for the product style classification task.

[0121] The classification output module is used to input the preprocessed user product image and the optimized prompt into the Qwen-VL model and output the product style classification result.

[0122] Example 3:

[0123] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores the computer program capable of running on the processor. When the processor loads and executes the computer program, it adopts the product style classification method of the optimized multimodal large model Prompt described in Embodiment 1.

[0124] It should be noted that the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server, and the terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may also include input / output devices, network access devices, and buses.

[0125] Furthermore, the processor can be a central processing unit (CPU). Of course, depending on the actual use, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be used. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it in this regard.

[0126] Example 4:

[0127] The present invention provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the product style classification method of the optimized multimodal large model Prompt described in Embodiment 1.

[0128] The computer program can be stored in a computer-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or certain middleware. The computer-readable medium includes any entity or device capable of carrying computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the computer-readable medium includes, but is not limited to, the above-mentioned components.

[0129] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0130] For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances. When an element is referred to as being "assembled on," "mounted on," "fixed to," or "set on" another element, it may be directly on the other element or there may be an intermediate element present. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be an intermediate element present. The terms "vertical," "horizontal," "upper," "lower," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only possible embodiments.

[0131] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

[0132] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

Claims

1. An optimized product style classification method for multimodal large-scale model Prompt, characterized in that, Includes the following steps: Receive user product images and perform preprocessing; A classification model framework is constructed using the Qwen-VL multimodal large model to extract visual features of product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter. The Qwen-VL model is based on the introduction of a prompt optimization mechanism, which consists of two parts: instruction and demonstration. Demonstration is constructed using the mind chain technology to inject expert knowledge in the product style domain into the Qwen-VL model and generate examples that include the reasoning process. The instruction part of the prompt is optimized using a genetic algorithm to generate the optimal instruction for the product style classification task. The preprocessed user product images and the optimized prompts are input into the Qwen-VL model to output product style classification results.

2. The product style classification method for optimizing multimodal large model Prompt according to claim 1, characterized in that: A classification model framework is constructed using the Qwen-VL multimodal large model to extract visual features from product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter, as detailed below: (21) Large language model: Qwen-VL uses a large language model as its basic component. The large language model is initialized with pre-trained weights from Qwen-7B to process and analyze the input text information, providing a language foundation for subsequent visual language interaction. (22) Visual encoder: The visual encoder of Qwen-VL uses the Vision Transformer (ViT) architecture and is initialized with the pre-trained weights of ViT-bigg from Openclip. During training and inference, the input image is adjusted to a specific resolution and then the image is segmented into small blocks with a stride of 14 to process the image and generate a set of image features, thereby transforming the visual information into a feature representation that the Qwen-VL model can process, enabling the Qwen-VL model to understand and process visual signals. (23) Position-aware visual language adapter: The position-aware visual language adapter includes a randomly initialized single-layer cross-attention module that uses a set of trainable vectors as query vectors and uses the image features generated by the visual encoder as the key to the cross-attention operation. It compresses the visual feature sequence to 256 dimensions and introduces 2D absolute position encoding to mitigate the position details that may be lost during the compression process.

3. The product style classification method for optimizing multimodal large model Prompt according to claim 2, characterized in that: The demonstration is constructed using the mind chain technique, as follows: (31) Define the key dimension set D = {d1, d2, ..., d...} for product style classification. n }, where d i This indicates specific dimensions, including style, pattern, and color combination; (32) For each dimension d i Establish style tags through expert knowledge j The associative mapping function: f(d) i →{s1,s2,…,s j }; (33) Select representative product examples. For each example product x, extract the feature vector V from the key dimension D. x =(v x,1 ,v x,2 ,…,v x,n ), where v x,i This indicates that the dress x is in dimension d. i Eigenvalues ​​on; (33) Through the reasoning process R, the feature vector V x Mapping to style tags j : Where w i It is dimension d i The weights are denoted by sim, where sim represents the similarity measure function. (34) Based on the reasoning process R, generate style tags s for each example product x. j And reason about the basis of the record: Where θ is the confidence threshold; (35) Input the example containing the reasoning process into Qwen-VL, guide the language big model to focus on the specific style features of the product, generate the product style reasoning process, and obtain the final result.

4. The product style classification method for optimizing multimodal large model Prompt according to claim 3, characterized in that: The instruction part of the prompt is optimized using a genetic algorithm, as follows: (41) Initializing the population: The initial design prompts include automatically generated model prompts and expert-built prompts; (42) Selection: Imitating the natural selection process, the roulette wheel selection method is adopted. The accuracy of the initial group prompts on the development set is used as the fitness of the corresponding prompts. Based on the fitness of each prompt, a certain number of prompts are selected as the parents of the next generation. (43) Crossover and mutation: In each iteration, the large model is treated as a genetic operator to generate a new population of prompts from the parent prompts in step (42); Crossover: Cross over parts of the text or semantics of the parent prompts: Mutation: Semantically delete or replace the statements that constitute the prompt in the offspring prompts population, with a certain probability to achieve mutation; (43) New population: In each iteration, N new prompts are generated and merged with the existing N prompts population. They are evaluated on the development set, and the top N prompts are retained according to the scores to form a new generation population. This process is repeated until the number of iterations reaches a predetermined value, and the optimal instruction is obtained.

5. The product style classification method for optimizing multimodal large model Prompt according to claim 4, characterized in that: The preprocessed user product image and the optimized prompt are input into the Qwen-VL model, and the product style classification results are output as follows: (51) The received user product images and the best prompt are used as input, and the Qwen-VL multimodal large model is used for image style classification. (52) Qwen-VL first extracts the visual features of the image through Vision Transformer, segments the image into patches and extracts spatial features. The first four layers use a full attention mechanism to build a global semantic map, and the subsequent layers use a window attention mechanism to focus on local details. The text description is transformed into semantic features through Token embedding similar to BERT. (53) The extracted visual features are projected onto the same dimension as the semantic features through a linear layer for alignment, and then deep interactive fusion is performed through MLP-based compression and fusion and a position-aware cross-attention module. (54) Finally, the fused features are input into the classifier, and the classification result is output by the fully connected layer to realize product style classification.

6. A product style classification system for a multimodal large model Prompt, used to implement the product style classification method for the optimized multimodal large model Prompt as described in any one of claims 1 to 5, characterized in that, include: The receiving module is used to receive user product images and perform preprocessing. The building module is used to construct a classification model framework using the Qwen-VL multimodal large model, which is used to extract visual features of product images and fuse multimodal information. The Qwen-VL model includes a large language model, a visual encoder, and a position-aware visual language adapter. The Prompt optimization module is used to introduce the Prompt optimization mechanism into the constructed Qwen-VL model. The Prompt consists of two parts: instruction and demonstration. Demonstration is constructed using the mind chain technology to inject expert knowledge in the product style domain into the Qwen-VL model and generate examples that include the reasoning process. The instruction part of the prompt is optimized using a genetic algorithm to generate the optimal instruction for the product style classification task. The classification output module is used to input the preprocessed user product image and the optimized prompt into the Qwen-VL model and output the product style classification result.

7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it employs the product style classification method of the optimized multimodal large model Prompt as described in any one of claims 1 to 5.

8. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the product style classification method of the optimized multimodal large model Prompt as described in any one of claims 1 to 5.