Large model-based scene retrieval method and terminal
Generate the question-answer text through a large language model and train the multimodal matching model, the problem that traditional single-modal retrieval cannot obtain a complete description across media types is solved, and multimodal retrieval with high accuracy is achieved.
Patent Information
- Application Number
- CN202510563679.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional single-modal retrieval methods cannot meet the multimodal retrieval requirements in complex environments, and cannot efficiently obtain complete description information of the scene between different media types.
By obtaining feature data sets, using large language models to generate question-answer text, extracting image and text feature vectors, training multimodal matching models, and realizing cross-modal retrieval.
High accuracy retrieval between different media types is achieved, and the complete description information of the scene can be obtained through image or text input.
Smart Images

Figure CN120508614A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent retrieval, and in particular to a scene retrieval method and terminal based on a large model. Background Art
[0002] In scenarios such as construction sites where complex environmental monitoring is required, it is often necessary to search for specific scenes and find the scene content that needs to be viewed. Traditional single-modal retrieval, such as image search, text search, or retrieval where the query and candidate set belong to the same modality, can no longer meet the retrieval needs in similar scenarios. Users hope to be able to search between modalities without being restricted by media type to obtain more complete description information about the scene. For example, they can find an image of a scene and the required content through a descriptive text. Therefore, cross-modal retrieval technology in the monitoring field can effectively utilize information between multiple modalities to assist users in retrieving the required content from massive amounts of monitoring data, which has important research significance and practical value. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a scene retrieval method and terminal based on a large model to achieve more accurate multi-modal retrieval.
[0004] In order to solve the above technical problems, a technical solution adopted by the present invention is:
[0005] A scene retrieval method based on a large model comprises the following steps:
[0006] Acquire a feature data set, wherein the feature data set includes a plurality of training samples, each of the training samples includes descriptive text data and a scene image corresponding to the descriptive text data;
[0007] Inputting the scene image into a large language model to obtain a plurality of question-answer texts corresponding to the scene image;
[0008] Extracting an image feature vector corresponding to the scene image and a text feature vector corresponding to the description text and the question-answer text;
[0009] Training a preset multimodal matching model based on the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model;
[0010] Inputting the target images in the database into the multimodal matching model one by one to obtain text information corresponding to each target image;
[0011] receiving a pending image and obtaining description information corresponding to the pending image;
[0012] The retrieval is performed based on the similarity between the description information and the text information.
[0013] In order to solve the above technical problems, another technical solution adopted by the present invention is:
[0014] A large model-based scene retrieval terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0015] Acquire a feature data set, wherein the feature data set includes a plurality of training samples, each of the training samples includes descriptive text data and a scene image corresponding to the descriptive text data;
[0016] Inputting the scene image into a large language model to obtain a plurality of question-answer texts corresponding to the scene image;
[0017] Extracting an image feature vector corresponding to the scene image and a text feature vector corresponding to the description text and the question-answer text;
[0018] Training a preset multimodal matching model based on the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model;
[0019] Inputting the target images in the database into the multimodal matching model one by one to obtain text information corresponding to each target image;
[0020] receiving a pending image and obtaining description information corresponding to the pending image;
[0021] The retrieval is performed based on the similarity between the description information and the text information.
[0022] The beneficial effects of the present invention are: first, a feature data set is obtained, and each training sample in the feature data set includes descriptive text data and a corresponding scene image; this is the final training target; then, the scene image is input into a large language model to obtain multiple question-answer texts, and an initial description of the scene image is obtained using the question-answering mechanism of the large language model; then, the scene image, descriptive text data and question-answer text are extracted as image feature vectors and text feature vectors respectively, and the image and text are unified into vector dimensions for easy comparison; finally, the image feature vector, text feature vector and descriptive text data are input into a preset multimodal matching model for training, so that the trained multimodal matching model can complete the process of obtaining the corresponding description based on the input image; then, according to the trained multimodal matching model, the text information corresponding to the target image already in the database can be extracted; then, if the user needs to match the image by inputting an image or retrieve the image by inputting text, matching can be achieved through the associated text features, thereby realizing cross-modal retrieval by unifying to the text dimension. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A flowchart of a scene retrieval method based on a large model according to an embodiment of the present invention;
[0024] Figure 2 Flowchart of another step of a scene retrieval method based on a large model according to an embodiment of the present invention;
[0025] Figure 3 Schematic diagram of a scene retrieval device based on a large model according to an embodiment of the present invention;
[0026] Figure 4 A schematic diagram of a scene retrieval terminal based on a large model according to an embodiment of the present invention;
[0027] Description of labels:
[0028] 1. A scene retrieval terminal based on a large model; 2. A processor; 3. A memory. DETAILED DESCRIPTION
[0029] To illustrate the technical content, achieved objectives and effects of the present invention in detail, the following description is given in conjunction with the embodiments and accompanying drawings.
[0030] Please refer to Figure 1 , a scene retrieval method based on a large model, comprising the steps of:
[0031] Acquire a feature data set, wherein the feature data set includes a plurality of training samples, each of the training samples includes descriptive text data and a scene image corresponding to the descriptive text data;
[0032] Inputting the scene image into a large language model to obtain a plurality of question-answer texts corresponding to the scene image;
[0033] Extracting an image feature vector corresponding to the scene image and a text feature vector corresponding to the description text and the question-answer text;
[0034] Training a preset multimodal matching model based on the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model;
[0035] Inputting the target images in the database into the multimodal matching model one by one to obtain text information corresponding to each target image;
[0036] receiving a pending image and obtaining description information corresponding to the pending image;
[0037] The retrieval is performed based on the similarity between the description information and the text information.
[0038] From the above description, it can be seen that the beneficial effects of the present invention are: first, a feature data set is obtained, and each training sample in the feature data set includes descriptive text data and a corresponding scene image; then, the scene image is input into a large language model to obtain multiple question-answer texts, and an initial description of the scene image is obtained using the question-answering mechanism of the large language model; then, the scene image, descriptive text data and question-answer text are extracted as image feature vectors and text feature vectors respectively, and the image and text are unified into vector dimensions for easy comparison; finally, the image feature vector, text feature vector and descriptive text data are input into a preset multimodal matching model for training, so that the trained multimodal matching model can complete the process of obtaining the corresponding description based on the input image; then, according to the trained multimodal matching model, the text information corresponding to the target image already in the database can be extracted; then, if the user needs to match the image by inputting an image or retrieve the image by inputting text, matching can be achieved through the associated text features, thereby realizing cross-modal retrieval by unifying to the text dimension.
[0039] Furthermore, inputting the scene image into a large language model to obtain a plurality of question-answer texts corresponding to the scene image includes:
[0040] Encoding the scene image into a first tag sequence by a recognition module in the large language model;
[0041] Associating the first token sequence with the current question and inputting the resultant tokens into the decoder of the large language model to obtain a current answer associated with the current question; combining the current question and the current answer to obtain context information;
[0042] Repeat the following steps until all questions are answered:
[0043] A new current question is added to the context information and then input into the decoder to obtain a current answer associated with the context information, and the current answer is added to the context information.
[0044] From the above description, it can be seen that the scene image is first encoded to obtain the first tag sequence, and the current question and the first tag sequence are input into the large language model together to obtain the current answer corresponding to the current question. Since the first tag sequence extracted from the image is also used as input, the large language model will output the current answer based on the current question and scene image, thereby realizing question and answer for the scene image. In the process of continuous questioning, the previous questions and answers are used as a reference for obtaining answers later. In the process of outputting answers, the large language model is based not only on the initial scene image, but also on the context generated in the subsequent question and answer process, thereby improving the depth and breadth of the mining of the description of the scene image, so that the final context information including multiple questions and answers can cover as comprehensive scene image content as possible.
[0045] Furthermore, the following steps are executed in a loop until all answers to the questions are obtained:
[0046] Inputting the context information into an artificial intelligence model to obtain a text summary;
[0047] A data set is generated according to the text summary, wherein the data set includes a plurality of question-answer texts, each of which includes an image, a question, and an answer corresponding to the question.
[0048] From the above description, it can be seen that after obtaining the context information through the large language model, the context information is extracted as a text summary through the artificial intelligence model. The artificial intelligence model can filter out the redundant parts in the context information and retain the characteristic parts while retaining the question-answer correspondence, thereby reducing the amount of information in the subsequent training process and improving the speed of subsequent model training.
[0049] Furthermore, the step of training a preset multimodal matching model based on the descriptive text data, the image feature vector, and the text feature vector to obtain the trained multimodal matching model includes:
[0050] Build image feature extraction models and text feature extraction models;
[0051] Constructing a feature enhancement module connected to the image feature extraction model and the text feature extraction model respectively;
[0052] constructing a projection network connected to the feature enhancement module;
[0053] A visual language model connected to the projection network and the text extraction feature model is constructed to obtain a preset multimodal matching model.
[0054] From the above description, it can be seen that the multimodal matching model obtained by the image feature extraction model, the text feature extraction model, the feature enhancement module, the projection network and the visual language model can extract image features, text features, enhance and fuse the extracted image features and text features, and embed the enhanced and fused features into the spatial dimension, that is, mapping different types of features into the same feature space, so as to facilitate the subsequent model to learn the correlation between different types of features.
[0055] Furthermore, the training of a preset multimodal matching model according to the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model includes:
[0056] Inputting the scene image into an image feature extraction model to obtain an image label;
[0057] Inputting the expression text data into a text feature extraction model to obtain a text feature vector;
[0058] fusing the image tag and the text feature vector by the feature enhancement module to obtain an image fusion code;
[0059] Mapping the image fusion code through the projection network to obtain a mapping sequence;
[0060] The mapping sequence is connected with the text feature vector and input into a visual language model, and the question in the question-answer text corresponding to the scene image is input. The output of the visual language model is compared with the answer in the question-answer text to train the preset multimodal matching model to obtain a trained multimodal matching model.
[0061] From the above description, it can be seen that the image tag is obtained through the image feature extraction model and the text feature vector is obtained through the question feature extraction model. The image tag and the text feature vector are fused, which can integrate the features on the scene image and the features of the corresponding text description, so that the subsequent output description is more comprehensive, thereby improving the accuracy of the retrieval, and the image fusion encoding after fusion through projection network mapping obtains a mapping sequence, so that the model can learn the association between different types of features; finally, the mapping sequence and the text feature vector are connected and input into the visual language model, and the question in the question-answer text is input, and the preset multimodal matching model is trained with the answer in the question-answer text as a reference. The question provides a reference feature for the text direction that can be generated, so that in the process of generating the description corresponding to the scene image, the features in the image can be fully mined for conversion.
[0062] Furthermore, the training of a preset multimodal matching model according to the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model includes:
[0063] Keeping the image feature extraction model and the visual language model unchanged, training the projection network;
[0064] After the projection network is trained, the image feature extraction model and the trained projection network are kept unchanged, and the visual language model is trained.
[0065] From the above description, it can be seen that since the multimodal matching model is composed of multiple models with different functions, if multiple models are trained at the same time, a large number of parameters need to be adjusted, and the adjustment of the parameters will affect the results of each other, making it difficult to locate the effect of the parameter adjustment. Therefore, when training the projection network, the image feature extraction model and the visual language model are kept unchanged. After the training is completed, the visual language model is trained again, and the image feature extraction model and the projection network are kept unchanged, thereby reducing the number of parameters that need to be processed during a training process, improving the efficiency of model iteration and improving the accuracy of the final trained model.
[0066] Furthermore, the training of the visual language model includes:
[0067] Obtaining a weight matrix in the visual language model;
[0068] Performing singular value decomposition on the weight matrix to obtain a low-rank matrix of m singular values, where m is the dimension of the weight matrix;
[0069] Select r main singular values in the low-rank matrix to form an initialization matrix for fine-tuning weights, and obtain k initialization matrices; r<m;
[0070] Constructing a first orthogonal matrix, a singular value matrix and a second orthogonal matrix according to the r main singular values;
[0071] Constructing a first low-rank matrix based on the first orthogonal matrix and the singular value matrix; constructing a second low-rank matrix based on the singular value matrix and the second orthogonal matrix;
[0072] During the process of training the chat model, the first low-rank matrix and the second low-rank matrix are iterated.
[0073] From the above description, it can be seen that in the process of training the visual language model, if the complete weight matrix is trained directly, it takes a long time to calculate the result each time. Here, the first low-rank matrix and the second low-rank matrix are obtained based on the weight matrix, that is, only the changed part of the weight matrix is calculated, and then the final result is deduced. During the calculation process, calculations are performed based on the changed part instead of directly calculating the complete weight matrix, thereby further reducing the amount of calculation in the iterative process and accelerating the iteration process of the model.
[0074] Furthermore, the visual language model is a large language model.
[0075] From the above description, we can see that using a large language model as a visual language model can fully understand natural language, thereby integrating features to obtain corresponding outputs. As the main model for extracting picture information, the extraction results obtained are comprehensive.
[0076] Furthermore, the receiving of the pending image and obtaining description information corresponding to the pending image includes:
[0077] Generate description information of the pending image features through machine model question answering.
[0078] From the above description, it can be seen that the description information of the undetermined image features is generated through machine model question answering, and the description information of the image to be retrieved is automatically extracted without manual image description.
[0079] Please refer to the figure, a scene retrieval terminal based on a large model includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step in the above-mentioned scene detection terminal based on a large model is implemented.
[0080] The above-mentioned large model-based scene detection method and terminal of the present invention are applicable to, and are described below through specific implementation methods.
[0081] Please refer to Figure 1 , embodiment 1 of the present invention is:
[0082] A scene detection method based on a large model comprises the following steps:
[0083] S1. Acquire a feature data set, wherein the feature data set includes a plurality of training samples, each of the training samples including descriptive text data and a scene image corresponding to the descriptive text data;
[0084] The description text data may describe scene information in the scene image or other information in the scene image, which may be determined according to the retrieval requirements;
[0085] S2. Inputting the scene image into a large language model to obtain multiple question-answer texts corresponding to the scene image, including:
[0086] S21, encoding the scene image into a first tag sequence through a recognition module in the large language model;
[0087] In some embodiments, a scene image is encoded into a first tag sequence f(X) by using ViT-base-16 (an image encoder) in a CLIP (Contrastive Language-Image Pre-training) model as a recognition module, where X represents the scene image and f() generally refers to the encoding process.
[0088] S22: Associating the first tag sequence with the current question Q1 and inputting the result into the decoder of the large language model to obtain a current answer A1 associated with the current question; combining the current question and the current answer to obtain context information;
[0089] In some embodiments, the Q1-QN questions include preset fixed questions, such as "Please describe the scene in the image," "What does the image mainly contain?", "What objects does the image mainly include?", etc. The preset fixed questions are generally common to different types of images. Extended questions can also be obtained through the model. Extended questions can be obtained through different machine models to increase randomness and thus enrich the training materials.
[0090] In some embodiments, the concatenated first tag sequence and Q1 are input into the autoregressive distil-GPT-2 decoder to predict the current answer A1 corresponding to Q1 in an autoregressive manner;
[0091] S23, looping through the following steps until all answers associated with the questions are obtained: adding a new current question to the context information and inputting the result into the decoder to obtain a current answer associated with the context information, and adding the current answer to the context information;
[0092] In some embodiments, if there are N questions in total, the context information is finally obtained [f(X), Q1, A1, Q2, ..., QN-1, AN-1, QN, AN];
[0093] Using a large language model, we can further explore and mine information in images through dialogue, and the questions and answers are interconnected, facilitating the subsequent training of a multimodal matching model.
[0094] In some embodiments, after S23, the method further includes:
[0095] S24, inputting the context information into an artificial intelligence model to obtain a text summary;
[0096] In some embodiments, context information is input into ChatGPT to obtain a text summary {X, U, T}, where X represents the scene image, which can be directly represented by the first tag sequence; U represents the question, and T represents the answer corresponding to question U; the question and answer here are the key content extracted from Q and A in the previous step, i.e., the summary;
[0097] S25. Generate a data set based on the text summary, wherein the data set includes multiple question-answer texts, each of which includes an image, a question, and an answer corresponding to the question;
[0098] S3, extracting the image feature vector corresponding to the scene image and the text feature vector corresponding to the description text data and the question-answer text;
[0099] S4. Training a preset multimodal matching model according to the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model;
[0100] In some embodiments, S4 includes:
[0101] S401, constructing an image feature extraction model and a text feature extraction model;
[0102] In some embodiments, the image feature extraction model is a pre-trained visual feature encoding network;
[0103] S402: Constructing feature enhancement modules connected to the image feature extraction model and the text feature extraction model respectively;
[0104] S403, constructing a projection network connected to the feature enhancement module;
[0105] The projection network connects the image feature extraction model to the text feature extraction model;
[0106] S404: constructing a visual language model connected to the projection network and the text extraction feature model to obtain a preset multimodal matching model;
[0107] The visual language model generates textual information by describing the content of the input image features. This can extract key information from the image to support scene category determination and retrieval. The visual language model takes an image as input and outputs a textual description.
[0108] Then S4 includes:
[0109] S41, inputting the scene image into an image feature extraction model to obtain an image label;
[0110] The visual feature encoding network starts from the input scene image X∈R H×W×C The visual mark is extracted from the scene image, where H represents the height of the scene image, W represents the width of the scene image, and C represents the number of channels of the scene image; the feature encoding network encodes the scene image into an image mark I X ∈R N×D ; N represents the length of the tag sequence, that is, the length of the image after encoding, and D represents the dimension of the image encoder of the image feature extraction model, that is, after encoding the image, a two-dimensional matrix of size N×D is obtained;
[0111] The R used in this embodiment represents the spatial dimension, for example, R H×W×C Indicates that the matrix dimension is a three-dimensional matrix of H×W×C, R N×D represents a two-dimensional matrix with a matrix dimension of N×D. Other Rs in the embodiments have the same meanings and are not repeated here.
[0112] S42, inputting the expression text data into a text feature extraction model to obtain a text feature vector;
[0113] For the input text data S = s1 + s2 + s3 + ... + sM; M represents the total length of the sentence, s represents the sentence segmentation, that is, the effective segmentation unit token of the sentence; use the pre-trained DistilBERT model to encode each token of the text (the unit after text segmentation) to obtain the encoding F I (S), length is S1, there are M in total, so the matrix is a matrix of size M×S1;
[0114] F I (S) = DistilBERT(S);
[0115] Finally, we get the encoding matrix F corresponding to the sentence I (S)=[h1,h2,h3,…,hM]∈R M×S1; The encoding matrix is the text feature vector; where M represents the total length of the sentence, i.e. the number of tokens, and S1 represents the encoding length of each token.
[0116] S43, fusing the image tag and the text feature vector through the feature enhancement module to obtain an image fusion code;
[0117] To enhance the connection between images and text, a feature enhancement module is added after the image feature extraction model and the text feature extraction model to fuse the cross-modal residual features between images and text.
[0118] In some embodiments, the image tags and text feature vectors are input into a multi-head self-attention network to capture global contextual information within the modality. In the self-attention mechanism, the input features and three trainable parameter matrices are multiplied to obtain query, key, and value matrices (Q, K, V), where the self-attention network formula is as follows:
[0119]
[0120] Where (Q, K, V) are the query, key, and value matrices corresponding to the image tags and text feature vectors, d k represents the dimension of the key vector K;
[0121] Q, K, and V are all linearly transformed from the same input matrix X. We can simply understand it as:
[0122] Q=XW Q
[0123] K=XW K
[0124] V=XW V
[0125] Where W Q 、W K 、W V There are three trainable parameter matrices. The Attention mechanism does not use them directly, but uses these three matrices generated by matrix multiplication. This is because the use of three trainable parameter matrices can enhance the fitting ability of the model. Here, X represents the input matrix, which refers to the image label and text feature vector, and Q I K I V I The image label is obtained by linear transformation of the input matrix X, Q T K T V T The text feature vector is transformed as the input matrix;
[0126] The features of the image and text output by the multi-head attention network are input into the cross attention network respectively, where the cross attention network formula is as follows
[0127]
[0128] The matrix of image features is represented by Q I , K I , V I , the matrix expression of text features is Q T , K T , V T ; F I-T and F T-I Represent the image features and text features of the feature enhancement part respectively;
[0129] The image features and text features are fused across modal feature residuals according to the fusion ratio to obtain the final fusion enhanced features. The fusion formula is as follows:
[0130] I enhance =F I +γ1F I-T ;
[0131] T enhance =F T +γ2F T-I ;
[0132] Among them F I and F T The last layer representing image features and text features, F I-T and T T-I Represent the image features and text features of the feature enhancement part respectively. γ1 and γ2 represent the parameters that control the residual fusion ratio, which are initialized to very small numbers. enhance and T enhance Represents the image and text features after fusion of residual features for enhancement;
[0133] S44, mapping the image fusion code through the projection network to obtain a mapping sequence;
[0134] The projection network consists of two layers of GELU activations, which embed the visual tag map into the spatial dimension S to form a mapping sequence Fx. The mapped image feature is Fx∈R N×S2 ; That is, mapping image features to the space where text features are located; where N represents the length of the tag sequence in the image encoding, and S2 represents the dimension of the mapped space;
[0135] S45. Concatenating the mapping sequence with the text feature vector and inputting the resultant data into a visual language model, and inputting the question in the question-answer text corresponding to the scene image, comparing the output of the visual language model with the answer in the question-answer text to train the preset multimodal matching model, thereby obtaining a trained multimodal matching model; using the description expanded by the question-answer text as training content, thereby enriching and broadening the description content of the image, increasing the text features of the image, and thus improving the training effect;
[0136] In some embodiments, the mapped image features Fx∈R N×S and text feature vector F I ∈R M×S Connect to form F LLM ∈R K×S Input, where K = M + N;
[0137] In some embodiments, the visual language model is an LLM (Large Language Model), which is a visual language model based on the transformer architecture. The model takes a sequence of visual and language tokens as input and begins to generate responses in an autoregressive manner; given an image instruction token (token1), it can maximize the probability distribution of generating the correct response. This probability distribution can be expressed as follows:
[0138]
[0139] Where K represents the length of the response sequence, which is used here to explain the principle of forming a large language model. It is independent of the feature length mentioned above. Here we explain how to maximize the probability distribution of generating the correct response (answer); P(T k |T 1......(k-1) ,U,X) represents the probability of getting the kth answer when the answer, question and scene image before the kth answer are determined; token1 here is just a general description and has no connection with the tokens in the previous text. It is only used to illustrate the probability distribution of the model generating a response under given data such as scene image and question instructions; this describes the principle of generating responses (answers) by the large language model;
[0140] In some embodiments, S45 includes:
[0141] S451, keeping the image feature extraction model and the visual language model unchanged, and training the projection network;
[0142] In some embodiments, the projection network is trained using a general image-language dataset of text-image pairs; the text in the text-image pairs includes descriptive text data and question-answer text, i.e., a text dataset expanded by a large language model;
[0143] S452: After the projection network is trained, the image feature extraction model and the trained projection network are kept unchanged, and the visual language model is trained;
[0144] Since the image feature extraction model and the text feature extraction model in the multimodal matching model are less coupled with the subsequent modules, the feature extraction effect can be achieved through separate training. Therefore, they are trained separately and the parameters of the feature extraction model are not adjusted during the iteration process of the entire multimodal matching model, thereby reducing the number of parameters that need to be iterated and improving the training speed. Since both the projection network and the visual language model will directly affect the final output results, the projection network is trained first and then the visual language model is trained in this order, which reduces the parameters that need to be adjusted in a single training process and avoids the problem that the parameters in the two modules cannot be adjusted together to locate the specific module corresponding to the changes. This speeds up the model iteration process and improves the accuracy of the final model.
[0145] Due to the large number of parameters and high computational cost in the visual language model, we use the improved LoRA (LoRA constrains the update amount to a low-rank matrix to reduce the number of parameters during training) to fine-tune the visual language model;
[0146] In some embodiments, S452 includes:
[0147] S4521, obtain the weight matrix W0∈R in the visual language model m×n , select 1 / 10 of the original matrix dimension as the low rank r, and when updating the parameters, use the low rank decomposition representation of W0+ΔW=W0+BA, B∈R m×r , A∈R r×n ;
[0148] S4522. Perform singular value decomposition on the weight matrix to obtain a low-rank matrix of m singular values, where m is the dimension of the weight matrix;
[0149] In some embodiments, the decomposition formula is as follows:
[0150] SVD(W0)=UΣV T ;
[0151] In the above formula, SVD() represents the decomposition operation; U represents an m×m orthogonal matrix, also known as the left singular vector; Σ represents an m×n diagonal matrix, that is, only the elements on the diagonal are singular values, and the other elements are 0, that is, a singular value matrix composed of singular values; V Trepresents the transpose of an n×n orthogonal matrix, also known as the right singular vector; let L = [L1, L2, L3…, Lm] represent the singular value matrix of the weight matrix W0; usually the singular values and singular vectors in Σ are in a one-to-one correspondence;
[0152] S4523, selecting r main singular values in the low-rank matrix to form an initialization matrix for fine-tuning weights, to obtain k initialization matrices; r<m;
[0153] Select r principal singular values (submatrices of Σ) and singular vectors (U and V T submatrix) as the initialization matrix for fine-tuning weights; the larger singular values in the front of the r singular values correspond to the main part, and the smaller singular values in the back correspond to the details. Each singular value and singular vector are one-to-one corresponding, just like the eigenvalues and eigenvectors of the matrix; just like PCA decomposition, the larger ones in the front are the principal components, and the larger singular values and singular vectors also correspond to the main part of the matrix; if only a part of them is selected, it may cause poor model calibration and poor accuracy; here, multiple groups of singular values are selected to form k initialization matrices, which are fine-tuned separately to obtain matrices of k fine-tuning parts, where the number of selected groups k < (N / r), where N represents the total number of singular values; for example, the pre-training weight dimension is 1000, the low rank r of the low rank matrix is 100, and the SVD decomposition singular values are generally much larger than 100. Multiple groups are selected proportionally, with the first 100 singular values being the first group, the 100th to 200th singular values being the second group, and so on to obtain k initialization matrices;
[0154] Because large singular values represent the main structural information of the data, and smaller singular values represent different degrees of detail information, if only the first r singular values are retained as the low-rank matrix for initialization, the influence of other weights in the pre-trained weights will be directly discarded. Therefore, the LoRA integration method is proposed here to select singular values of different parts and combine multiple initialization methods to further ensure the accuracy of model prediction and reduce the impact of overconfidence;
[0155] S4524. Construct a first orthogonal matrix, a singular value matrix, and a second orthogonal matrix according to the r main singular values;
[0156] That is, according to the selected r main singular values, the submatrix Ur of U is obtained as the first orthogonal matrix; according to the selected r main singular values, the singular value matrix Lr is obtained, and Lr is a diagonal matrix with r singular values as diagonals; according to the selected r main singular values, V is obtained. T The sub-matrix is used as the second orthogonal matrix Vr T ;
[0157] S4525. Construct a first low-rank matrix A based on the first orthogonal matrix and the singular value matrix; construct a second low-rank matrix B based on the singular value matrix and the second orthogonal matrix;
[0158] A=Ur×Lr 1 / 2 ∈R m×r ;
[0159] B=Lr 1 / 2 ×Vr T ∈R m×r ;
[0160] Where, Lr 1 / 2 Indicates the square root of the Lr matrix;
[0161] S4526. During the process of training the chat model, iterate the first low-rank matrix and the second low-rank matrix;
[0162] In some embodiments, the visual language model is a large language model;
[0163] Because W0 is a pre-trained weight, it generally has many parameters, possibly reaching tens of billions; so you need to use your own dataset for fine-tuning. Generally, you freeze W0, then fine-tune it (that is, focus only on the change term ΔW), and then use W0 + ΔW to get the updated weight parameters. ΔW is represented by the product of the first low-rank matrix A and the second low-rank matrix B, which has fewer parameters and is easier to train.
[0164] During training, W0 is frozen and does not receive gradient updates, while A and B contain fine-tuned weights, representing the differences to be added to the original weights of the visual language model. Finally, during inference, the fine-tuned weights are combined with the original pre-trained weights to obtain W0+ΔW. For input x, the modified forward pass can be expressed as y=W0×x+ΔW×x=W0×x+BAx. For each of the k ΔW matrices obtained by fine-tuning different initialization matrices, forward inference is performed separately to obtain k results y, which are then integrated to obtain the final result y';
[0165]
[0166] Where y' represents the expected output of the visual language model, α i represents the weight of the i-th group of low-rank adapters, that is, ΔW trained by initializing the i-th group of singular value matrices, and represents the change part of fine-tuning; represents the sum of r singular values corresponding to the i-th group of adapters, represents the sum of all singular values of weight W0;
[0167] Forward reasoning is to get the weight W0+ΔW after training yourself, and then use the input to get the output y=W0×x+ΔW×x;
[0168] S5, inputting the target images in the database into the multimodal matching model one by one to obtain text information corresponding to each target image;
[0169] During the formal process of matching text information to the target image, there is no need to input questions again. This is because the scene description has been enriched by inputting questions during the model training phase. Multimodal training is to train images and scene descriptions, and the resulting model can then perform forward reasoning, providing a better feature description of the input image. Therefore, only the target image needs to be input during forward reasoning.
[0170] S6. Receive a pending image and obtain description information corresponding to the pending image;
[0171] In some embodiments, the description information of the pending image is generated by machine model question answering; for example, the multimodal matching model trained in step S4 may be used, or other machine models capable of reading image features, such as the openly available ChatGPT, etc. may be used;
[0172] In some embodiments, the description information input by the operator through the operation terminal can be received by sending the pending image to the operation terminal; thus, the accuracy of the description information of the pending image used in the retrieval can be improved;
[0173] In some embodiments, the user may also directly input description information to retrieve images associated with the description information in the database;
[0174] S7. Complete the search based on the similarity between the description information and the text information, that is, search the database for text information whose similarity with the description information exceeds a threshold, and output the text information and the target image corresponding to the text information to complete the search.
[0175] Please refer to Figure 2 , the second embodiment of the present invention is:
[0176] A large model-based scene detection terminal 1 includes a processor 2, a memory 3, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, each step in the first embodiment is implemented.
[0177] In summary, the present invention provides a scene retrieval method and terminal based on a large model. This solution addresses the limitations of traditional single-modal scene retrieval and traditional visual algorithms that require manual feature design, and proposes a scene retrieval device based on multi-modal large model fine-tuning. It can not only use the text description of the scene content for intelligent retrieval to obtain the scene data of the required content, but also use context information combined with machine question answering to mine image information, which can fully extract the feature representation of the image; the constructed visual language model can fully extract and effectively train the instructions and responses of the image, wherein the feature fusion enhancement module can perform local enhancement of the features themselves and cross-modal joint enhancement; the integrated fine-tuning using the LoRA method with an improved initialization method can not only converge quickly, but also obtain better results for its specific application, reduce the prediction uncertainty of the model, and improve the calibration effect. The model adopts a large model based on deep learning. The above technical solution can process data of multiple modalities and adapt to different scenes and target objects. With the addition of new labeled data and feature information, the model can continuously optimize and improve its recognition performance through continuous training. Specifically, the first part of this disclosure is to use a machine question-answering system to supplement and enrich the features of the annotated text; the second part uses the supplemented text and the original image to fine-tune the large model (multimodal matching); a model that can describe the scene image in detail is obtained, and finally this model is used for reasoning, that is, scene images are collected and text is generated using the model; because of the fine-tuning and text supplementation, the generated text description effect can be improved.
[0178] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A scene retrieval method based on a large model, characterized in that: Including steps: Acquire a feature data set, wherein the feature data set includes a plurality of training samples, each of the training samples includes descriptive text data and a scene image corresponding to the descriptive text data; Inputting the scene image into a large language model to obtain a plurality of question-answer texts corresponding to the scene image; Extracting an image feature vector corresponding to the scene image and a text feature vector corresponding to the description text and the question-answer text; Training a preset multimodal matching model based on the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model; Inputting the target images in the database into the multimodal matching model one by one to obtain text information corresponding to each target image; receiving a pending image and obtaining description information corresponding to the pending image; The retrieval is performed based on the similarity between the description information and the text information.
2. A scene retrieval method based on a large model according to claim 1, characterized in that: Inputting the scene image into a large language model to obtain a plurality of question-answer texts corresponding to the scene image includes: Encoding the scene image into a first tag sequence by a recognition module in the large language model; Associating the first token sequence with the current question and inputting the resultant tokens into the decoder of the large language model to obtain a current answer associated with the current question; combining the current question and the current answer to obtain context information; Repeat the following steps until all questions are answered: A new current question is added to the context information and then input into the decoder to obtain a current answer associated with the context information, and the current answer is added to the context information.
3. A scene retrieval method based on a large model according to claim 2, characterized in that: The loop executes the following steps until all the answers associated with the questions are obtained: Inputting the context information into an artificial intelligence model to obtain a text summary; A data set is generated according to the text summary, wherein the data set includes a plurality of question-answer texts, each of which includes an image, a question, and an answer corresponding to the question.
4. A scene retrieval method based on a large model according to claim 1, characterized in that: The method includes: training a preset multimodal matching model according to the description text data, the image feature vector, and the text feature vector to obtain the trained multimodal matching model; Build image feature extraction models and text feature extraction models; Constructing a feature enhancement module connected to the image feature extraction model and the text feature extraction model respectively; constructing a projection network connected to the feature enhancement module; A visual language model connected to the projection network and the text feature extraction model is constructed to obtain a preset multimodal matching model.
5. A scene retrieval method based on a large model according to claim 4, characterized in that: The step of training a preset multimodal matching model according to the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model includes: Inputting the scene image into an image feature extraction model to obtain an image label; Inputting the description text data into a text feature extraction model to obtain a text feature vector; fusing the image tag and the text feature vector by the feature enhancement module to obtain an image fusion code; Mapping the image fusion code through the projection network to obtain a mapping sequence; The mapping sequence is connected with the text feature vector and input into a visual language model, and the question in the question-answer text corresponding to the scene image is input. The output of the visual language model is compared with the answer in the question-answer text to train the preset multimodal matching model to obtain a trained multimodal matching model.
6. A scene retrieval method based on a large model according to claim 4, characterized in that: The step of training a preset multimodal matching model according to the descriptive text data, the image feature vector, and the text feature vector to obtain a trained multimodal matching model includes: Keeping the image feature extraction model and the visual language model unchanged, training the projection network; After the projection network is trained, the image feature extraction model and the trained projection network are kept unchanged, and the visual language model is trained.
7. A scene retrieval method based on a large model according to claim 6, characterized in that: The training of the visual language model comprises: Obtaining a weight matrix in the visual language model; Performing singular value decomposition on the weight matrix to obtain a low-rank matrix of m singular values, where m is the dimension of the weight matrix; Select r main singular values in the low-rank matrix to form an initialization matrix for fine-tuning weights, and obtain k initialization matrices; r<m; Constructing a first orthogonal matrix, a singular value matrix and a second orthogonal matrix according to the r main singular values; Constructing a first low-rank matrix based on the first orthogonal matrix and the singular value matrix; constructing a second low-rank matrix based on the singular value matrix and the second orthogonal matrix; During the process of training the visual language model, the first low-rank matrix and the second low-rank matrix are iterated.
8. A scene retrieval method based on a large model according to claim 4, characterized in that: The visual language model is an LLM model.
9. The scene retrieval method based on a large model according to claim 1, characterized in that: The receiving the pending image and obtaining description information corresponding to the pending image includes: Generate description information of the undetermined image features through the large model.
10. A scene retrieval terminal based on a large model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the processor implements the steps of the large model-based scene retrieval method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Question and answer method and question and answer model training method
CN116561270A
Model training method and device, electronic equipment and storage medium
CN117521771A
Question-answer processing method and apparatus
WO2024220031A1
Cited By
Training sample construction method and device, electronic equipment and readable storage medium
CN120780819A
Virtual environment data processing method, electronic equipment and storage medium
CN120848743A
Virtual environment data processing method, electronic device, and storage medium
CN120848743B