Image search method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202510899761.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-06-30
AI Technical Summary
[0003]本发明实施例提供一种图像搜索方法,旨在解决现有图像搜索方法存在对于某个目标进行搜索时,需要了解目标在数据库中对应的检索要素,才能准确地在数据库中查询到对应的图像,增加了图像数据的查询难度,导致图像搜索效率低的问题
[0042]本发明实施例中,获取用户的输入数据;对输入数据进行语义提取处理,得到输入数据对应的全局语义向量和至少一个目标语义向量;基于全局语义向量和至少一个目标语义向量,在图像数据库中进行语义搜索,得到输入数据对应的搜索结果;将搜索结果进行返回。本发明通过对用户的输入数据进行语义提取处理,得到输入数据对应的全局语义向量和至少一个目标语义向量,根据全局语义向量和至少一个目标语义向量,在图像数据库中进行语义搜索,得到输入数据对应的搜索结果,并将搜索结果进行返回,解决了现有图像搜索方法存在对于某个目标进行搜索时,需要了解目标在数据库中对应的检索要素,才能准确地在数据库中查询到对应的图像,增加了图像数据的查询难度,导致图像搜索效率低的问题。
Smart Images

Figure CN120950716B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image search technology, and in particular to an image search method, apparatus, electronic device, and storage medium. Background Technology
[0002] When searching image data, existing methods can retrieve matching images based on user-input criteria and database query rules. However, this requires users to understand and learn these rules. To accurately find the corresponding image in the database, users need to know the relevant search elements for that target, increasing the difficulty of image data retrieval. Therefore, there is an urgent need for a more accurate image search method to address the problem of low efficiency caused by the requirement to understand the relevant search elements in the database for accurate image retrieval. Summary of the Invention
[0003] This invention provides an image search method aimed at solving the problem that existing image search methods require knowledge of the corresponding retrieval elements in a database to accurately find the image, increasing the difficulty of image data retrieval and resulting in low image search efficiency. This invention performs semantic extraction processing on user input data to obtain a global semantic vector and at least one target semantic vector. Based on the global semantic vector and at least one target semantic vector, a semantic search is performed in the image database to obtain the search results corresponding to the input data, and the search results are returned. This solves the problem that existing image search methods require knowledge of the corresponding retrieval elements in a database to accurately find the image, increasing the difficulty of image data retrieval and resulting in low image search efficiency.
[0004] In a first aspect, embodiments of the present invention provide an image search method, the method comprising the following steps:
[0005] Obtain user input data;
[0006] The input data is subjected to semantic extraction processing to obtain a global semantic vector and at least one target semantic vector corresponding to the input data;
[0007] Based on the global semantic vector and at least one of the target semantic vectors, a semantic search is performed in the image database to obtain the search results corresponding to the input data;
[0008] Return the search results.
[0009] Optionally, the step of performing semantic extraction processing on the input data to obtain a global semantic vector corresponding to the input data and at least one target semantic vector includes:
[0010] The input data is input into a trained multimodal large model, and the trained multimodal large model outputs multiple candidate global semantic vectors and multiple candidate target semantic vectors.
[0011] Clustering is performed on multiple candidate target semantic vectors to obtain target semantic vector clusters;
[0012] Based on the distance between each candidate global semantic vector and each target semantic vector cluster, the global semantic vector corresponding to the input data and at least one target semantic vector are selected.
[0013] Optionally, selecting the global semantic vector corresponding to the input data and at least one target semantic vector based on the distance between each candidate global semantic vector and each target semantic vector cluster includes:
[0014] Calculate the center vector of each of the target semantic vector clusters, and calculate the compactness of each of the target semantic vector clusters;
[0015] Calculate the Euclidean distance between each candidate global semantic vector and each of the center vectors to obtain the distance value between each candidate global semantic vector and each of the target semantic vector clusters;
[0016] Based on the distance value and the density, the global semantic vector and at least one target semantic vector corresponding to the input data are determined.
[0017] Optionally, the step of performing a semantic search in the image database based on the global semantic vector and at least one of the target semantic vectors to obtain the search results corresponding to the input data includes:
[0018] Calculate the first semantic similarity between the global semantic vector and the image semantic vectors of each image data in the image database;
[0019] And, calculate the second semantic similarity between the image semantic vectors of each image data for each target semantic vector;
[0020] Based on the first semantic similarity and the second semantic similarity, target image data is selected from the image database;
[0021] Based on the target image data, the search results corresponding to the input data are determined.
[0022] Optionally, selecting target image data from the image database as the search result corresponding to the input data based on the first semantic similarity and the second semantic similarity includes:
[0023] Select target image data from the image database whose first semantic similarity is greater than a first threshold and whose second semantic similarity is greater than a second threshold.
[0024] Alternatively, for each image data, calculate the comprehensive similarity with the second semantic similarity corresponding to all the target semantic vectors;
[0025] Target image data with a first semantic similarity greater than a first threshold and a comprehensive similarity greater than a third threshold are selected from the image database.
[0026] Optionally, before inputting the input data into the trained multimodal large model and outputting multiple candidate global semantic vectors and multiple candidate target semantic vectors through the trained multimodal large model, the method further includes:
[0027] Acquire a training dataset and a pre-trained multimodal large model. The training dataset contains multimodal sample data, global semantic category labeling data, and target semantic category labeling data corresponding to the sample data. The output of the pre-trained multimodal large model is a global semantic prediction vector and the global semantic category data corresponding to the global semantic prediction vector, and a target semantic prediction vector and the target semantic category data corresponding to the target semantic prediction vector.
[0028] The pre-trained multimodal large model is fine-tuned based on the training dataset. After training, the trained multimodal large model is obtained.
[0029] Optionally, the step of fine-tuning the pre-trained multimodal large model based on the training dataset, and obtaining a trained multimodal large model after training, includes:
[0030] The sample data is input into the pre-trained multimodal large model, and the pre-trained multimodal large model outputs global semantic category data and target semantic category data.
[0031] The first error loss between the category data of the global semantics and the category labeling data of the global semantics is calculated using a first loss function;
[0032] And, a second error loss is calculated between the category data of the target semantics and the category annotation data of the target semantics using a second loss function;
[0033] The total error loss is calculated by weighting the first error loss and the second error loss.
[0034] With minimizing the total error loss as the optimization objective, the model parameters of the pre-trained multimodal large model are adjusted through the backpropagation algorithm. The adjustment process of the model parameters is iterated until the total error loss is less than a preset value or the number of iterations reaches a preset number, at which point training stops and a trained multimodal large model is obtained.
[0035] Secondly, embodiments of the present invention provide an image search device, the image search device comprising:
[0036] The acquisition module is used to acquire user input data;
[0037] The processing module is used to perform semantic extraction processing on the input data to obtain a global semantic vector and at least one target semantic vector corresponding to the input data;
[0038] The search module is used to perform a semantic search in the image database based on the global semantic vector and at least one of the target semantic vectors to obtain the search results corresponding to the input data;
[0039] The return module is used to return the search results.
[0040] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the image search method provided in embodiments of the present invention.
[0041] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the image search method provided in the embodiments of the present invention.
[0042] In this embodiment of the invention, user input data is acquired; semantic extraction processing is performed on the input data to obtain a global semantic vector and at least one target semantic vector corresponding to the input data; based on the global semantic vector and at least one target semantic vector, a semantic search is performed in an image database to obtain the search results corresponding to the input data; and the search results are returned. This invention solves the problem in existing image search methods where, when searching for a target, it is necessary to know the corresponding retrieval elements in the database to accurately find the corresponding image, increasing the difficulty of image data retrieval and leading to low image search efficiency. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart of an image search method provided in an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram of the structure of an image search device provided in an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] like Figure 1 As shown, Figure 1 This is a flowchart of an image search method provided by an embodiment of the present invention. The image search method includes the following steps:
[0049] 101. Obtain user input data.
[0050] In this embodiment of the invention, the image search method described above can be applied to an image search platform, which can be built on a server-based or distributed platform. The image search platform includes a data interface (for sensor or user uploads), a knowledge database, and a knowledge database construction program. The data interface can be used to acquire user input data, and the knowledge database construction program can be used to construct the knowledge database. The knowledge database is specifically designed to provide additional relational information for identified data entities, thereby improving the depth of the data recognition system's understanding of the content.
[0051] The users mentioned above can be understood as people who use the image search platform, including administrators.
[0052] The above input data can be used to understand the text content entered by a user during the image search process.
[0053] 102. Perform semantic extraction processing on the input data to obtain the global semantic vector and at least one target semantic vector corresponding to the input data.
[0054] In this embodiment of the invention, the above semantic extraction process can be understood as a process of extracting semantic features from the input data by analyzing the input data.
[0055] The semantic vectors mentioned above are used to represent the semantic features in the input data.
[0056] The aforementioned global semantic vector can be understood as the key features extracted from the input data, representing the theme and content of the entire input data. The global semantic vector corresponding to the input data is used to capture the complete meaning expressed by the input data; the global semantic vector includes all the information and concepts contained in the input data.
[0057] At least one of the above can be understood as the image search containing at least one target semantic vector that matches the input data.
[0058] The target semantic vector mentioned above can be understood as the semantic vector corresponding to the input data.
[0059] It should be noted that semantic extraction of input data can be performed using natural language processing (NLP) techniques to obtain semantic features of the input data. Based on these semantic features, a global semantic vector and at least one target semantic vector corresponding to the input data can be obtained. The aforementioned NLP techniques aim to enable machines to understand, interpret, and generate human language, achieving effective communication between humans and machines, and enabling computers to perform tasks such as language translation, sentiment analysis, and text summarization.
[0060] 103. Based on the global semantic vector and at least one target semantic vector, perform a semantic search in the image database to obtain the search results corresponding to the input data.
[0061] In this embodiment of the invention, the image database described above is a database used for storing, retrieving, and managing image data.
[0062] The semantic search described above can be understood as a process of performing a semantic search in an image database based on the global semantic vector corresponding to the input data and at least one target semantic vector.
[0063] The search results mentioned above can be obtained by performing a semantic search in an image database based on the global semantic vector corresponding to the input data and at least one target semantic vector, resulting in search results corresponding to the input data.
[0064] Specifically, based on the global semantic vector corresponding to the input data and at least one target semantic vector, a semantic search can be performed in the image database to obtain the search results corresponding to the input data.
[0065] 104. Return the search results.
[0066] In this embodiment of the invention, the search results mentioned above are the search results corresponding to the input data.
[0067] The above-mentioned return can be understood as, in image search, the list of content related to the user's input data displayed by the system. Specifically, search results can be returned to the user in list format.
[0068] In this embodiment of the invention, user input data is acquired; semantic extraction processing is performed on the input data to obtain a global semantic vector and at least one target semantic vector corresponding to the input data; based on the global semantic vector and at least one target semantic vector, a semantic search is performed in an image database to obtain the search results corresponding to the input data; and the search results are returned. This invention solves the problem in existing image search methods where, when searching for a target, it is necessary to know the corresponding retrieval elements in the database to accurately find the corresponding image, increasing the difficulty of image data retrieval and leading to low image search efficiency.
[0069] It is understood that in the specific implementation of this application, data such as input data, image data, semantic data, knowledge data, and user data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required. Furthermore, the collection, use, and processing of related data, as well as the training, deployment, and invocation of algorithm models, must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0070] Optionally, in the step of semantic extraction processing of the input data to obtain the global semantic vector corresponding to the input data and at least one target semantic vector, the input data can be input into a trained multimodal large model, and the trained multimodal large model can output multiple candidate global semantic vectors and multiple candidate target semantic vectors; the multiple candidate target semantic vectors can be clustered to obtain target semantic vector clusters; based on the distance between each candidate global semantic vector and each target semantic vector cluster, the global semantic vector corresponding to the input data and at least one target semantic vector can be selected.
[0071] In this embodiment of the invention, the trained multimodal large model can be a multimodal large model built based on deep learning or machine learning, such as CLIP, LLM, etc. This multimodal large model is capable of recognizing the semantic features of the input data.
[0072] The semantic vectors mentioned above are used to represent the semantic features in the input data.
[0073] The global semantic vector mentioned above can represent the theme and content of the entire input data. The global semantic vector is used to capture the complete meaning expressed by the input data.
[0074] The target semantic vector mentioned above can be understood as the semantic vector corresponding to the input data.
[0075] The aforementioned candidate global semantic vectors can be understood as alternative global semantic vectors.
[0076] The aforementioned candidate target semantic vectors can be understood as alternative target semantic vectors.
[0077] The clustering process described above can be understood as classifying multiple candidate target semantic vectors and grouping similar data into one category. The K-means algorithm or other clustering algorithms can be used to cluster multiple candidate target semantic vectors.
[0078] The aforementioned target semantic vector cluster can be understood as a set of target semantic vectors.
[0079] At least one of the above can be understood as the image search containing at least one target semantic vector that matches the input data.
[0080] The distance between each candidate global semantic vector and each target semantic vector cluster can be understood as the Euclidean distance between them. The Euclidean distance is the true distance between two points in Euclidean space. A smaller Euclidean distance indicates a smaller distance between each candidate global semantic vector and each target semantic vector cluster, signifying a smaller difference and higher similarity. Conversely, a larger Euclidean distance indicates a larger distance between each candidate global semantic vector and each target semantic vector cluster, signifying a greater difference and lower similarity.
[0081] It should be noted that the global semantic vector corresponding to the input data and at least one target semantic vector can be selected based on the distance between each candidate global semantic vector and each target semantic vector cluster.
[0082] Optionally, in the step of selecting the global semantic vector and at least one target semantic vector corresponding to the input data based on the distance between each candidate global semantic vector and each target semantic vector cluster, the center vector of each target semantic vector cluster can be calculated, as well as the compactness of each target semantic vector cluster can be calculated; the Euclidean distance between each candidate global semantic vector and each center vector can be calculated to obtain the distance value between each candidate global semantic vector and each target semantic vector cluster; based on the distance value and the compactness, the global semantic vector and at least one target semantic vector corresponding to the input data can be determined.
[0083] In this embodiment of the invention, the aforementioned center vector can be understood as the average value of all vectors in each target semantic vector cluster, used to represent the representative features of each target semantic vector cluster. Specifically, the center vector can be obtained by calculating the average value of all vectors in each target semantic vector cluster.
[0084] The aforementioned density can be understood as the degree of similarity between all vectors within each target semantic vector cluster, which can be measured by calculating the distance or angle between vectors.
[0085] The Euclidean distance mentioned above can be understood as the straight-line distance between each candidate global semantic vector and each center vector in space. The similarity between each candidate global semantic vector and each center vector can be measured by calculating the Euclidean distance between vectors.
[0086] The above distance values can be understood as the Euclidean distance between each candidate global semantic vector and each center vector.
[0087] It should be noted that the smaller the Euclidean distance value, the smaller the distance between each candidate global semantic vector and each center vector, which means that the difference between each candidate global semantic vector and each center vector is smaller and the similarity is higher; conversely, the larger the Euclidean distance value, the larger the distance between each candidate global semantic vector and each center vector, which means that the difference between each candidate global semantic vector and each center vector is greater and the similarity is lower.
[0088] The global semantic vector mentioned above can represent the theme and content of the entire input data. The global semantic vector is used to capture the complete meaning expressed by the input data.
[0089] The target semantic vector mentioned above can be understood as the semantic vector corresponding to the input data.
[0090] Specifically, the center vector of each target semantic vector cluster can be calculated, as well as the compactness of each target semantic vector cluster. For each candidate global semantic vector, the Euclidean distance between each candidate global semantic vector and each target semantic vector cluster needs to be calculated. Based on the Euclidean distance, the distance value between each target semantic vector cluster and each target semantic vector cluster is obtained. Based on the Euclidean distance and compactness, the global semantic vector and at least one target semantic vector corresponding to the input data are determined.
[0091] Optionally, in the step of performing a semantic search in an image database based on a global semantic vector and at least one target semantic vector to obtain the search results corresponding to the input data, a first semantic similarity between the global semantic vector and the image semantic vectors of each image data in the image database can be calculated; and a second semantic similarity between the image semantic vectors of each image data for each target semantic vector can be calculated; target image data is selected in the image database based on the first semantic similarity and the second semantic similarity; and the search results corresponding to the input data are determined based on the target image data.
[0092] In this embodiment of the invention, the aforementioned first semantic similarity can be understood as the semantic similarity between the global semantic vector and the image semantic vectors of each image data in the image database. The degree of similarity between the global semantic vector and the image semantic vectors of each image data in the image database can be evaluated by calculating the cosine similarity between the global semantic vector and the image semantic vectors of each image data in the image database.
[0093] The aforementioned second semantic similarity can be understood as the semantic similarity between the image semantic vectors of each image data for each target semantic vector. The degree of similarity between the image semantic vectors of each image data for each target semantic vector can be evaluated by calculating the cosine similarity between them.
[0094] The cosine similarity mentioned above is an indicator used to measure the degree of similarity between the image semantic vectors of each target semantic vector in each image data. It can be judged by the cosine value of the angle between the image semantic vectors of each target semantic vector in each image data. The smaller the angle between the image semantic vectors of each target semantic vector in each image data, the greater the similarity. Conversely, the larger the angle between the image semantic vectors of each target semantic vector in each image data, the smaller the similarity.
[0095] The aforementioned image database is used for storing, retrieving, and managing image data.
[0096] The target image data mentioned above is selected from the image database based on the first semantic similarity and the second semantic similarity.
[0097] It should be noted that the search results corresponding to the input data can be determined based on the target image data.
[0098] Optionally, in the step of selecting target image data as the search result corresponding to the input data from the image database based on the first semantic similarity and the second semantic similarity, target image data with a first semantic similarity greater than a first threshold and a second semantic similarity greater than the second threshold can be selected from the image database; or, for each image data, the comprehensive similarity with the second semantic similarity corresponding to all target semantic vectors can be calculated; and target image data with a first semantic similarity greater than the first threshold and a comprehensive similarity greater than the third threshold can be selected from the image database.
[0099] In this embodiment of the invention, the first threshold is a semantic similarity threshold between the global semantic vector pre-set by the system and the image semantic vectors of each image data in the image database.
[0100] The second threshold mentioned above is a semantic similarity threshold between the image semantic vectors of each target semantic vector and each image data set, which is preset by the system.
[0101] The target image data mentioned above is the number of images selected from the image database based on the first semantic similarity and the second semantic similarity.
[0102] The aforementioned third threshold is a pre-set comprehensive similarity threshold that corresponds to the second semantic similarity of all target semantic vectors.
[0103] Specifically, target image data with a first semantic similarity greater than a first threshold and a second semantic similarity greater than a second threshold can be selected from the image database. Alternatively, for each image data, the comprehensive similarity with the second semantic similarity corresponding to all target semantic vectors can be calculated, and target image data with a first semantic similarity greater than a first threshold and a comprehensive similarity greater than a third threshold can be selected from the image database.
[0104] Optionally, before inputting the input data into the trained multimodal large model and outputting multiple candidate global semantic vectors and multiple candidate target semantic vectors through the trained multimodal large model, a training dataset and a pre-trained multimodal large model can be obtained; the pre-trained multimodal large model can be fine-tuned based on the training dataset, and after training, a trained multimodal large model can be obtained.
[0105] In this embodiment of the invention, the training dataset includes multimodal sample data, as well as global semantic category labeling data and target semantic category labeling data corresponding to the sample data. The sample data can be understood as sample input data. The global semantics can be understood as the complete meaning expressed by the sample input data, including all information and concepts contained therein. The categories can be understood as the labels assigned to the global and target semantics of the sample data. The target semantics can be understood as the target semantics of the sample data.
[0106] The labeled data mentioned above can be understood as adding structured labels to the raw data, enabling machine learning models to recognize and process both the raw and labeled data. Through the labeled data, the model can establish a mapping relationship between the input data and the correct output labels.
[0107] The output of the pre-trained multimodal large model is a global semantic prediction vector and the category data of the global semantics corresponding to the global semantic prediction vector, as well as a target semantic prediction vector and the category data of the target semantics corresponding to the target semantic prediction vector.
[0108] The global semantic category data corresponding to the above global semantic prediction vector can be understood as:
[0109] The aforementioned pre-trained multimodal large models can be multimodal large models built based on deep learning or machine learning, such as CLIP, LLM, etc.
[0110] The fine-tuning training described above can be supervised training or model parameter fine-tuning. Supervised training uses a set of data with known labels to train the model, optimizing the model parameters so that the model can predict the labels of new data or make decisions based on the characteristics of existing data. During training, the minimum loss function can be used to adjust the model parameters to minimize the difference between the model's output label and the input data. The loss function measures the difference between the model's prediction and the true result, and its purpose is to improve prediction accuracy by minimizing the loss function value by adjusting the model parameters. The aforementioned loss function can be the mean squared error loss function, cross-entropy loss function, etc.
[0111] The trained multimodal model described above can identify the semantic features of the input data.
[0112] Optionally, in the step of fine-tuning the pre-trained multimodal large model based on the training dataset to obtain the trained multimodal large model, sample data can be input into the pre-trained multimodal large model, and the pre-trained multimodal large model can output global semantic category data and target semantic category data; a first error loss between the global semantic category data and the global semantic category label data can be calculated using a first loss function; and a second error loss between the target semantic category data and the target semantic category label data can be calculated using a second loss function; the total error loss is obtained by weighted calculation based on the first error loss and the second error loss; with minimizing the total error loss as the optimization objective, the model parameters of the pre-trained multimodal large model are adjusted using the backpropagation algorithm, and the adjustment process of the model parameters is iterated until the total error loss is less than a preset value or the number of iterations reaches a preset number, at which point training stops, and the trained multimodal large model is obtained.
[0113] In this embodiment of the invention, the above-mentioned sample data can be understood as sample input data.
[0114] The aforementioned pre-trained multimodal large models can be multimodal large models built based on deep learning or machine learning, such as CLIP, LLM, etc.
[0115] The category data of the global semantics mentioned above can be an understanding and description of the overall sample data, while the category data of the target semantics mentioned above can be an understanding and description of the key data of the sample data.
[0116] The first loss function mentioned above is used to calculate the error loss between the global semantic category data and the global semantic category labeled data. The first loss function can be the cross-entropy loss function, which can be used to calculate the error loss between the global semantic category data and the global semantic category labeled data.
[0117] The aforementioned first error loss can be the error loss between the global semantic category data and the global semantic category annotation data.
[0118] The aforementioned second error loss is calculated using a second loss function to determine the error between the target semantic category data and the target semantic category-labeled data. This second loss function can be a mean squared error loss function, which can be used to calculate the error between the target semantic category data and the target semantic category-labeled data.
[0119] The above weighted calculation can be understood as a calculation scheme that assigns corresponding weights to the first error loss and the second error loss based on their importance or influence.
[0120] The total error loss mentioned above is a measure of the difference between the model's predicted values and the true values, used to guide the adjustment of model parameters to minimize the error.
[0121] Specifically, the optimization objective can be minimizing the total error loss. This can be achieved by using the backpropagation algorithm to adjust the model parameters, iterating through this process until the total error loss is less than a preset value, or the number of iterations reaches a preset number. At this point, the training process ends, resulting in a well-trained multimodal large model. The backpropagation algorithm described above can be understood as an algorithm that updates the weights by calculating the gradient of the loss function to minimize the loss.
[0122] like Figure 2 As shown, an embodiment of the present invention provides an image search device, which includes:
[0123] Module 201 is used to acquire user input data;
[0124] Processing module 202 is used to perform semantic extraction processing on the input data to obtain a global semantic vector and at least one target semantic vector corresponding to the input data;
[0125] Search module 203 is used to perform semantic search in the image database based on the global semantic vector and at least one of the target semantic vectors to obtain the search results corresponding to the input data;
[0126] The return module 204 is used to return the search results.
[0127] Optionally, the processing module 202 is further configured to input the input data into a trained multimodal large model, output multiple candidate global semantic vectors and multiple candidate target semantic vectors through the trained multimodal large model; perform clustering processing on the multiple candidate target semantic vectors to obtain target semantic vector clusters; and select the global semantic vector corresponding to the input data and at least one target semantic vector based on the distance between each candidate global semantic vector and each target semantic vector cluster.
[0128] Optionally, the processing module 202 is further configured to calculate the center vector of each target semantic vector cluster and the compactness of each target semantic vector cluster; calculate the Euclidean distance between each candidate global semantic vector and each center vector to obtain the distance value between each candidate global semantic vector and each target semantic vector cluster; and determine the global semantic vector and at least one target semantic vector corresponding to the input data based on the distance value and the compactness.
[0129] Optionally, the search module 203 calculates a first semantic similarity between the global semantic vector and the image semantic vectors of each image data in the image database; and calculates a second semantic similarity between the image semantic vectors of each image data of each target semantic vector; selects target image data in the image database based on the first semantic similarity and the second semantic similarity; and determines the search results corresponding to the input data based on the target image data.
[0130] Optionally, the search module 203 selects target image data from the image database whose first semantic similarity is greater than a first threshold and whose second semantic similarity is greater than a second threshold; or, for each image data, calculates the comprehensive similarity with the second semantic similarity corresponding to all the target semantic vectors; and selects target image data from the image database whose first semantic similarity is greater than a first threshold and whose comprehensive similarity is greater than a third threshold.
[0131] Optionally, the device is further configured to acquire a training dataset and a pre-trained multimodal large model, wherein the training dataset contains multimodal sample data, and the global semantic category labeling data and target semantic category labeling data corresponding to the sample data; the output of the pre-trained multimodal large model is a global semantic prediction vector and the global semantic category data corresponding to the global semantic prediction vector, and a target semantic prediction vector and the target semantic category data corresponding to the target semantic prediction vector; and to fine-tune the pre-trained multimodal large model based on the training dataset, thereby obtaining a trained multimodal large model after training is completed.
[0132] Optionally, the device is further configured to input the sample data into the pre-trained multimodal large model, output global semantic category data and target semantic category data through the pre-trained multimodal large model; calculate a first error loss between the global semantic category data and the global semantic category labeling data through a first loss function; and calculate a second error loss between the target semantic category data and the target semantic category labeling data through a second loss function; calculate a total error loss based on the first error loss and the second error loss using a weighted average; and adjust the model parameters of the pre-trained multimodal large model using a backpropagation algorithm with the goal of minimizing the total error loss, iterating the adjustment process of the model parameters until the total error loss is less than a preset value or the number of iterations reaches a preset number, then stop training to obtain a trained multimodal large model.
[0133] like Figure 3 As shown, this embodiment of the invention also provides an electronic device, including a processor, which can execute any of the above-described image search methods.
[0134] Specifically, it includes a processor 301 and a memory 302, as well as a computer program stored in the memory 302 and capable of running on the processor 301 to execute the image search method, wherein:
[0135] Processor 301 executes the calculator program for the image search method stored in memory 302, performing the following steps:
[0136] Obtain user input data;
[0137] The input data is subjected to semantic extraction processing to obtain a global semantic vector and at least one target semantic vector corresponding to the input data;
[0138] Based on the global semantic vector and at least one of the target semantic vectors, a semantic search is performed in the image database to obtain the search results corresponding to the input data;
[0139] Return the search results.
[0140] Optionally, the semantic extraction processing performed by the processor 301 on the input data to obtain a global semantic vector corresponding to the input data and at least one target semantic vector includes:
[0141] The input data is input into a trained multimodal large model, and the trained multimodal large model outputs multiple candidate global semantic vectors and multiple candidate target semantic vectors.
[0142] Clustering is performed on multiple candidate target semantic vectors to obtain target semantic vector clusters;
[0143] Based on the distance between each candidate global semantic vector and each target semantic vector cluster, the global semantic vector corresponding to the input data and at least one target semantic vector are selected.
[0144] Optionally, the step of selecting the global semantic vector corresponding to the input data and at least one target semantic vector based on the distance between each of the candidate global semantic vectors and each of the target semantic vector clusters, performed by the processor 301, includes:
[0145] Calculate the center vector of each of the target semantic vector clusters, and calculate the compactness of each of the target semantic vector clusters;
[0146] Calculate the Euclidean distance between each candidate global semantic vector and each of the center vectors to obtain the distance value between each candidate global semantic vector and each of the target semantic vector clusters;
[0147] Based on the distance value and the density, the global semantic vector and at least one target semantic vector corresponding to the input data are determined.
[0148] Optionally, the process executed by processor 301 to perform a semantic search in the image database based on the global semantic vector and at least one of the target semantic vectors to obtain the search results corresponding to the input data includes:
[0149] Calculate the first semantic similarity between the global semantic vector and the image semantic vectors of each image data in the image database;
[0150] And, calculate the second semantic similarity between the image semantic vectors of each image data for each target semantic vector;
[0151] Based on the first semantic similarity and the second semantic similarity, target image data is selected from the image database;
[0152] Based on the target image data, the search results corresponding to the input data are determined.
[0153] Optionally, the step of selecting target image data as the search result corresponding to the input data based on the first semantic similarity and the second semantic similarity, performed by the processor 301, includes:
[0154] Select target image data from the image database whose first semantic similarity is greater than a first threshold and whose second semantic similarity is greater than a second threshold.
[0155] Alternatively, for each image data, calculate the comprehensive similarity with the second semantic similarity corresponding to all the target semantic vectors;
[0156] Target image data with a first semantic similarity greater than a first threshold and a comprehensive similarity greater than a third threshold are selected from the image database.
[0157] Optionally, before inputting the input data into the trained multimodal large model and outputting multiple candidate global semantic vectors and multiple candidate target semantic vectors through the trained multimodal large model, the method executed by the processor 301 further includes:
[0158] Acquire a training dataset and a pre-trained multimodal large model. The training dataset contains multimodal sample data, global semantic category labeling data, and target semantic category labeling data corresponding to the sample data. The output of the pre-trained multimodal large model is a global semantic prediction vector and the global semantic category data corresponding to the global semantic prediction vector, and a target semantic prediction vector and the target semantic category data corresponding to the target semantic prediction vector.
[0159] The pre-trained multimodal large model is fine-tuned based on the training dataset. After training, the trained multimodal large model is obtained.
[0160] Optionally, the processor 301 performs fine-tuning training on the pre-trained multimodal large model based on the training dataset. After training, a trained multimodal large model is obtained, including:
[0161] The sample data is input into the pre-trained multimodal large model, and the pre-trained multimodal large model outputs global semantic category data and target semantic category data.
[0162] The first error loss between the category data of the global semantics and the category labeling data of the global semantics is calculated using a first loss function;
[0163] And, a second error loss is calculated between the category data of the target semantics and the category annotation data of the target semantics using a second loss function;
[0164] The total error loss is calculated by weighting the first error loss and the second error loss.
[0165] With minimizing the total error loss as the optimization objective, the model parameters of the pre-trained multimodal large model are adjusted through the backpropagation algorithm. The adjustment process of the model parameters is iterated until the total error loss is less than a preset value or the number of iterations reaches a preset number, at which point training stops and a trained multimodal large model is obtained.
[0166] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the image search method provided in this invention and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0167] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0168] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. An image search method, characterized in that, The method includes the following steps: Obtain user input data; The input data is subjected to semantic extraction processing to obtain a global semantic vector and at least one target semantic vector corresponding to the input data. Specifically, the input data is input into a trained multimodal large model, which outputs multiple candidate global semantic vectors and multiple candidate target semantic vectors. The multiple candidate target semantic vectors are clustered to obtain target semantic vector clusters. The center vector of each target semantic vector cluster is calculated, as well as the compactness of each target semantic vector cluster. The Euclidean distance between each candidate global semantic vector and each center vector is calculated to obtain the distance value between each candidate global semantic vector and each target semantic vector cluster. Based on the distance value and the compactness, the global semantic vector and at least one target semantic vector corresponding to the input data are determined. Based on the global semantic vector and at least one target semantic vector, a semantic search is performed in the image database to obtain the search results corresponding to the input data. Specifically, a first semantic similarity is calculated between the global semantic vector and the image semantic vectors of each image data in the image database; and a second semantic similarity is calculated between the image semantic vectors of each image data for each target semantic vector; based on the first semantic similarity and the second semantic similarity, target image data is selected in the image database; and the search results corresponding to the input data are determined according to the target image data. Return the search results.
2. The image search method as described in claim 1, characterized in that, The step of selecting target image data from the image database as the search result corresponding to the input data based on the first semantic similarity and the second semantic similarity includes: Select target image data from the image database whose first semantic similarity is greater than a first threshold and whose second semantic similarity is greater than a second threshold. Alternatively, for each image data, calculate the comprehensive similarity with the second semantic similarity corresponding to all the target semantic vectors; Target image data with a first semantic similarity greater than a first threshold and a comprehensive similarity greater than a third threshold are selected from the image database.
3. The image search method as described in any one of claims 1 to 2, characterized in that, Before inputting the input data into the trained multimodal large model and outputting multiple candidate global semantic vectors and multiple candidate target semantic vectors through the trained multimodal large model, the method further includes: Acquire a training dataset and a pre-trained multimodal large model. The training dataset contains multimodal sample data, global semantic category labeling data, and target semantic category labeling data corresponding to the sample data. The output of the pre-trained multimodal large model is a global semantic prediction vector and the global semantic category data corresponding to the global semantic prediction vector, and a target semantic prediction vector and the target semantic category data corresponding to the target semantic prediction vector. The pre-trained multimodal large model is fine-tuned based on the training dataset. After training, the trained multimodal large model is obtained.
4. The image search method as described in claim 3, characterized in that, The process of fine-tuning the pre-trained multimodal large model based on the training dataset, and obtaining a trained multimodal large model after training, includes: The sample data is input into the pre-trained multimodal large model, and the pre-trained multimodal large model outputs global semantic category data and target semantic category data. The first error loss between the category data of the global semantics and the category labeling data of the global semantics is calculated using a first loss function; And, a second error loss is calculated between the category data of the target semantics and the category annotation data of the target semantics using a second loss function; The total error loss is calculated by weighting the first error loss and the second error loss. With minimizing the total error loss as the optimization objective, the model parameters of the pre-trained multimodal large model are adjusted through the backpropagation algorithm. The adjustment process of the model parameters is iterated until the total error loss is less than a preset value or the number of iterations reaches a preset number, at which point training stops and a trained multimodal large model is obtained.
5. An image search device, characterized in that, The image search device includes: The acquisition module is used to acquire user input data; The processing module is used to perform semantic extraction processing on the input data to obtain a global semantic vector and at least one target semantic vector corresponding to the input data. Specifically, the input data is input into a trained multimodal large model, which outputs multiple candidate global semantic vectors and multiple candidate target semantic vectors. The multiple candidate target semantic vectors are clustered to obtain target semantic vector clusters. The center vector of each target semantic vector cluster is calculated, and the compactness of each target semantic vector cluster is calculated. The Euclidean distance between each candidate global semantic vector and each center vector is calculated to obtain the distance value between each candidate global semantic vector and each target semantic vector cluster. Based on the distance value and the compactness, the global semantic vector and at least one target semantic vector corresponding to the input data are determined. The search module is used to perform a semantic search in an image database based on the global semantic vector and at least one target semantic vector to obtain search results corresponding to the input data. Specifically, it calculates a first semantic similarity between the global semantic vector and the image semantic vectors of each image data in the image database; and calculates a second semantic similarity between the image semantic vectors of each image data for each target semantic vector; selects target image data in the image database based on the first semantic similarity and the second semantic similarity; and determines the search results corresponding to the input data based on the target image data. The return module is used to return the search results.
6. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the image search method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image search method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Product recommendation method and device, computer equipment and storage medium
CN120163634A
Method and apparatus for unsupervised learning of multi-resolution user profile from text analysis
US20140229486A1