Image searching method and device, electronic equipment and storage medium
By combining image search methods with small-scale models and multimodal large-scale models, the problem of querying elements that require users to understand database retrieval elements is solved, achieving efficient and accurate image search.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
- Filing Date
- 2025-11-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing image search methods require users to know the search elements for the target in the database, which increases the difficulty of the query and reduces work efficiency.
By obtaining the first feature vector to be searched, a first search is performed in the first event database using the small model algorithm to obtain candidate search results. Then, based on the candidate search results, a second feature vector to be searched is determined. Finally, a second search is performed in the second event database using the multimodal large model algorithm to obtain the target search results.
It reduces the difficulty for users to query image data and improves the accuracy and efficiency of image search.
Smart Images

Figure CN121901446A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image search technology, and in particular relates to an image search method, apparatus, electronic device and storage medium. Background Technology
[0002] When searching image data, the process typically involves retrieving matching images based on user input and database query rules. This requires users to understand and learn these rules, needing to know the corresponding search elements in the database to accurately find the relevant image. This increases the difficulty of image data retrieval for users. Therefore, there is an urgent need for a highly accurate image search method to address the problem of existing methods requiring knowledge of the target's corresponding search elements in the database to accurately retrieve the image, thus increasing the difficulty of image data retrieval and leading to low efficiency. Summary of the Invention
[0003] This application provides an image search method that solves the problem that existing search methods require knowing the corresponding search elements in the database to accurately find the corresponding image, which increases the difficulty of image data retrieval for users and leads to low work efficiency.
[0004] In a first aspect, embodiments of this application provide an image search method, the method comprising the following steps:
[0005] Obtain the first feature vector to be searched;
[0006] Based on the first feature vector to be searched, a first search is performed in the first event database to obtain candidate search results corresponding to the first feature vector to be searched. The first event database includes the first feature vector extracted by the small model algorithm.
[0007] Based on the first feature vector to be searched and the candidate search results, a second feature vector to be searched is determined;
[0008] Based on the second feature vector to be searched, a second search is performed in the second event database to obtain the target search result. The second event database includes the second feature vector extracted by the multimodal large model algorithm.
[0009] Optionally, obtaining the first feature vector to be searched includes:
[0010] Obtain the first data to be searched;
[0011] The first search data is subjected to feature extraction processing to obtain the first search feature vector.
[0012] Optionally, determining the second search feature vector based on the first search feature vector and the candidate search results includes:
[0013] Determine the hyperparameters of the candidate search results;
[0014] Based on the hyperparameters, a weighted calculation is performed on the first feature vector to be searched and the candidate search results to obtain the second feature vector to be searched.
[0015] Optionally, before performing a first search in the first event database based on the first search data to obtain candidate search results corresponding to the first search data, the method further includes:
[0016] Retrieve base database event data;
[0017] The first event data is obtained by performing event detection on the base database event data using a preset small model algorithm;
[0018] The first event data is subjected to feature extraction processing to obtain the first feature vector corresponding to the first event data;
[0019] Based on the first feature vector, a first event library is constructed.
[0020] Optionally, before performing a second search in the second event database based on the second search feature vector to obtain the target search result, the method further includes:
[0021] The second event data is obtained by performing event detection on the base database event data using a preset multimodal large model algorithm;
[0022] Multimodal feature extraction processing is performed on the second event data to obtain the second feature vector corresponding to the second event data;
[0023] Based on the second feature vector, a second event library is constructed.
[0024] Optionally, the step of constructing the second event library based on the second feature vector includes:
[0025] The second feature vector is clustered according to time and location to obtain the clustering results;
[0026] The clustering results are deduplicated to obtain the duplicated second feature vector;
[0027] Based on the deduplicated second feature vector, a second event library is constructed.
[0028] Optionally, building the second event library further includes:
[0029] The first event data is processed by multimodal feature extraction using a pre-defined multimodal large model algorithm to obtain the second feature vector corresponding to the first event.
[0030] A second event library is constructed based on the second feature vector corresponding to the first event.
[0031] Secondly, embodiments of this application provide an image search device, the image search device comprising:
[0032] The acquisition module is used to acquire the first feature vector to be searched.
[0033] The first search module is used to perform a first search in the first event library based on the first feature vector to be searched, and to obtain candidate search results corresponding to the first feature vector to be searched. The first event library includes the first feature vector extracted by the small model algorithm.
[0034] The determining module is used to determine a second feature vector to be searched based on the first feature vector to be searched and the candidate search results;
[0035] The second search module is used to perform a second search in the second event library based on the second feature vector to be searched, and to obtain the target search result. The second event library includes the second feature vector extracted by the multimodal large model algorithm.
[0036] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the image search method provided in embodiments of the present invention.
[0037] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the image search method provided in the embodiments of the present invention.
[0038] The above-described solution of this application has the following beneficial effects: First, a first feature vector to be searched is obtained; a first search is performed in a first event database based on the first feature vector to be searched to obtain candidate search results corresponding to the first feature vector to be searched, wherein the first event database includes the first feature vector extracted by a small model algorithm; a second feature vector to be searched is determined based on the first feature vector to be searched and the candidate search results; a second search is performed in a second event database based on the second feature vector to be searched to obtain the target search result, wherein the second event database includes the second feature vector extracted by a multimodal large model algorithm. This invention solves the problem that existing search methods require understanding the corresponding search elements in the database to accurately retrieve the corresponding image, increasing the difficulty of image data retrieval for users and leading to low work efficiency.
[0039] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A flowchart illustrating an image search method provided in one embodiment of this application;
[0042] Figure 2 This is a schematic diagram of the structure of an image search device provided in an embodiment of this application;
[0043] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0044] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0045] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0046] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0047] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0048] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0049] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0050] like Figure 1 As shown, Figure 1 This is a flowchart of an image search method provided by an embodiment of the present invention. The image search method includes the following steps:
[0051] 101. Obtain the first feature vector to be searched.
[0052] In this embodiment of the invention, the image search method described above can be applied to an image search platform, which can be built on a server-based or distributed platform. The image search platform includes a data interface (for sensors or users to upload data), a knowledge database, and a knowledge database construction program. The data interface can be used to obtain a first feature vector to be searched, and the knowledge database construction program can be used to construct the knowledge database. The knowledge database is specifically used to provide additional association information for the identified data entities, thereby improving the depth of the data recognition system's understanding of the content.
[0053] The first feature vector to be searched can be understood as a feature description of the image or object input by the user. This first feature vector can be a vector feature of the image being searched, such as color, shape, and texture. The feature vector is a concept in linear algebra; it refers to a non-zero vector whose direction remains unchanged or is only scaled after a matrix transformation.
[0054] It should be noted that a first searchable data vector can be obtained by acquiring the user's input as initial search data and performing feature extraction on it. This initial search data can be the text content entered by the user during the image search process. The feature extraction process described above can be understood as the process of analyzing the initial search data and extracting its feature vectors.
[0055] 102. Based on the first feature vector to be searched, perform a first search in the first event database to obtain the candidate search results corresponding to the first feature vector to be searched.
[0056] In this embodiment of the invention, the first event library includes a first feature vector extracted by a small-model algorithm. The small-model can be understood as a model with a small parameter size that focuses on a specific task. The small-model algorithm can be understood as an algorithm that optimizes search efficiency; it can search the small-model event library using the first feature vector to be searched to obtain candidate search results corresponding to the first feature vector, thereby reducing the search space.
[0057] The aforementioned small model is obtained by training an untrained small model using a training dataset. This untrained small model can be a small model built based on deep learning or machine learning, such as a CNN or RNN. The training dataset includes sample search data and corresponding search feature vector annotations. These annotations add structured labels to the original data, enabling the machine learning model to recognize and process both the original and labeled data. Through the annotations, the model can establish a mapping between input data and the correct output label. The training can be supervised, which uses a set of data with known labels to train the model. By optimizing the model parameters, the model can predict the label of new data or make decisions based on the characteristics of existing data. During training, a minimum loss function can be used to adjust the model parameters to minimize the difference between the model's output label and the input data. The loss function measures the difference between the model's prediction and the true result, aiming to improve prediction accuracy by minimizing the loss function value through adjusting the model parameters. The aforementioned loss function can be a mean squared error loss function, a cross-entropy loss function, etc. The aforementioned small model can identify the feature vectors of the search data.
[0058] The aforementioned first event library can be a small model event library, a database used to store and manage events.
[0059] The aforementioned first search can be understood as a process of comparing the similarity between the first feature vector to be searched and the corresponding first feature vector in the first event database to obtain the candidate search results corresponding to the first feature vector to be searched.
[0060] The aforementioned candidate search results can be obtained by performing a first search in the first event database based on the first feature vector to be searched, and obtaining the search results corresponding to the first feature vector to be searched.
[0061] Specifically, the first feature vector to be searched can be compared with the first feature vector in the first event database. M first feature vectors whose similarity to the first feature vector is greater than a preset similarity threshold are selected as candidate search results. The preset similarity threshold can be a similarity threshold pre-set by the system.
[0062] 103. Based on the first feature vector to be searched and the candidate search results, determine the second feature vector to be searched.
[0063] In this embodiment of the invention, a second feature vector to be searched can be determined based on a first feature vector to be searched and candidate search results.
[0064] Furthermore, the hyperparameters of the candidate search results can be determined first, and the first search feature vector can be weighted and calculated using the hyperparameters of the candidate search results to obtain the second search feature vector.
[0065] The hyperparameters for the candidate search results mentioned above are pre-set. It should be noted that these hyperparameters can be adjusted according to actual needs. The weighted calculation described above can be understood as assigning different weights to different data points to calculate a comprehensive result. Specifically, each data point can be multiplied by its corresponding weight, and then all products can be summed to obtain the final result.
[0066] The second feature vector to be searched is a feature vector determined based on the first feature vector to be searched and the candidate search results.
[0067] 104. Based on the second feature vector to be searched, perform a second search in the second event database to obtain the target search result.
[0068] In this embodiment of the invention, the second event library includes a second feature vector extracted by a multimodal large model algorithm. The multimodal large model possesses cross-modal information fusion capabilities, enabling it to simultaneously process data from different modalities, such as text descriptions and image verification, capturing richer features and correlations, and enhancing the depth of understanding of complex scenes. The multimodal large model algorithm can be understood as a large model algorithm capable of processing multiple different modalities of data, such as text descriptions and image verification.
[0069] The aforementioned multimodal large model is obtained by training a pre-trained multimodal large model on a multimodal training set. This pre-trained multimodal large model can be a multimodal large model built based on deep learning or machine learning, such as CLIP, MLLMs, VSM, etc. The multimodal training set includes sample multimodal search data and corresponding feature vector annotation data. The annotation data can be structured labels added to the original data, enabling the machine learning model to recognize and process both the original and labeled data. Through the annotation data, the model can establish a mapping relationship between input data and the correct output label. The training can be supervised training, which uses a set of data with known labels to train the model. By optimizing the model parameters, the model can predict the label of new data or make decisions based on the characteristics of existing data. During training, the minimum loss function can be used to adjust the model parameters to minimize the difference between the model's output label and the input data. The loss function measures the difference between the model's prediction and the true result, aiming to improve prediction accuracy by minimizing the loss function value through adjusting the model parameters. The loss function can be the mean squared error loss function, cross-entropy loss function, etc. The aforementioned multimodal large model can identify the feature vectors of multimodal search data.
[0070] The aforementioned second event library can be a multimodal large model database, used for recording, organizing, and managing data.
[0071] The aforementioned second search can be understood as a process of comparing the similarity between the second feature vector to be searched and the corresponding second feature vector in the second event database to obtain the target search result.
[0072] Specifically, the second feature vector to be searched can be compared with the second feature vectors in the second event database. The K second feature vectors whose similarity to the second feature vector is greater than a preset similarity threshold are selected as the target search results. The preset similarity threshold can be a similarity threshold pre-set by the system.
[0073] The target search result is obtained by performing a second search on the second feature vector to be searched and the corresponding second feature vector in the second event library.
[0074] Furthermore, multiple feature vectors from the target search results can be queried in an image database to retrieve corresponding image data, which is then displayed to the user in a list format. This image database includes the feature vector corresponding to each image data.
[0075] In this embodiment of the invention, by combining a large multimodal model and a small multimodal model, the invention provides a large model event library and a small model event library for searching in the multimodal search task, which can reduce the vector offset between image modalities and thus improve the accuracy of image search.
[0076] In this embodiment of the invention, a first feature vector to be searched is obtained; a first search is performed on a first event database based on the first feature vector to be searched to obtain candidate search results corresponding to the first feature vector to be searched. The first event database includes the first feature vector extracted by a small model algorithm; a second feature vector to be searched is determined based on the first feature vector to be searched and the candidate search results; a second search is performed on a second event database based on the second feature vector to be searched to obtain the target search result. The second event database includes the second feature vector extracted by a multimodal large model algorithm. This invention solves the problem that existing search methods require understanding the corresponding search elements in a database to accurately retrieve the corresponding image, increasing the difficulty of image data retrieval for users and leading to low work efficiency.
[0077] It is understood that in the specific implementation of this application, data such as search data, feature vector data, knowledge data, and user data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required. Furthermore, the collection, use, and processing of related data, as well as the training, deployment, and invocation of algorithm models, must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0078] Optionally, in the step of obtaining the first feature vector to be searched, the first data to be searched can be obtained; feature extraction processing can be performed on the first data to be searched to obtain the first feature vector to be searched.
[0079] In this embodiment of the invention, the first search data can be understood as the text content entered by the user during the image search process.
[0080] The above feature extraction process can be understood as the process of extracting feature vectors from the first search data by analyzing the first search data.
[0081] The first feature vector to be searched can be the image vector features to be searched, such as color, shape, texture, etc.
[0082] It should be noted that the first feature vector to be searched can effectively represent the first data to be searched.
[0083] Optionally, in the step of determining the second feature vector to be searched based on the first feature vector to be searched and the candidate search results, the hyperparameters of the candidate search results can be determined; based on the hyperparameters, the first feature vector to be searched and the candidate search results are weighted and calculated to obtain the second feature vector to be searched.
[0084] In this embodiment of the invention, the aforementioned candidate search results can be obtained by comparing the first feature vector to be searched with a first feature vector in a first event database, and selecting multiple first feature vectors whose similarity to the first feature vector is greater than a preset similarity threshold as candidate search results. The candidate search results include multiple first feature vectors. The aforementioned first event database includes first feature vectors extracted by a small model algorithm.
[0085] The hyperparameters mentioned above can be pre-set. It should be noted that hyperparameters can be set according to actual conditions.
[0086] The weighted calculation described above can be understood as calculating a comprehensive result by assigning different weights to different data. Specifically, each data point can be multiplied by its assigned weight, and then all the products can be added together to obtain the weighted calculation result.
[0087] The aforementioned second feature vector to be searched is obtained by weighting the first feature vector to be searched and the candidate search results according to the hyperparameters of the candidate search results.
[0088] Specifically, the second feature vector to be searched can be calculated using the following formula:
[0089]
[0090] in, This represents the second feature vector to be searched. This represents the first feature vector to be searched. Denotes the hyperparameter, where K represents the number of first feature vectors in the candidate search results. This represents the first eigenvector.
[0091] Optionally, before performing a first search in the first event database based on the first search data to obtain candidate search results corresponding to the first search data, it is also possible to obtain base database event data; perform event detection on the base database event data using a preset small model algorithm to obtain first event data; perform feature extraction processing on the first event data to obtain the first feature vector corresponding to the first event data; and construct the first event database based on the first feature vector.
[0092] In this embodiment of the invention, the aforementioned base database event data can be understood as data used by the base database for recording and management. The aforementioned base database may be a database responsible for data storage, retrieval, and transaction management.
[0093] The aforementioned small models can be understood as models with a small parameter scale that focus on a specific task. The aforementioned preset small model algorithm can be understood as an algorithm that optimizes search efficiency, capable of quickly processing data and providing results.
[0094] The aforementioned small model is obtained by training an untrained small model using a training dataset. This untrained small model can be a small model built based on deep learning or machine learning, such as a CNN or RNN. The training dataset includes sample search data and corresponding search feature vector annotations. These annotations add structured labels to the original data, enabling the machine learning model to recognize and process both the original and labeled data. Through the annotations, the model can establish a mapping between input data and the correct output label. The training can be supervised, which uses a set of data with known labels to train the model. By optimizing the model parameters, the model can predict the label of new data or make decisions based on the characteristics of existing data. During training, a minimum loss function can be used to adjust the model parameters to minimize the difference between the model's output label and the input data. The loss function measures the difference between the model's prediction and the true result, aiming to improve prediction accuracy by minimizing the loss function value through adjusting the model parameters. The aforementioned loss function can be a mean squared error loss function, a cross-entropy loss function, etc. The aforementioned small model can identify the feature vectors of the search data. Specifically, the optimization objective can be minimizing error loss. This can be achieved by adjusting model parameters using the backpropagation algorithm, iterating until the loss is less than a preset value or the preset number of iterations is reached, at which point the training process ends, resulting in a trained small model. The backpropagation algorithm described above updates the weights by calculating the gradient of the loss function to minimize the loss.
[0095] The above event detection can be understood as the process of detecting events in the base database event data through a preset small model algorithm.
[0096] The aforementioned first event data can be understood as event data detected through a preset small model algorithm.
[0097] The above feature extraction process can be understood as the process of extracting feature vectors that can represent the first event data from the first event data.
[0098] The aforementioned first event library includes the first feature vector corresponding to the first event data.
[0099] It should be noted that the first event data can be processed by a preset small model algorithm to obtain the first feature vector corresponding to the first event data, and the first event library can be constructed from the first feature vector.
[0100] Optionally, before performing a second search in the second event database based on the second search feature vector to obtain the target search result, an event detection can be performed on the base database event data using a preset multimodal large model algorithm to obtain second event data; multimodal feature extraction processing can be performed on the second event data to obtain the second feature vector corresponding to the second event data; and the second event database can be constructed based on the second feature vector.
[0101] In this embodiment of the invention, the aforementioned preset multimodal large model algorithm can be understood as a large model algorithm capable of processing multiple different modalities of data, such as text description + image verification. The aforementioned multimodal large model possesses cross-modal information fusion capabilities, enabling it to simultaneously process data from different modalities, such as text description + image verification, capturing richer features and correlations, and enhancing the depth of understanding of complex scenarios.
[0102] The aforementioned multimodal large model is obtained by training a pre-trained multimodal large model on a multimodal training set. This pre-trained multimodal large model can be a multimodal large model built based on deep learning or machine learning, such as CLIP or LLM. The multimodal training set includes sample multimodal search data and corresponding feature vector annotation data. The annotation data can be structured labels added to the original data, enabling the machine learning model to recognize and process both the original and labeled data. Through the annotation data, the model can establish a mapping relationship between input data and the correct output label. The training can be supervised training, which uses a set of data with known labels to train the model. By optimizing the model parameters, the model can predict the label of new data or make decisions based on the characteristics of existing data. During training, the minimum loss function can be used to adjust the model parameters to minimize the difference between the model's output label and the input data. The loss function measures the difference between the model's prediction and the true result, aiming to improve prediction accuracy by minimizing the loss function value through adjusting the model parameters. The aforementioned loss function can be the mean squared error loss function, cross-entropy loss function, etc. The aforementioned multimodal large model can identify the feature vectors of multimodal search data.
[0103] The aforementioned base database event data can be understood as the data used by the base database for recording and management. This base database can be a database responsible for data storage, retrieval, and transaction management.
[0104] The above event detection process involves processing the base database event data using a pre-defined multimodal model algorithm.
[0105] The aforementioned second event data can be understood as event data detected through a preset multimodal large model.
[0106] The above multimodal feature extraction process can be understood as the process of extracting multimodal feature vectors that can represent the second event data from the second event data.
[0107] The aforementioned second event library includes the second feature vector corresponding to the second event data.
[0108] It should be noted that the base event data can be detected by a preset multimodal large model algorithm to obtain the second event data. Multimodal feature extraction processing is then performed on the second event data to obtain the second feature vector corresponding to the second event data. The second event database is then constructed based on the second feature vector.
[0109] Optionally, in the step of constructing the second event library based on the second feature vector, the second feature vector can be clustered according to time and location to obtain clustering results; the clustering results can be deduplicated to obtain duplicated second feature vectors; and the second event library can be constructed based on the deduplicated second feature vectors.
[0110] In this embodiment of the invention, the above clustering can be understood as a process of dividing the second feature vector into clustering results composed of similar feature vectors according to time and location.
[0111] The above clustering result is a dataset obtained by clustering the second feature vector according to time and location.
[0112] The deduplication process described above can be understood as the process of removing duplicate second feature vectors from the clustering results.
[0113] The aforementioned second event library is constructed using the deduplicated second feature vector.
[0114] It should be noted that the second feature vector can be clustered according to time and location to obtain clustering results. The clustering results can then be deduplicated to obtain duplicate second feature vectors. The second event database can then be constructed based on these deduplicated second feature vectors.
[0115] Optionally, the second event library can also be constructed by performing multimodal feature extraction processing on the first event data using a preset multimodal large model algorithm to obtain the second feature vector corresponding to the first event; and constructing the second event library based on the second feature vector corresponding to the first event.
[0116] In this embodiment of the invention, the aforementioned preset multimodal large model algorithm can be understood as a large model algorithm capable of processing various different modal data, such as text description + image verification.
[0117] The aforementioned first event data can be understood as event data detected through a preset small model algorithm.
[0118] The multimodal feature extraction process described above can be understood as the process of extracting multimodal features from the first event data using a pre-defined multimodal large model algorithm. Multimodal features can be text features, image features, etc.
[0119] The second feature vector mentioned above can be the feature vector corresponding to the first event.
[0120] The aforementioned second event library can be a multimodal large-scale event library.
[0121] Specifically, the first event data can be processed by multimodal feature extraction using a pre-defined multimodal large model algorithm to obtain the second feature vector corresponding to the first event, and the second event library can be constructed from the second feature vector corresponding to the first event.
[0122] like Figure 2 As shown, an embodiment of the present invention provides an image search device, which includes:
[0123] The acquisition module 201 is used to acquire the first feature vector to be searched;
[0124] The first search module 202 is used to perform a first search in a first event library based on the first feature vector to be searched, and to obtain candidate search results corresponding to the first feature vector to be searched. The first event library includes the first feature vector extracted by the small model algorithm.
[0125] The determining module 203 is used to determine a second search feature vector based on the first search feature vector and the candidate search results;
[0126] The second search module 204 is used to perform a second search in the second event library based on the second feature vector to be searched, and to obtain the target search result. The second event library includes the second feature vector extracted by the multimodal large model algorithm.
[0127] Optionally, the acquisition module 201 is further configured to acquire first search data; perform feature extraction processing on the first search data to obtain a first search feature vector.
[0128] Optionally, the determining module 203 is further configured to determine the hyperparameters of the candidate search results; based on the hyperparameters, a weighted calculation is performed on the first search feature vector and the candidate search results to obtain a second search feature vector.
[0129] Optionally, the device is further configured to acquire base database event data; perform event detection on the base database event data using a preset small model algorithm to obtain first event data; perform feature extraction processing on the first event data to obtain a first feature vector corresponding to the first event data; and construct a first event database based on the first feature vector.
[0130] Optionally, the device is further configured to perform event detection on the base event data using a preset multimodal large model algorithm to obtain second event data; perform multimodal feature extraction processing on the second event data to obtain a second feature vector corresponding to the second event data; and construct a second event database based on the second feature vector.
[0131] Optionally, the device is further configured to cluster the second feature vector according to time and location to obtain a clustering result; to perform deduplication processing on the clustering result to obtain a duplicated second feature vector; and to construct a second event library based on the deduplicated second feature vector.
[0132] Optionally, the device is further configured to perform multimodal feature extraction processing on the first event data using a preset multimodal large model algorithm to obtain a second feature vector corresponding to the first event; and to construct a second event library based on the second feature vector corresponding to the first event.
[0133] like Figure 3 As shown, this embodiment of the invention also provides an electronic device, including a processor, which can execute any of the above-described image search methods.
[0134] Specifically, it includes a processor 301 and a memory 302, as well as a computer program stored in the memory 302 and capable of running on the processor 301 to execute the image search method, wherein:
[0135] Processor 301 executes the calculator program for the image search method stored in memory 302, performing the following steps:
[0136] Obtain the first feature vector to be searched;
[0137] Based on the first feature vector to be searched, a first search is performed in the first event database to obtain candidate search results corresponding to the first feature vector to be searched. The first event database includes the first feature vector extracted by the small model algorithm.
[0138] Based on the first feature vector to be searched and the candidate search results, a second feature vector to be searched is determined;
[0139] Based on the second feature vector to be searched, a second search is performed in the second event database to obtain the target search result. The second event database includes the second feature vector extracted by the multimodal large model algorithm.
[0140] Optionally, the process of obtaining the first feature vector to be searched, performed by processor 301, includes:
[0141] Obtain the first data to be searched;
[0142] The first search data is subjected to feature extraction processing to obtain the first search feature vector.
[0143] Optionally, the process of determining the second search feature vector based on the first search feature vector and the candidate search results executed by the processor 301 includes:
[0144] Determine the hyperparameters of the candidate search results;
[0145] Based on the hyperparameters, a weighted calculation is performed on the first feature vector to be searched and the candidate search results to obtain the second feature vector to be searched.
[0146] Optionally, before performing a first search in the first event database based on the first search data to obtain candidate search results corresponding to the first search data, the method executed by the processor 301 further includes:
[0147] Retrieve base database event data;
[0148] The first event data is obtained by performing event detection on the base database event data using a preset small model algorithm;
[0149] The first event data is subjected to feature extraction processing to obtain the first feature vector corresponding to the first event data;
[0150] Based on the first feature vector, a first event library is constructed.
[0151] Optionally, before performing a second search in the second event database based on the second search feature vector to obtain the target search result, the method executed by the processor 301 further includes:
[0152] The second event data is obtained by performing event detection on the base database event data using a preset multimodal large model algorithm;
[0153] Multimodal feature extraction processing is performed on the second event data to obtain the second feature vector corresponding to the second event data;
[0154] Based on the second feature vector, a second event library is constructed.
[0155] Optionally, the process executed by processor 301 to construct a second event library based on the second feature vector includes:
[0156] The second feature vector is clustered according to time and location to obtain the clustering results;
[0157] The clustering results are deduplicated to obtain the duplicated second feature vector;
[0158] Based on the deduplicated second feature vector, a second event library is constructed.
[0159] Optionally, the construction of the second event library executed by processor 301 further includes:
[0160] The first event data is processed by multimodal feature extraction using a pre-defined multimodal large model algorithm to obtain the second feature vector corresponding to the first event.
[0161] A second event library is constructed based on the second feature vector corresponding to the first event.
[0162] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the image search method provided in this invention and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0163] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An image search method, characterized in that, The method includes the following steps: Obtain the first feature vector to be searched; Based on the first feature vector to be searched, a first search is performed in the first event database to obtain candidate search results corresponding to the first feature vector to be searched. The first event database includes the first feature vector extracted by the small model algorithm. Based on the first feature vector to be searched and the candidate search results, a second feature vector to be searched is determined; Based on the second feature vector to be searched, a second search is performed in the second event database to obtain the target search result. The second event database includes the second feature vector extracted by the multimodal large model algorithm.
2. The image search method as described in claim 1, characterized in that, The step of obtaining the first feature vector to be searched includes: Obtain the first data to be searched; The first search data is subjected to feature extraction processing to obtain the first search feature vector.
3. The image search method as described in claim 2, characterized in that, The step of determining the second search feature vector based on the first search feature vector and the candidate search results includes: Determine the hyperparameters of the candidate search results; Based on the hyperparameters, a weighted calculation is performed on the first feature vector to be searched and the candidate search results to obtain the second feature vector to be searched.
4. The image search method as described in claim 1, characterized in that, Before performing a first search in the first event database based on the first search data to obtain candidate search results corresponding to the first search data, the method further includes: Retrieve base database event data; The first event data is obtained by performing event detection on the base database event data using a preset small model algorithm; The first event data is subjected to feature extraction processing to obtain the first feature vector corresponding to the first event data; Based on the first feature vector, a first event library is constructed.
5. The image search method as described in claim 4, characterized in that, Before performing a second search on the second event database based on the second search feature vector to obtain the target search result, the method further includes: The second event data is obtained by performing event detection on the base database event data using a preset multimodal large model algorithm; Multimodal feature extraction processing is performed on the second event data to obtain the second feature vector corresponding to the second event data; Based on the second feature vector, a second event library is constructed.
6. The image search method as described in claim 5, characterized in that, The second event library is constructed based on the second feature vector, including: The second feature vector is clustered according to time and location to obtain the clustering results; The clustering results are deduplicated to obtain the second feature vector after deduplication; Based on the deduplicated second feature vector, a second event library is constructed.
7. The image search method as described in any one of claims 4-6, characterized in that, The construction of the second event library also includes: The first event data is processed by multimodal feature extraction using a pre-defined multimodal large model algorithm to obtain the second feature vector corresponding to the first event. A second event library is constructed based on the second feature vector corresponding to the first event.
8. An image search device, characterized in that, The image search device includes: The acquisition module is used to acquire the first feature vector to be searched. The first search module is used to perform a first search in the first event library based on the first feature vector to be searched, and to obtain candidate search results corresponding to the first feature vector to be searched. The first event library includes the first feature vector extracted by the small model algorithm. The determining module is used to determine a second feature vector to be searched based on the first feature vector to be searched and the candidate search results; The second search module is used to perform a second search in the second event library based on the second feature vector to be searched, and to obtain the target search result. The second event library includes the second feature vector extracted by the multimodal large model algorithm.
9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the image search method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image search method as described in any one of claims 1 to 7.