Knowledge base retrieval method and device, electronic equipment and storage medium
Through the combination of the target visual model and the cross-modal mapping module, the knowledge base search is used to use image and text features to perform knowledge base search, which solves the problem of insufficient accuracy of knowledge base search under low-quality image conditions, and realizes efficient multimodal information utilization and rapid identification of new categories of objects.
Patent Information
- Application Number
- CN202510364318.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, when the image quality is low, the accuracy of knowledge base retrieval is low, especially under the shooting angle with large inclination, rotation or extreme lighting conditions, it is difficult to accurately identify the category of the target object.
The target visual model and the target cross-modal mapping module are adopted to search through visual feature extraction and text feature mapping, combined with the multimodal knowledge base, and the multimodal information of images and text is used to improve the retrieval accuracy.
By fully utilizing the pixel and semantic information of the image, the accuracy and efficiency of knowledge base retrieval can be improved under low-quality image conditions, and adapted to the rapid identification and detection of new categories of objects.
Smart Images

Figure CN120492659A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a knowledge base retrieval method and device, an electronic device, and a storage medium. Background Art
[0002] Related technologies use images for knowledge base retrieval to determine the object category of target objects in the image. When the camera's shooting angle is significantly tilted or rotated relative to a reference angle, or when the camera is exposed to extreme lighting conditions, the quality of the image captured by the camera is low, resulting in lower accuracy in knowledge base retrieval. Therefore, improving the accuracy of knowledge base retrieval has become a pressing issue. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to propose a knowledge base retrieval method and device, an electronic device and a storage medium, aiming to improve the accuracy of knowledge base retrieval.
[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a knowledge base retrieval method, the method comprising:
[0005] Acquire a target retrieval image; the target retrieval image includes a target object;
[0006] Inputting the target retrieval image into a target retrieval model, wherein the target retrieval model includes a target visual macro model and a target cross-modal mapping module;
[0007] Extracting visual features from the target retrieval image using the target visual macro model to obtain target retrieval visual features;
[0008] Mapping the target retrieval visual features into text features through the target cross-modal mapping module to obtain target retrieval text features;
[0009] Searching a pre-built multimodal knowledge base according to the target retrieval visual features and the target retrieval text features to obtain an initial object category of the target object and a retrieval score of the initial object category;
[0010] A target object category of the target object is determined according to the initial object category and the retrieval score.
[0011] In some embodiments, before searching a pre-built multimodal knowledge base based on the target retrieval visual features and the target retrieval text features to obtain a retrieval score, the knowledge base retrieval method further includes:
[0012] Constructing the multimodal knowledge base includes:
[0013] Acquire a first sample image and a sample object category of the first sample image;
[0014] Inputting the sample object category into a diffusion generation model to generate an image to obtain a second sample image;
[0015] Inputting the first sample image and the second sample image into a text description model for text generation to obtain a first sample text description of the first sample image and a second sample text description of the second sample image;
[0016] The multimodal knowledge base is constructed according to the first sample image, the first sample text description, the second sample image and the second sample text description.
[0017] In some embodiments, the target retrieval model is trained according to the following steps:
[0018] Acquire a third sample image and a third sample text description of the third sample image;
[0019] Inputting the third sample image and the third sample text description into an initial retrieval model, the initial retrieval model comprising an initial visual large model, an initial language large model, an initial cross-modal mapping module, and an initial cross-modal prediction module; the target visual large model is trained based on the initial visual large model; and the target cross-modal mapping module is trained based on the initial cross-modal mapping module;
[0020] Performing visual feature extraction on the third sample image using the initial large visual model to obtain sample visual features;
[0021] Performing text feature extraction on the third sample text description using the initial language large model to obtain first sample text features;
[0022] Performing feature mapping on the sample visual features by the initial cross-modal mapping module to obtain second sample text features;
[0023] Performing text generation on the sample visual features by the initial cross-modal prediction module to obtain a fourth sample text description;
[0024] determining first target loss data according to the third sample text description, the first sample text feature, the second sample text feature, and the fourth sample text description;
[0025] The model parameters of the initial retrieval model are adjusted according to the first target loss data to obtain the target retrieval model.
[0026] In some embodiments, determining first target loss data based on the third sample text description, the first sample text feature, the second sample text feature, and the fourth sample text description includes:
[0027] Calculate a first loss based on the third sample text description and the fourth sample text description to obtain a first sub-loss;
[0028] Calculate a second loss based on the first sample text feature and the second sample text feature to obtain a second sub-loss;
[0029] Performing text feature extraction on the fourth sample text description using the initial language large model to obtain predicted text features;
[0030] Calculate a third loss based on the first sample text feature and the predicted text feature to obtain a third sub-loss;
[0031] The first sub-loss, the second sub-loss, and the third sub-loss are summed to obtain the first target loss data.
[0032] In some embodiments, the target visual large model includes multiple network layers. Before extracting visual features from the target retrieval image using the target visual large model to obtain target retrieval visual features, the knowledge base retrieval method further includes:
[0033] Adding a low-rank adaptation layer after each of the network layers to obtain a new target visual model;
[0034] Acquire a fourth sample image and an object position label of the fourth sample image;
[0035] Inputting the fourth sample image into a new target visual model to obtain a predicted object position;
[0036] Performing loss calculation based on the object position label and the predicted object position to obtain second target loss data;
[0037] Parameters of the low-rank adaptation layer are adjusted according to the second target loss data.
[0038] In some embodiments, adjusting parameters of the low-rank adaptation layer according to the second target loss data includes:
[0039] Calculate the absolute value of the gradient of the low-rank adaptation layer according to the second target loss data;
[0040] Classifying the low-rank adaptive layer according to the absolute value of the gradient to obtain a first low-rank layer and a second low-rank layer; the absolute value of the gradient of the first low-rank layer is greater than the absolute value of the gradient of the second low-rank layer;
[0041] Performing a rank halving operation on the second low-rank layer to obtain a target low-rank layer;
[0042] Parameters of the first low-rank layer and the target low-rank layer are adjusted according to the second target loss data.
[0043] In some embodiments, the target retrieval model further includes a target language large model, and determining the target object category of the target object based on the initial object category and the retrieval score includes:
[0044] If the retrieval score is greater than a first preset score threshold, taking the initial object category as the target object category;
[0045] If the retrieval score is greater than or equal to the second preset score threshold, and the retrieval score is less than or equal to the first preset score threshold, the initial object category is eliminated; the target retrieval image is input into the text description model for text generation to obtain a target text description; and the target text description is predicted by the target language large model to obtain the target object category.
[0046] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a knowledge base retrieval device, the device comprising:
[0047] An acquisition module is used to acquire a target retrieval image; the target retrieval image includes a target object;
[0048] An input module, configured to input the target retrieval image into a target retrieval model, wherein the target retrieval model includes a target visual macro model and a target cross-modal mapping module;
[0049] A feature extraction module is used to extract visual features of the target retrieval image using the target visual large model to obtain target retrieval visual features;
[0050] a feature mapping module, configured to map the target retrieval visual features into text features through the target cross-modal mapping module to obtain target retrieval text features;
[0051] a retrieval module, configured to search a pre-built multimodal knowledge base based on the target retrieval visual features and the target retrieval text features to obtain an initial object category of the target object and a retrieval score of the initial object category;
[0052] A determination module is configured to determine a target object category of the target object according to the initial object category and the retrieval score.
[0053] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the knowledge base retrieval method of the above-mentioned first aspect is implemented.
[0054] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the knowledge base retrieval method of the above-mentioned first aspect.
[0055] The knowledge base retrieval method, knowledge base retrieval device, electronic device and computer-readable storage medium proposed in this application obtain a target retrieval image so as to use the target retrieval image to perform knowledge base retrieval. The target retrieval image is input into a target retrieval model, which includes a target visual large model and a target cross-modal mapping module. The target visual large model is used to extract visual features of the target retrieval image to fully extract the pixel information of the image, identify the complex patterns and structures in the image, and obtain the target retrieval visual features. In order to make full use of the semantic information contained in the image, the target retrieval visual features are mapped to text features through the target cross-modal mapping module to obtain target retrieval text features. In order to improve the accuracy of knowledge base retrieval, a pre-constructed multimodal knowledge base is retrieved based on the target retrieval visual features and the target retrieval text features, so as to use the multimodal information based on images and texts to perform knowledge base retrieval, reduce the impact of low-quality images on retrieval performance, and obtain the initial object category of the target object and the retrieval score of the initial object category. In order to further improve the accuracy of knowledge base retrieval, the target object category of the target object is determined based on the initial object category and the retrieval score. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a flow chart of the knowledge base retrieval method provided in an embodiment of the present application;
[0057] Figure 2 is another flow chart of the knowledge base retrieval method provided in an embodiment of the present application;
[0058] Figure 3 yes Figure 2 Flowchart of step S250 in FIG.
[0059] Figure 4 is another flow chart of the knowledge base retrieval method provided in an embodiment of the present application;
[0060] Figure 5 This is a flowchart of the training process of the target retrieval model provided in the embodiment of the present application;
[0061] Figure 6 yes Figure 5Flowchart of step S570 in FIG.
[0062] Figure 7 yes Figure 1 Flowchart of step S160 in FIG.
[0063] Figure 8 This is a schematic diagram of the structure of the knowledge base retrieval device provided in an embodiment of the present application;
[0064] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0066] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0068] Related technologies use visual representation models, such as large visual models, to extract and compare features between images and knowledge base samples to determine the object category of the target object in the image. However, this object recognition method only utilizes the pixel information of the image and fails to fully exploit the semantic information contained in the image. Therefore, when the shooting angle is significantly tilted or rotated relative to the reference angle, or under extreme lighting conditions, object recognition performance is poor, resulting in low knowledge base retrieval accuracy. Therefore, improving the accuracy of knowledge base retrieval has become a pressing issue.
[0069] Based on this, the embodiments of the present application provide a knowledge base retrieval method, a knowledge base retrieval device, an electronic device and a computer-readable storage medium, aiming to improve the accuracy of knowledge base retrieval.
[0070] The knowledge base retrieval method, knowledge base retrieval device, electronic device and computer-readable storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the knowledge base retrieval method in the embodiments of the present application is described.
[0071] The knowledge base retrieval method provided in the embodiment of the present application relates to the field of data processing technology. The knowledge base retrieval method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the knowledge base retrieval method, etc., but is not limited to the above forms.
[0072] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0073] Figure 1 This is an optional flowchart of a knowledge base retrieval method provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S110 to S160.
[0074] Step S110, obtaining a target retrieval image; the target retrieval image includes a target object;
[0075] Step S120: input the target retrieval image into a target retrieval model, which includes a target visual macro model and a target cross-modal mapping module;
[0076] Step S130, extracting visual features of the target retrieval image using the target visual large model to obtain target retrieval visual features;
[0077] Step S140, mapping the target retrieval visual features into text features through a target cross-modal mapping module to obtain target retrieval text features;
[0078] Step S150 , searching a pre-built multimodal knowledge base based on the target retrieval visual features and the target retrieval text features to obtain an initial object category of the target object and a retrieval score of the initial object category;
[0079] Step S160 : determining the target object category of the target object according to the initial object category and the retrieval score.
[0080] In step S110 of some embodiments, the target retrieval image is the image to be detected that is input into the retrieval system. The target retrieval image may also be an image of a candidate area detected using a target detection model. The target retrieval image may be adjusted according to user needs. In medical image analysis, the target retrieval image may be a medical image. In a pedestrian recognition scenario, the target retrieval image may be a street view image. The target retrieval image includes a target object, which is a person, object, or other object that appears in the target retrieval image. The target object varies with the image content of the target retrieval image. If the target retrieval image is a pathological section image, the target object may be a pathological tissue structure. If the target retrieval image is a street view image, the target object may be a vehicle, a pedestrian, a traffic sign, or the like.
[0081] In step S120 of some embodiments, the target retrieval image is input into a target retrieval model, which is used to detect the object category of the target object. The target retrieval model includes a target visual macromodel and a target cross-modal mapping module. The target visual macromodel is used to extract pixel information from the target retrieval image, and the target cross-modal mapping module is used to extract semantic information from the target retrieval image, thereby fully utilizing multimodal information to improve the accuracy of object category detection.
[0082] In step S130 of some embodiments, visual features of the target retrieval image are extracted through the target visual big model to obtain target retrieval visual features. The target visual big model is a pre-trained visual big model, such as Vi T, convolutional neural network, etc. The target retrieval visual features are image features of the target retrieval image, such as color features, texture features, pixel space distribution features, object position features, etc.
[0083] In order to improve the accuracy of target detection and enable the target vision large model to detect new categories of objects, it is necessary to update the model parameters of the target vision large model. The target vision large model usually contains a complex network structure and a large number of model parameters. If all the model parameters of the target vision large model are updated, it will consume a lot of time and computing resource costs. In order to improve the efficiency of target detection, the embodiment of the present application uses LoRA technology to fine-tune the target vision large model.
[0084] See also Figure 2 In some embodiments, the target visual model includes multiple network layers. Before step S130, the knowledge base retrieval method may include but is not limited to steps S210 to S250:
[0085] Step S210, adding a low-rank adaptation layer after each network layer to obtain a new target visual model;
[0086] Step S220, obtaining a fourth sample image and an object position label of the fourth sample image;
[0087] Step S230 , inputting the fourth sample image into the new target visual model to obtain a predicted object position;
[0088] Step S240 , performing loss calculation based on the object position label and the predicted object position to obtain second target loss data;
[0089] Step S250: Adjust the parameters of the low-rank adaptation layer according to the second target loss data.
[0090] In step S210 of some embodiments, as the number of new category objects increases, the volume of the multimodal knowledge base will grow linearly, which will significantly reduce the efficiency of knowledge base retrieval, and thus cause the performance and efficiency of the entire target retrieval model to decay over time. In order to improve the recognition ability and recognition efficiency of the target retrieval model for new image samples, when the number of new category samples in the multimodal knowledge base exceeds a preset number threshold or the fine-tuning time interval between two adjacent target visual large models exceeds a preset interval threshold, it is necessary to fine-tune the LoRA weight parameters of the target visual large model. Specifically, before performing LoRA fine-tuning, a low-rank adaptation layer (LoRA) is added after each network layer to obtain a new target visual large model. The LoRA weight parameters are the weight parameters of the low-rank adaptation layer. It should be noted that the ranks of the low-rank adaptation layers are equal initially.
[0091] In step S220 of some embodiments, a fourth sample image and an object position label for the fourth sample image are obtained from the multimodal knowledge base. The fourth sample image is a sample image used to fine-tune the target large visual model. The multimodal knowledge base is a sample image of a new category generated by a diffusion generative model or a real sample image of a new category captured by a photographic device. The fourth sample image contains a target object, and the object position label is used to characterize the position of the target object in the fourth sample image. Targeted fine-tuning of the target large visual model using the new sample allows the target large visual model to absorb knowledge from the knowledge base into the model, thereby improving the model's overall detection capability for new samples.
[0092] In step S230 of some embodiments, the fourth sample image is input into the new target visual model for position prediction to obtain a predicted object position, which is the position of the target object in the fourth sample image predicted by the new target visual model.
[0093] In step S240 of some embodiments, based on the L2 loss function (mean square error loss function), a loss calculation is performed on the object position label and the predicted object position to obtain second target loss data. The second target loss data is the loss value output by the L2 loss function.
[0094] In step S250 of some embodiments, the model parameters of the target visual large model network layer are kept unchanged, the second target loss data is minimized, and the parameters of the low-rank adaptation layer are adjusted.
[0095] In the above steps S210 to S250, the target visual large model is fine-tuned through the new samples, so that the target visual large model has the ability to recognize the new samples, thereby improving the accuracy of knowledge base retrieval.
[0096] Some low-rank adaptation layers require higher ranks, some require lower ranks, and some network layers do not even require low-rank adaptation layers. If each network layer of the target vision model is connected to a low-rank adaptation layer with the same rank, the inherent differences between the low-rank adaptation layers will be ignored. In order to further improve the accuracy and efficiency of target detection, the LoRA weight parameters need to be redistributed.
[0097] See also Figure 3 In some embodiments, step S250 may include but is not limited to steps S310 to S320:
[0098] Step S310, calculating the absolute value of the gradient of the low-rank adaptation layer according to the second target loss data;
[0099] Step S320, classifying the low-rank adaptive layer according to the absolute value of the gradient to obtain a first low-rank layer and a second low-rank layer; the absolute value of the gradient of the first low-rank layer is greater than the absolute value of the gradient of the second low-rank layer;
[0100] Step S330, performing a rank halving operation on the second low-rank layer to obtain a target low-rank layer;
[0101] Step S340: Adjust parameters of the first low-rank layer and the target low-rank layer according to the second target loss data.
[0102] In step S310 of some embodiments, the embodiments of the present application use a gradient-based LoRA clipping method to adaptively allocate LoRA weights, calculate the derivative of the second target loss data with respect to the LoRA weight parameter of each low-rank adaptation layer, and obtain the absolute value of the derivative to obtain the absolute value of the gradient of each LoRA weight parameter of each low-rank adaptation layer. For each low-rank adaptation layer, the absolute value of the gradient of each LoRA weight parameter in the low-rank adaptation layer is added to obtain the absolute value of the gradient of the low-rank adaptation layer. The absolute value of the gradient can be used as an importance score to reflect the importance of the low-rank adaptation layer. The larger the absolute value of the gradient, the more important the low-rank adaptation layer.
[0103] In step S320 of some embodiments, all low-rank adaptive layers are sorted from largest to smallest according to the absolute value of the gradient, i.e., the importance score, and the first first preset number of low-rank adaptive layers after the sorting are all used as the first low-rank layer, and the second second preset number of low-rank adaptive layers after the sorting are all used as the second low-rank layer. The sum of the first preset number and the second preset number equals the total number of low-rank adaptive layers. The first preset number and the second preset number can be set according to actual conditions, for example, the first preset number and the second preset number are both 50% of the total number.
[0104] In step S330 of some embodiments, the rank of the LoRA weight parameter in the second low-rank layer is reduced to half of the original rank to obtain a target low-rank layer.
[0105] In step S340 of some embodiments, the model parameters of the target visual large model network layer are kept unchanged, and the LoRA weight parameters of the first low-rank layer and the LoRA weight parameters of the target low-rank layer are adjusted by minimizing the second target loss data.
[0106] It should be noted that the LoRA weight parameters can be pruned periodically, that is, the above steps S310 to S340 are performed in each training cycle (epoch). Each training cycle is performed based on the LoRA weight parameters after the end of the previous training cycle until the loss function converges. After periodically optimizing the large visual model, the sample images and object location labels of the new categories generated by the diffusion generative model that have been used to train the target large visual model can be deleted from the multimodal knowledge base to pruned the multimodal knowledge base and improve the efficiency of knowledge base retrieval.
[0107] Through the above steps S310 to S340, the LoRA weight parameters can be adaptively allocated, reflecting the inherent differences between the low-rank adaptation layers, enabling the target large visual model to detect new objects, thereby improving the model performance of the target large visual model.
[0108] In step S140 of some embodiments, the target cross-modal mapping module is used to map image features to text features. The target cross-modal mapping module maps the target search visual features to text features to obtain target search text features. The target search text features are used to represent the semantic information of the target search image, such as object category, image style and emotion, object behavior, and object attributes (such as color and shape).
[0109] In step S150 of some embodiments, object detection has always been a key topic in computer vision and related applications. With the development and maturity of deep learning technology, object detection can be effectively solved given a large number of labeled samples. For example, given a large number of labeled vehicle image samples, vehicle detection can be effectively solved by training a deep learning model. However, in real applications, some categories or image samples not seen during training often appear. For example, a new brand of car appears, and the vehicle needs to be detected and its category identified. In this case, only a small number of samples are available, which requires the object detection algorithm to have the ability to learn from a small number of samples. In related art, detection methods that adapt to new categories are divided into two types. The first type involves model training and adaptation based on a large number of training samples. This type of method requires accumulating a large number of samples of the new category and using these samples as training samples to optimize the existing model, enabling the existing model to detect and identify objects of the new category. This type of method uses a large number of samples for model training, resulting in stronger model performance and generalization, enabling detection of new categories of objects from different angles and environments. However, this type of method requires the accumulation of a large number of samples for a long period of training, has a long adaptation cycle, and has low object detection efficiency. It cannot meet the needs of some highly dynamic scenes, such as scenes where new categories appear frequently and have a short adaptation cycle. The second type is a method based on few-sample retrieval. This type of method first detects the area where objects may exist through a general detection model, and uses this area as a candidate area. Then, a comparison model is used to compare the candidate area in the image with a small number of samples of new objects. If the comparison score reaches a certain threshold, the object in the candidate area is considered to be an object of a new category. This type of method does not require training and adaptation, but when the objects that appear and the provided samples have some differences in posture, angle, lighting and other conditions, missed detections will occur, resulting in poor object detection performance.
[0110] In order to resolve the contradiction between adaptation efficiency and adaptation performance, this application utilizes the generalization ability of the large model and the ability of the large model to use an external knowledge base to adapt new categories of objects quickly and with high precision. The pre-built multimodal knowledge base includes image-text pairs, the text is a text description of the image content, and the image-text pairs have object category labels. The target retrieval model also includes a target language large model. For each image-text pair in the multimodal knowledge base, the image is feature extracted by the target visual large model to obtain reference visual features, and the text is feature extracted by the target language large model to obtain reference text features. By comparing the target retrieval visual features and each reference visual feature, the target retrieval text features and each reference text feature, the first similarity between the target retrieval visual features and each reference visual feature can be calculated using metrics such as Euclidean distance, Mahalanobis distance, cosine similarity, and the second similarity between the target retrieval text features and each reference text feature, and the average of the first similarity and the second similarity is taken as the target similarity. If the first similarity is represented by x1 and the second similarity is represented by x2, then the target similarity is (x1+x2) / 2, where / represents a division operation.
[0111] The object category labels corresponding to the reference visual features and reference text features with the highest target similarity are selected as the initial object category of the target object, and the highest target similarity is used as the retrieval score.
[0112] The embodiment of the present application utilizes a small number of samples provided for sample expansion to construct a high-quality, high-rich multimodal knowledge base. The specific process of constructing the multimodal knowledge base is as follows.
[0113] See also Figure 4 In some embodiments, before step S150, the knowledge base retrieval method includes constructing a multimodal knowledge base. The process of constructing the multimodal knowledge base may include but is not limited to steps S410 to S440:
[0114] Step S410, obtaining a first sample image and a sample object category of the first sample image;
[0115] Step S420, inputting the sample object category into the diffusion generation model to generate an image to obtain a second sample image;
[0116] Step S430: Input the first sample image and the second sample image into a text description model to generate text, thereby obtaining a first sample text description of the first sample image and a second sample text description of the second sample image;
[0117] Step S440 : constructing a multimodal knowledge base based on the first sample image, the first sample text description, the second sample image, and the second sample text description.
[0118] In step S410 of some embodiments, the first sample image is a sample image used to construct a knowledge base, the first sample image is an image of a new category object, namely a target object, and the sample object category is the category text name of the new category object, such as brand A sports car.
[0119] In step S420 of some embodiments, when a new object category first appears, there are often only a small number of samples, insufficient to construct a sufficient knowledge base. Therefore, it is necessary to expand the image samples using a diffusion generative model to simulate more image samples of the new category. The diffusion generative model is a Vincent graph diffusion model, which can be a VAE, UNET, CLIP, or other model. The sample object category is input into the diffusion generative model for Vincent graph generation, resulting in a second sample image. The first sample image is a real sample image, and the second sample image is a generated sample image.
[0120] The training process for the diffusion generative model is as follows: a sample image, a preset object category label for the sample image, and random noise are obtained. The preset object category label and random noise are input into the initial diffusion model to generate an image, thereby obtaining a predicted image with the preset object category label. A loss value is calculated based on the sample image and the predicted image. The model parameters of the initial diffusion model are adjusted to minimize the loss value, thus obtaining the diffusion generative model.
[0121] In step S430 of some embodiments, to obtain an image-text pair, the first sample image and the second sample image are input into a text description model for text generation, thereby obtaining a first sample text description and a second sample text description. The text description model has a GPT structure, where the first sample text description is a description of the image content of the first sample image, and the second sample text description is a description of the image content of the second sample image.
[0122] The training process for the text description model is as follows: Sample images are manually annotated with text descriptions, such as object color, lighting, and background, to obtain text description labels. The sample images can be of any type, not limited to new categories. The sample images can be RGB images. The text description labels describe the image content of the sample images. For example, the text description label could be "a slowly setting sun in the background, with yellow sunlight shining in from the left side of the image." The sample images are input into the initial description model for text generation, resulting in a predicted text description. The loss between the text description labels and the predicted text description is calculated. The model parameters of the initial description model are adjusted to minimize the loss, resulting in a text description model.
[0123] In step S440 of some embodiments, image-text pairs are constructed based on the first sample image, the first sample text description, the second sample image, and the second sample text description to obtain a multimodal knowledge base.
[0124] Through the above steps S410 to S440, a multimodal knowledge base with sufficient samples can be obtained to search the multimodal knowledge base, thereby improving the accuracy of knowledge base retrieval.
[0125] See also Figure 5 In some embodiments, the training process of the target retrieval model may include but is not limited to steps S510 to S580:
[0126] Step S510, obtaining a third sample image and a third sample text description of the third sample image;
[0127] Step S520: Input the third sample image and the third sample text description into an initial retrieval model, where the initial retrieval model includes an initial visual large model, an initial language large model, an initial cross-modal mapping module, and an initial cross-modal prediction module; the target visual large model is trained based on the initial visual large model; and the target cross-modal mapping module is trained based on the initial cross-modal mapping module.
[0128] Step S530, extracting visual features from the third sample image using the initial large visual model to obtain sample visual features;
[0129] Step S540, extracting text features from the third sample text description using the initial language large model to obtain first sample text features;
[0130] Step S550: performing feature mapping on the sample visual features through the initial cross-modal mapping module to obtain second sample text features;
[0131] Step S560 , performing text generation on the sample visual features through the initial cross-modal prediction module to obtain a fourth sample text description;
[0132] Step S570, determining first target loss data based on the third sample text description, the first sample text feature, the second sample text feature, and the fourth sample text description;
[0133] Step S580: Adjust the model parameters of the initial retrieval model according to the first target loss data to obtain a target retrieval model.
[0134] In step S510 of some embodiments, the third sample image is a sample image used to train the retrieval model. The third sample image can be a real sample image or a sample image generated by a diffusion generative model. The first sample image, the third sample image, and the fourth sample image can be the same or different. The third sample text description is a description of the image content of the third sample image. The third sample text description can be manually annotated or generated using a text description model.
[0135] In step S520 of some embodiments, the third sample image and the third sample text description are input as training samples into the initial retrieval model for model training. The initial retrieval model includes an initial visual large model, an initial language large model, an initial cross-modal mapping module, and an initial cross-modal prediction module. The initial visual large model is used to extract image features, the initial language large model is used to extract text features, the initial cross-modal mapping module is used to map image features to text features, and the initial cross-modal prediction module is used to generate text descriptions based on image features. The target visual large model is obtained by training the initial visual large model. When the model fine-tuning step of the target visual large model is not performed, the model parameters of the target visual large model are the same as those of the initial visual large model, that is, the model parameters of the visual large model remain unchanged during the retrieval model training phase. The model parameters of the target language large model are the same as those of the initial language large model, that is, the model parameters of the language large model remain unchanged during the retrieval model training phase. The target cross-modal mapping module is obtained by training the initial cross-modal mapping module.
[0136] In step S530 of some embodiments, visual features are extracted from the third sample image using the initial visual macro model to obtain sample visual features, which are image features of the third sample image.
[0137] In step S540 of some embodiments, the initial language model is used to extract text features from the third sample text description to obtain first sample text features. The first sample text features are text features of the third sample text description (original sample) and are used to represent text semantic information of the third sample image.
[0138] In step S550 of some embodiments, the sample visual features are mapped from image features to text features by an initial cross-modal mapping module to obtain second sample text features.
[0139] In step S560 of some embodiments, the initial cross-modal prediction module generates text from the sample visual features, converts the image features into a text description, and obtains a fourth sample text description. The fourth sample text description is a predicted text description.
[0140] In step S570 of some embodiments, the loss function provides a quantitative indicator of model performance, which can be used to guide the update of model parameters. To optimize the initial retrieval model, it is necessary to determine the loss value of the loss function. The loss is calculated based on the third sample text description, the first sample text feature, the second sample text feature, and the fourth sample text description, and the calculated loss value is used as the first target loss data.
[0141] In step S580 of some embodiments, the first target loss data is minimized, and the model parameters of the initial cross-modal mapping module and the initial cross-modal prediction module are adjusted to obtain a target visual large model, a target language large model, a target cross-modal mapping module, and an adjusted cross-modal prediction module, and the adjusted four models are used as target retrieval models.
[0142] In the above steps S510 to S580, based on the visual big model, the cross-modal mapping module is trained based on the multimodal information of images and text descriptions, so that the semantic information contained in the image can be fully utilized in the retrieval stage, thereby improving the accuracy of knowledge base retrieval.
[0143] See also Figure 6 In some embodiments, step S570 may include but is not limited to steps S610 to S650:
[0144] Step S610, calculating a first loss based on the third sample text description and the fourth sample text description to obtain a first sub-loss;
[0145] Step S620, calculating a second loss based on the first sample text feature and the second sample text feature to obtain a second sub-loss;
[0146] Step S630, extracting text features from the fourth sample text description using the initial language large model to obtain predicted text features;
[0147] Step S640: Calculate a third loss based on the first sample text feature and the predicted text feature to obtain a third sub-loss;
[0148] Step S650 , summing the first sub-loss, the second sub-loss, and the third sub-loss to obtain first target loss data.
[0149] In step S610 of some embodiments, the predicted text description and the original text description should be consistent. A first loss may be calculated based on the third sample text description and the fourth sample text description based on a cross-entropy loss function to obtain a first sub-loss. The first sub-loss is a cross-modal text description loss.
[0150] In step S620 of some embodiments, the predicted cross-modal features and the text features of the original sample should be consistent, and a second loss may be calculated based on the cosine similarity loss function for the first sample text features and the second sample text features to obtain a second sub-loss. The second sub-loss is the cross-modal feature loss.
[0151] In step S630 of some embodiments, the text features corresponding to the predicted text description should be consistent with the text features of the original sample. The text features of the fourth sample text description are extracted using the initial language model to obtain the predicted text features.
[0152] In step S640 of some embodiments, a third loss may be calculated based on the cosine similarity loss function for the first sample text feature and the predicted text feature to obtain a third sub-loss, which is a text feature loss.
[0153] In step S650 of some embodiments, the first sub-loss, the second sub-loss, and the third sub-loss are added or weighted to obtain first target loss data.
[0154] Through the above steps S610 to S650, the first target loss data can be obtained to train the retrieval model based on the first target loss data.
[0155] See also Figure 7 In some embodiments, step S160 may include but is not limited to step S710 or step S720:
[0156] Step S710: If the search score is greater than a first preset score threshold, the initial object category is used as the target object category;
[0157] Step S720: If the retrieval score is greater than or equal to the second preset score threshold, and the retrieval score is less than or equal to the first preset score threshold, the initial object category is eliminated; the target retrieval image is input into the text description model for text generation to obtain the target text description; the target text description is predicted by the target language model to obtain the target object category.
[0158] In step S710 of some embodiments, if the retrieval score is greater than a first preset score threshold, indicating that the target retrieval image has a high similarity with the image-text pair in the multimodal knowledge base, the initial object category is used as the target object category.
[0159] In step S720 of some embodiments, the second preset score threshold is less than the first preset score threshold. The second preset score threshold is greater than or equal to 0. If the retrieval score is greater than or equal to the second preset score threshold and the retrieval score is less than or equal to the first preset score threshold, it indicates that the target retrieval image has a low similarity with the image-text pair in the multimodal knowledge base and the target retrieval image is a new category sample, and the initial object category is eliminated. The target retrieval image is input into a text description model for text generation, or the target retrieval visual features are input into a trained initial cross-modal prediction module for text generation to obtain a target text description, which is a description of the image content of the target retrieval image. The target text description is predicted by the target language macro model to obtain the target object category. Alternatively, the target retrieval image is input into a universal detection model to obtain the target object category. If the retrieval score is less than the second preset score threshold, the target retrieval image is reacquired until the number of acquisitions of the target retrieval image exceeds a preset number threshold, and the target object category is obtained using the target language macro model, the trained initial cross-modal prediction module, or the universal detection model.
[0160] It should be noted that during object category detection, the knowledge base can be dynamically updated to ensure its richness and accuracy. Specifically, during detection, if a new category of object is detected and the detection confidence exceeds a certain threshold, the image of the new category object is added to the multimodal knowledge base and a detailed text description is generated. During the detection process, if a new sample is obtained, a generated sample of that category in the knowledge base is eliminated to ensure the authenticity of the samples in the knowledge base, thereby improving retrieval efficiency and accuracy.
[0161] In the above steps S710 to S720, the object categories are screened again using the retrieval scores, thereby avoiding incorrect initial object categories, improving the model's ability to recognize samples of new categories, and further improving the accuracy of knowledge base retrieval.
[0162] The embodiment of the present application first expands the provided samples based on the diffusion generation model and the text generation model to obtain a multimodal knowledge base with image-text pairs, and dynamically updates the knowledge base during the detection process to ensure the richness and accuracy of the knowledge base. Secondly, the multimodal feature mapping model is trained, and the multimodal knowledge base is retrieved during the detection process. Then, as the number of new categories of objects increases to a certain extent, the large model is efficiently fine-tuned so that it itself has the ability to detect new categories of objects, thereby improving the overall ability of the large model. Finally, the model obtained by the above training is deployed and applied. For example, the large visual model, cross-modal mapping module and knowledge base can be deployed to network devices (cloud servers, terminal devices, etc.), thereby improving the efficiency and accuracy of knowledge base retrieval.
[0163] See also Figure 8 The present application also provides a knowledge base search device that can implement the above-mentioned knowledge base search method. The knowledge base search device includes:
[0164] The acquisition module 810 is used to acquire a target retrieval image; the target retrieval image includes a target object;
[0165] An input module 820 is used to input the target retrieval image into the target retrieval model, which includes a target visual model and a target cross-modal mapping module;
[0166] The feature extraction module 830 is used to extract visual features of the target retrieval image using the target visual large model to obtain target retrieval visual features;
[0167] A feature mapping module 840 is configured to map the target retrieval visual features into text features through a target cross-modal mapping module to obtain target retrieval text features;
[0168] A retrieval module 850 is configured to search a pre-built multimodal knowledge base based on target retrieval visual features and target retrieval text features to obtain an initial object category of the target object and a retrieval score of the initial object category;
[0169] The determination module 860 is configured to determine the target object category of the target object according to the initial object category and the retrieval score.
[0170] The specific implementation of the knowledge base retrieval device is basically the same as the specific embodiment of the above-mentioned knowledge base retrieval method, and will not be repeated here.
[0171] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned knowledge base search method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.
[0172] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0173] The processor 910 may be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0174] The memory 920 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 920 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 920 and are called by the processor 910 to execute the knowledge base retrieval method of the embodiments of this application.
[0175] Input / output interface 930, used to implement information input and output;
[0176] Communication interface 940, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.);
[0177] bus 950 , which transmits information between various components of the device (e.g., processor 910 , memory 920 , input / output interface 930 , and communication interface 940 );
[0178] The processor 910 , the memory 920 , the input / output interface 930 , and the communication interface 940 are connected to each other in communication within the device via a bus 950 .
[0179] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned knowledge base retrieval method is implemented.
[0180] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0181] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0182] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0183] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0184] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0185] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0186] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0187] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0188] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0189] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.
[0191] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A knowledge base retrieval method, characterized in that: The method comprises: Acquire a target retrieval image; the target retrieval image includes a target object; Inputting the target retrieval image into a target retrieval model, wherein the target retrieval model includes a target visual macro model and a target cross-modal mapping module; Extracting visual features from the target retrieval image using the target visual macro model to obtain target retrieval visual features; Mapping the target retrieval visual features into text features through the target cross-modal mapping module to obtain target retrieval text features; Searching a pre-built multimodal knowledge base according to the target retrieval visual features and the target retrieval text features to obtain an initial object category of the target object and a retrieval score of the initial object category; A target object category of the target object is determined according to the initial object category and the retrieval score.
2. The knowledge base retrieval method according to claim 1, characterized in that: Before searching the pre-built multimodal knowledge base according to the target retrieval visual features and the target retrieval text features to obtain a retrieval score, the knowledge base retrieval method further includes: Constructing the multimodal knowledge base includes: Acquire a first sample image and a sample object category of the first sample image; Inputting the sample object category into a diffusion generation model to generate an image to obtain a second sample image; Inputting the first sample image and the second sample image into a text description model for text generation to obtain a first sample text description of the first sample image and a second sample text description of the second sample image; The multimodal knowledge base is constructed according to the first sample image, the first sample text description, the second sample image and the second sample text description.
3. The knowledge base retrieval method according to claim 1, characterized in that: The target retrieval model is trained according to the following steps: Acquire a third sample image and a third sample text description of the third sample image; Inputting the third sample image and the third sample text description into an initial retrieval model, the initial retrieval model comprising an initial visual large model, an initial language large model, an initial cross-modal mapping module, and an initial cross-modal prediction module; the target visual large model is trained based on the initial visual large model; and the target cross-modal mapping module is trained based on the initial cross-modal mapping module; Performing visual feature extraction on the third sample image using the initial large visual model to obtain sample visual features; Performing text feature extraction on the third sample text description using the initial language large model to obtain first sample text features; Performing feature mapping on the sample visual features by the initial cross-modal mapping module to obtain second sample text features; Performing text generation on the sample visual features by the initial cross-modal prediction module to obtain a fourth sample text description; determining first target loss data according to the third sample text description, the first sample text feature, the second sample text feature, and the fourth sample text description; The model parameters of the initial retrieval model are adjusted according to the first target loss data to obtain the target retrieval model.
4. The knowledge base retrieval method according to claim 3, characterized in that: The determining of first target loss data according to the third sample text description, the first sample text feature, the second sample text feature, and the fourth sample text description includes: Calculate a first loss based on the third sample text description and the fourth sample text description to obtain a first sub-loss; Calculate a second loss based on the first sample text feature and the second sample text feature to obtain a second sub-loss; Performing text feature extraction on the fourth sample text description using the initial language large model to obtain predicted text features; Calculate a third loss based on the first sample text feature and the predicted text feature to obtain a third sub-loss; The first sub-loss, the second sub-loss, and the third sub-loss are summed to obtain the first target loss data.
5. The knowledge base retrieval method according to any one of claims 1 to 4, characterized in that: The target visual large model includes multiple network layers. Before extracting visual features from the target retrieval image using the target visual large model to obtain target retrieval visual features, the knowledge base retrieval method further includes: Adding a low-rank adaptation layer after each of the network layers to obtain a new target visual model; Acquire a fourth sample image and an object position label of the fourth sample image; Inputting the fourth sample image into a new target visual model to obtain a predicted object position; Performing loss calculation based on the object position label and the predicted object position to obtain second target loss data; Parameters of the low-rank adaptation layer are adjusted according to the second target loss data.
6. The knowledge base retrieval method according to claim 5, characterized in that: The adjusting parameters of the low-rank adaptation layer according to the second target loss data includes: Calculate the absolute value of the gradient of the low-rank adaptation layer according to the second target loss data; Classifying the low-rank adaptive layer according to the absolute value of the gradient to obtain a first low-rank layer and a second low-rank layer; the absolute value of the gradient of the first low-rank layer is greater than the absolute value of the gradient of the second low-rank layer; Performing a rank halving operation on the second low-rank layer to obtain a target low-rank layer; Parameters of the first low-rank layer and the target low-rank layer are adjusted according to the second target loss data.
7. The knowledge base retrieval method according to any one of claims 1 to 4, characterized in that: The target retrieval model further includes a target language large model, and the step of determining the target object category of the target object according to the initial object category and the retrieval score includes: If the retrieval score is greater than a first preset score threshold, taking the initial object category as the target object category; If the retrieval score is greater than or equal to the second preset score threshold, and the retrieval score is less than or equal to the first preset score threshold, the initial object category is eliminated; the target retrieval image is input into the text description model for text generation to obtain a target text description; and the target text description is predicted by the target language large model to obtain the target object category.
8. A knowledge base retrieval device, characterized in that: The device comprises: An acquisition module is used to acquire a target retrieval image; the target retrieval image includes a target object; An input module, configured to input the target retrieval image into a target retrieval model, wherein the target retrieval model includes a target visual macro model and a target cross-modal mapping module; A feature extraction module is used to extract visual features of the target retrieval image using the target visual large model to obtain target retrieval visual features; a feature mapping module, configured to map the target retrieval visual features into text features through the target cross-modal mapping module to obtain target retrieval text features; A retrieval module, configured to search a pre-built multimodal knowledge base based on the target retrieval visual features and the target retrieval text features to obtain an initial object category of the target object and a retrieval score of the initial object category; A determination module is configured to determine a target object category of the target object according to the initial object category and the retrieval score.
9. An electronic device, characterized in that The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the knowledge base retrieval method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the knowledge base retrieval method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Knowledge base retrieval device and method based on sentence pair embedding
CN121388086A