Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

657 results about "Image retrieval" patented technology

An image retrieval system is a computer system for browsing, searching and retrieving images from a large database of digital images. Most traditional and common methods of image retrieval utilize some method of adding metadata such as captioning, keywords, title or descriptions to the images so that retrieval can be performed over the annotation words. Manual image annotation is time-consuming, laborious and expensive; to address this, there has been a large amount of research done on automatic image annotation. Additionally, the increase in social web applications and the semantic web have inspired the development of several web-based image annotation tools.

Multi-dimensional data intelligent retrieval matching method and system for graphic and text features

The invention provides an intelligent retrieval matching method and system for multidimensional data of image-text features, and relates to the technical field of icon image retrieval. Comprising the following steps: extracting image features, text content and semantic features of an icon image by using a convolutional neural network, an image segmentation attention mechanism network, a converter optical character recognition model and a bidirectional semantic understanding model, constructing the extracted features into heterogeneous feature tensors, and performing singular value decomposition to obtain icon feature fingerprint vectors; and constructing a multi-level index based on locality sensitive hashing, realizing rapid retrieval, calculating visual, text and semantic similarities in combination with a deep metric learning model, weighting according to variances and discrimination coefficients of similarity features to obtain a comprehensive similarity score, and outputting a retrieval result with the highest similarity.
Owner:BEIJING YIZHUANG TECHNOLOGY INNOVATION CO LTD

Query evaluation for image retrieval and conditional image generation

Techniques are generally described for query evaluation for image retrieval and image generation. In various examples, a first encoded representation of first natural language input data may be generated. An image retrieval process may be selected from among the image retrieval process and an image generation process based at least in part on the first encoded representation of the first natural language input data. A second natural language encoder may generate a second encoded representation of the first natural language input data. The second encoded representation may be used to determine first image data stored in a first data repository. The first image data may be sent for output on a display of a first computing device.
Owner:AMAZON TECH INC

Enhanced LLM-RAG multi-hop question and answer method based on logic tree reasoning

The invention relates to an enhanced LLM-RAG multi-hop question and answer method based on logic tree reasoning, and belongs to the technical field of new-generation information, and the method comprises the following steps: inputting a multi-hop question and answer question into a computer system; the computer system calls a pre-training large language model LLM, the multi-hop question-answer question is decomposed into a hierarchical logic tree in a recursive mode, and each node of the logic tree comprises a sub-question and a corresponding hypothesis answer; performing image retrieval from a structured knowledge source Wikidata and performing text retrieval from an unstructured knowledge source Wikipedia on the basis of each node sub-question and the hypothesis answer to obtain corresponding evidence; traversing the logic tree, verifying the consistency between the hypothetical answer of each node and the evidence through LLM, if the contradiction exists, reconstructing the corresponding sub-tree, and dynamically correcting the reasoning path; and integrating the verified logic tree node information, and outputting an accurate answer to the multi-hop question and answer question.
Owner:GUIZHOU UNIV +1

Dynamic game difficulty self-adaptive adjustment method and system based on user behavior feedback

The invention provides a dynamic game difficulty self-adaptive adjustment method and system based on user behavior feedback, and relates to the technical field of game design, the method comprises the following steps: carrying out dynamic weighting calculation based on a difficulty adaptation index, generating a dynamic difficulty correction coefficient in combination with user real-time physiological feedback data, and adjusting the dynamic game difficulty according to the dynamic difficulty correction coefficient; the dynamic difficulty correction coefficient comprises a scene complexity adjustment parameter and an interaction response threshold adjustment parameter; inputting the user feedback parameter set into an image feature weight distributor, dynamically adjusting weight distribution of image retrieval feature vectors according to user attention distribution data and operation delay parameters, and generating an optimized scene element retrieval strategy; and performing feature space mapping on the dynamic difficulty correction coefficient and the optimized scene element retrieval strategy, and generating a final game difficulty control instruction set through nonlinear superposition. Game design can be optimized, and game adaptability and flexibility are enhanced.
Owner:LIANYUNGANG FEIYANG NETWORK TECH CO LTD

Atlas-driven intelligent medical image retrieval method and system

The invention discloses a graph-driven intelligent medical image retrieval method and system, and relates to the field of image retrieval, and the method comprises the steps: firstly, coding a text query input by a user through constructing a medical knowledge graph, querying to generate a graph-driven enhanced embedded vector containing an advanced medical concept, and querying a related graph concept set; then, performing preliminary vector similarity retrieval by using the enhanced vector, and quickly screening out Top-K preliminary candidate images; and finally, on the basis of querying a related map concept set, carrying out map concept secondary sorting on the preliminary candidate images, and accurately matching deep medical semantics of the images by calculating a concept coincidence degree. The mechanism of two-stage retrieval and map deep fusion ensures that the retrieval result is not only similar in vision, but also highly related in medical concept, so that the accuracy and intelligent level of medical image retrieval are remarkably improved.
Owner:ZHEJIANG FEITU IMAGING TECH CO LTD

System and Method for Enhancing Generative Artificial Intelligence (AI) Model-Based Document Search with Image Retrieval

A method, computer program product, and computing system for generating a plurality of chunks for a plurality of text portions of a document, wherein the document includes the plurality of text portions and a plurality of images. Each chunk is indexed using a word embedding. Each of the plurality of images is indexed based upon, at least in part, a position of a respective image relative to a corresponding chunk. An image placeholder is generated for each of the plurality of images. A plurality of image-enhanced embeddings is generated by inserting the image placeholder for each of the plurality of images into a respective word embedding for the corresponding chunk. The plurality of image-enhanced embeddings are provided for processing a query using a generative artificial intelligence (AI) model.
Owner:DELL PROD LP

Cross-modal joint contrast learning method and device and electronic equipment

The invention provides a cross-modal joint contrast learning method and device and electronic equipment, and the method comprises the steps: constructing a sample data set which covers a plurality of task types, such as a text retrieval image, an image retrieval text, a text retrieval text, an image retrieval image, an image-text joint retrieval image and an image-text joint retrieval text; and generating prompt words for identifying task types for each task sample, splicing the prompt words with sample data to form task input, and determining modal types of a retrieval object and a retrieval target. A retrieval object and a retrieval target are respectively input into encoders of corresponding modes to extract features, the features are mapped to the same semantic space through a unified projection layer to obtain an embedded vector, and the parameters of the encoders and the projection layer are optimized by utilizing a contrast learning loss function based on the similarity of the retrieval object and the retrieval target, so that multi-task unified training is realized. Various cross-modal retrieval tasks can be supported in a unified semantic space at the same time, and the overall retrieval effect is improved on the premise of ensuring multi-task performance balance.
Owner:SHANGHAI ANXINCHENG NETWORK TECHNOLOGY CO LTD

Cross-view-angle image geographic positioning method based on dynamic threshold value pseudo label self-training learning

The invention discloses a cross-view image geographic positioning method based on dynamic threshold pseudo tag self-training learning, and the method specifically comprises the following steps: introducing a difficult sample feature mining method, dynamically adjusting the loss weight of a sample according to the change of similarity, and building a dynamic difficult sample triple loss model; the method comprises the following steps: dynamically adjusting a confidence threshold value of a sample by adopting an index moving average weighting method, iteratively training and screening an unlabeled sample, namely a pseudo label, establishing a pseudo label self-training mechanism of a dynamic threshold value, mining and utilizing non-paired data, and solving the problem of high manual labeling cost; a reference image most similar to a query image is found through image retrieval, and the offset of a query position is predicted. Experiments on CVUSA and CVACT data sets show that as the distance threshold increases, the accuracy of the cross-view image geographic positioning method based on dynamic threshold pseudo tag self-training learning presents a stable rising trend, and the cross-view image geographic positioning method based on dynamic threshold pseudo tag self-training learning is superior to other methods under the same threshold condition.
Owner:HENAN UNIVERSITY

Video geographic positioning method and device based on iron tower monitoring

The invention discloses a video geographic positioning method and device based on iron tower monitoring, and the method comprises the steps: constructing a monitoring video homonymy point sample library, collecting the homonymy point data of a monitoring video and a remote sensing orthoimage, and storing a reference image; establishing an initial calibration mapping relation, and realizing coordinate transformation of the remote sensing image and the video image through projection transformation; matching a reference image closest to the real-time video frame from a reference image library through image retrieval; dense matching of inclined images is carried out by using a deep learning model, and automatic alignment of feature points is realized; and resolving a video image coordinate to a geographic coordinate through homography transformation and an initial calibration model to complete high-precision positioning. According to the method, the bidirectional conversion function of finding the land by the video and finding the video by the land is supported, the application efficiency of iron tower monitoring in the fields of cultivated land protection and the like is effectively improved, a good positioning effect is still achieved for complex scenes such as plains, mountainous regions and the like, and corresponding technical support is provided for wide application of iron tower video monitoring.
Owner:WUHAN UNIV

Method, device, and product for retrieval

The present disclosure provides a method, a device, and a product for retrieval. The method includes acquiring context information related to an image and determining a representation of the image based on image data and the context information of the image, where the context information includes at least one of environment parameters, user behavior data, time elements, or field metadata. The method further includes encoding the representation as an image vector in a high-dimensional vector space and storing it into an image vector database. When retrieval is performed, a query that includes text information and that is for the image vector database is received, and an image associated with the text information is determined from the image vector database. The method according to the present disclosure can improve accuracy and efficiency for image retrieval.
Owner:DELL PROD LP

Deep hash image retrieval method based on hierarchical multi-scale feature fusion

The invention provides a deep hash image retrieval method based on hierarchical multi-scale feature fusion. The method comprises the following steps: firstly, adopting a Nested Hierarchical Transform as a backbone feature extraction network, and synchronously capturing global context information and local detail features of an image through a nested local self-attention mechanism and a hierarchical feature aggregation module; then, constructing a multi-scale feature fusion module, extracting visual features of different receptive fields in parallel by using a multi-branch unequal-ratio expansion rate dilated convolution kernel, and adaptively weighting important features in combination with attention to generate discriminative multi-scale fusion features; and finally, designing a mixed loss function, optimizing intra-class feature compactness through center similarity loss, and realizing efficient and semantic similarity keeping Hash code generation in combination with a quantization loss constraint Hash code discretization process. According to the method, full expression of image features is realized through a collaborative mechanism of Transform and multi-scale feature fusion, and the precision and efficiency of large-scale image retrieval are improved through an end-to-end deep hash learning framework.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Complex text image retrieval method and system based on multi-modal large model thinking chain

The invention discloses a complex text image retrieval method and system based on a multi-modal large model thinking chain. The method comprises the following steps: S1, carrying out adaptive semantic disassembly; s2, reconstruction and optimization are carried out; s3, constructing a matching probability matrix; elements of the matching probability matrix represent probability scores of matching degrees between the candidate images and the matching texts; combining the matched text with the candidate image, and sending the combined image to a pre-training visual language model pair by pair to obtain a corresponding matching score; then a binary discriminant scoring mechanism is used to convert the matching score into a probability score that the cue word is' YES '; and S4, according to the matching probability matrix, calculating the matching degree between each candidate image and the original proposition, and selecting the candidate image with the highest matching degree as an image retrieval result. The invention aims to extract and optimize semantic features from complex text description by using a self-adaptive local deconstruction and global optimization method and a binary discriminant scoring mechanism guided by cue words, so as to improve the accuracy and generalization ability of image retrieval.
Owner:HANGZHOU DIANZI UNIV

Image retrieval method and device, computer equipment and storage medium

The invention discloses an image retrieval method and device, computer equipment and a storage medium, belongs to the technical field of artificial intelligence, and is provided with a retrieval system applied to an insurance marketing scene image. According to the method and the device, the to-be-processed image is segmented and coded, the content features of the image are combined with the position information to form high-quality image representation, and the corresponding image index information is generated by utilizing the preset image index generator, so that efficient indexing and retrieval of the image are realized. In the encoding process, local features of the image are reserved, spatial position information is fused, discrimination and uniqueness of image indexing are improved, generated image indexing information and an original image are stored in an image retrieval database in an associated mode, it is ensured that the image retrieval process has high matching precision and response speed, and the image retrieval efficiency is improved. The image retrieval accuracy and processing efficiency are effectively improved, and the method is suitable for large-scale image library management and rapid retrieval.
Owner:PING AN TECH (SHENZHEN) CO LTD

Cross-view image retrieval method for unmanned aerial vehicle navigation

A cross-view-angle image retrieval method for unmanned aerial vehicle navigation comprises the following steps: step 1, enhancing an image, step 2, based on a ConvNeXt network for cross-view-angle image retrieval for unmanned aerial vehicle navigation, performing multi-scale feature extraction on the enhanced image, and step 3, performing multi-scale feature extraction on the enhanced image based on a ConvNeXt network for cross-view-angle image retrieval for unmanned aerial vehicle navigation. 3, performing multi-dimensional feature fusion based on an attention mechanism; 4, after multi-dimensional feature fusion, calculating a loss value, and based on a determined weight, obtaining a detection model for image retrieval; through the multi-scale feature convolution and multi-dimensional feature fusion technology, the technical problems of insufficient matching precision, poor retrieval robustness and the like caused by view angle, scale and structure differences between the satellite image and the unmanned aerial vehicle image are solved; the invention further comprises a system, equipment and a medium for implementing the method.
Owner:XI AN JIAOTONG UNIV

Image retrieval method, device and equipment based on multi-modal semantics

The invention discloses an image retrieval method, device and equipment based on multi-modal semantics, which are applied to the technical field of multi-modal retrieval. According to the image retrieval method based on the multi-modal semantics, a retrieval request comprising a reference image and a modified text is obtained; based on the reference image and the modified text, a first global image feature including visual information of the reference image, an object feature including information of the modified object, and a description feature including semantic information of the modified text are extracted, respectively. Therefore, relatively complete visual information and semantic information can be extracted. And integrating the first global image feature, the object feature and the description feature to obtain a retrieval feature. And finally, determining a target image by utilizing the retrieval features, and generating a retrieval result. According to the method, retrieval is carried out by utilizing the retrieval characteristics fusing the visual information and the semantic information, so that the retrieval capability of cross-modal information can be enhanced, the accuracy and the effectiveness of the target image obtained through retrieval are improved, and the retrieval requirements of a user in a multi-modal image retrieval scene are met.
Owner:UNIV OF SCI & TECH OF CHINA

Image searching method and device, electronic equipment and storage medium

The invention provides an image searching method, which comprises the following steps of: when a new label is added to an image in a first image database, acquiring a first semantic feature of the new label; based on the first semantic feature, performing image retrieval in a first image database to obtain a first retrieval result and a confidence coefficient of the first retrieval result; if the confidence coefficient is smaller than the preset confidence coefficient, performing secondary retrieval in the first retrieval result based on the existing tag in the first retrieval result to obtain a second retrieval result; based on the second retrieval result, adding the newly added tag into an image corresponding to the secondary retrieval result to obtain a second image database; and after an image search instruction of the user is obtained, performing image retrieval in the second image database based on the image search instruction. The problems that in the database maintenance process, a new label needs to be added, the workload of adding the label again to the data is increased while the new label is added, and the database is difficult to maintain in an existing method are solved.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD +1

An Adaptive Method and System for Selecting UAV Image Matching Pairs

The present invention provides an adaptive drone image matching pair selection method and system, comprising: S1: acquiring drone images, performing feature extraction on the drone images using the SIFT algorithm, and obtaining a local feature set of the drone images; S2: constructing a vector codebook, performing feature aggregation on the local feature set using the vector codebook, and obtaining a feature matrix set; vectorizing the feature matrix set to obtain a global feature descriptor vector set; S3: performing image retrieval on the global feature descriptor vector set based on a graph index structure to obtain a drone image matching pair set; S4: performing three-dimensional reconstruction on the drone image matching pair set to obtain a three-dimensional reconstruction model. The present invention uses global feature descriptor vectors to replace word frequency calculation based on local features, and provides a graph-based indexing strategy to achieve efficient overlapping image search, making drone image matching pairs more efficient and accurate.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Extracting images and determining their meaning for semantic image retrieval and training a transformer-based multi-modal large language model to generate domain-aware images based on image meanings

The disclosure relates to systems and methods automatically extracting an image and related image components, computationally determining an understanding of the image, and generating mathematical vector embeddings via sentence encoders based on the computationally determined understanding. The mathematical vector embeddings may be used for semantic image retrieval that enables image searching based on a semantic understanding of input images and / or input text. The mathematical vector embeddings may be used for training and executing generative Artificial Intelligence (AI) models to create new content that includes retrieved images and / or generate new images.
Owner:ROHIRRIM INC

Method and system for quickly retrieving and matching inspection images of power distribution network

The invention relates to the technical field of power grid image retrieval, and discloses a power distribution network inspection image rapid retrieval matching method and system, and the method comprises the steps: obtaining a to-be-retrieved inspection image of power distribution network equipment, and extracting the equipment structure features of the inspection image through a hierarchical convolutional network; quantifying a surface texture attenuation index of the power distribution network equipment through fractal dimension based on the equipment structure characteristics; performing mapping relation coupling on the surface texture attenuation index and a space coordinate of an equipment connecting piece to generate a dynamic feature coding sequence containing an equipment structure topological relation; performing time sequence consistency matching on the dynamic feature coding sequence and a pre-constructed reference image library, and aligning an equipment aging track through a dynamic time warping algorithm to generate a similarity sorting result; according to the method, the problems that effective features cannot be extracted during retrieval matching and the retrieval precision is low are solved.
Owner:安徽明生恒卓科技有限公司 +1

Fine-grained clothing image retrieval method and device based on large language model common sense knowledge injection

The application discloses a fine-grained clothing image retrieval method and device for injecting common sense knowledge of a large language model. The method first extracts fine-grained visual features of an input image through an image encoder and optimizes them in combination with a low-rank adapter to improve the representation ability at the image patch level. Then, a pre-trained large language model generates attribute-enhanced common sense knowledge context to enrich the image attribute representation, thereby helping the model understand and reason about unknown attribute information in an open scenario. The application introduces a switchable modal prompt and interpolation mechanism to ensure that the agent embedding can be dynamically supplemented when attributes or text are missing. During the retrieval process, through an attribute-guided cross-modal attention mechanism, fine-grained image content matching is performed based on the relationship between image features and attribute-enhanced context. Through multi-modal feature alignment and optimization, the application improves the accuracy and robustness of clothing image retrieval in an open world scenario.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Cross-modal image-text pedestrian retrieval method based on generative model

The invention relates to a cross-modal retrieval technology, in particular to a cross-modal image-text pedestrian retrieval method based on a generative model, which comprises the following steps: acquiring text description information and a corresponding original pedestrian image; calling a diffusion generation model to generate an intermediate image based on the text description; original image features and text features are extracted through an image encoder and a text encoder respectively, and semantic-fused text feature representation is constructed; calculating the similarity between the image feature representation and the text feature representation to obtain an image-text matching score; introducing and generating an intermediate image as a semantic bridge, and realizing multi-modal fusion among the image, the text and the intermediate image through a cross attention mechanism to obtain fused image and text features; and using the fusion features to train a recognition model, and finally outputting a pedestrian image retrieval result corresponding to the text description. According to the method, the image-text alignment precision can be improved, the robustness of a scene with incomplete text description is enhanced, and the recognition accuracy in a cross-modal retrieval task is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Text-based image retrieval

A method, apparatus, non-transitory computer readable medium, and system for media processing include obtaining a text prompt describing content, generating, using a multi-modal encoder, a text embedding based on the text prompt, and obtaining an image depicting the content based on the text embedding. The multi-modal encoder is trained to encode image descriptions based on a similarity between a caption of a training image and a paraphrase of the caption.
Owner:ADOBE INC

Multi-modal unstructured content association retrieval method

The invention discloses a multi-modal unstructured content association retrieval method, which comprises the following steps of: performing feature extraction on multi-modal data to obtain features of the data in different modals; the multi-modal features are aligned, and multi-modal alignment features are obtained; random masking is carried out on the multi-modal alignment features, the multi-modal alignment features are sent into a cross-modal self-attention model to be fused, and multi-modal fusion feature vectors after masking are obtained; extracting enhanced features of each image; processing the enhanced features through a cross attention network to obtain cosine similarity between different images; performing image feature matching to obtain an image association result; a retrieval text is input into a large language model to obtain text features, through a multi-modal data embedding space, the most similar image is matched by using cosine similarity, and a retrieval result of the image is obtained. According to the method, the accuracy and the stability of multi-modal feature fusion are improved, and the accuracy and the efficiency of image retrieval are improved.
Owner:10TH RES INST OF CETC

Zero-sample multi-modal combined image retrieval method and device based on context awareness

The invention provides a zero-sample multi-modal combined image retrieval method and device based on context awareness, and relates to the technical field of image retrieval, and the method comprises the following steps: obtaining a training sample set; the training sample set is input into a target CLIP model, target image embedding and target text embedding of the training sample set are obtained, and learnable language prompt embedding and visual prompt embedding are introduced into a text encoder and an image encoder of the target CLIP model respectively; training the mapping network model according to the target image embedding, the target text embedding and the target loss function to obtain a target mapping network model; obtaining a target image retrieval model according to the target CLIP model and the target mapping network model, and inputting the reference image and the text description into the target image retrieval model to obtain a target query feature; and outputting a retrieval result according to the similarity between the target query feature and the image embedding of each candidate image. According to the image retrieval method, a multi-modal combined image retrieval task can be completed.
Owner:HUBEI UNIV

Image retrieval method and apparatus, and storage medium and electronic device

The present application relates to the technical field of computers. Disclosed are an image retrieval method and apparatus, and a storage medium and an electronic device. The method comprises: using a preset deep hashing model to process an image to be retrieved, so as to obtain an image hash code, wherein the preset deep hashing model is obtained by means of performing incremental learning training on the basis of a plurality of preset class centers and query data; and on the basis of a preset image corresponding to a preset hash code matching the image hash code, obtaining a matched image. The present application effectively improves the accuracy of image retrieval.
Owner:SHENZHEN TCL NEW-TECH CO LTD

Defect automatic positioning method based on BIM virtual image

The invention relates to the technical field of computer vision processing, in particular to an automatic defect positioning method based on a BIM virtual image. The method comprises the following steps: establishing a high-rise building model, and making a data set integrating an illumination condition and a full view angle; calculating camera parameters; performing cross-modal image retrieval: performing targeted optimization on the basis of a classical ResNet architecture, and constructing a backbone network structure suitable for a building image retrieval task; initializing position attitude estimation based on matching; and correcting the camera position posture. According to the method, a defect fixed frame based on a building BIM virtual image is provided, a building image data set BIM-Vision based on Revit is constructed, rich visual angles and illumination condition setting are achieved, accurate camera position postures and 3D labels between beam columns are provided, high-quality basic data support is provided for building visual research, and the method has the advantages of being high in practicability and high in practicability. And during inspection, accurate positioning of defect positions and component association can be completed only by shooting a field image, so that the field operation process is greatly simplified.
Owner:DALIAN NATIONALITIES UNIVERSITY

Optimized container image retrieval for devices at edge locations

Techniques are described for enabling computing devices in computing environments distinct from a cloud provider network to use temporally staggered container image pull requests upon initiating execution of a container-based task. For example, upon determining to launch a container on a computing device running in a computing environment that is distinct from a cloud provider network, an agent running on the computing device can select a time in the future at which to request the container image(s) to be used to launch the container from a container registry based on a randomized value. The randomized value, for example, can enable each computing device in the environment to request the container image(s) at a time that differs from times at which other computing devices in the environment generate similar requests (e.g., such that the request times are “staggered” relative to one another).
Owner:AMAZON TECH INC

Building indoor positioning method and system based on AI image retrieval

The invention discloses a building indoor positioning method and system based on AI image retrieval, and relates to the technical field of indoor positioning, and the method comprises the steps: constructing an environment prior knowledge base, obtaining a real-time image for positioning, and generating an image comprehensive quality score; processing the real-time image to obtain a candidate positioning result set; calculating spatial dispersion according to the candidate positioning result set; according to the initial positioning result, determining a target logic area, extracting a corresponding scene visual ambiguity index, and generating a mismatch diagnosis index; according to the mismatch diagnosis index, the image comprehensive quality score and the spatial dispersion, a final positioning state is generated, and a calibration positioning result or a user instruction is output. Through a multi-module cooperation and innovation algorithm, scene difficulty is quantified, a positioning problem root is diagnosed, a closed-loop decision process is formed, the positioning reliability is improved, and the positioning accuracy is improved. And mismatching between the original confidence and the scene positioning difficulty is avoided.
Owner:HEBEI UNIV OF TECH

Multi-modal visual position identification reordering method and system based on guidance

The invention relates to the technical field of visual position recognition, and particularly discloses a multi-modal visual position recognition reordering method and system based on guidance, and the method comprises the steps: obtaining a query image, and retrieving a plurality of candidate images based on a pre-trained visual basic model and the query image; constructing a composite multi-modal prompt object, wherein the composite multi-modal prompt object comprises an image pair formed by the query image and the current candidate image, and an instruction text used for guiding a multi-modal large language model to perform visual comparison; outputting a structured similarity judgment result, wherein the result comprises a quantitative similarity score; and sorting based on the similarity scores corresponding to all the candidate images, and determining the candidate image with the highest score as an optimal matching result. Through combination of guiding type prompt engineering and structured output, an intermediate text generation link is avoided fundamentally, and the calculation efficiency is improved while the fidelity of all original visual information is reserved.
Owner:SHENZHEN 1024 ROBOT TECHNOLOGY CO LTD

Comparison learning model training method, image retrieval method and device

The invention provides a training method of a contrast learning model and an image retrieval method and device, and relates to the technical field of image retrieval, and the method comprises the steps: carrying out the feature alignment of a document real-shot image sample and an electronic document image sample, and obtaining the feature-aligned electronic document image sample, constructing a positive and negative sample image pair based on the samples after feature alignment; and training the initial image contrast learning model based on the constructed positive and negative sample image pairs to obtain an image contrast learning model. And constructing a positive and negative sample image pair based on the samples after feature alignment, and training an image comparison learning model. Therefore, the trained image comparison learning model does not pay attention to the interference feature difference irrelevant to the semantic content, and the model is guided to focus attention on the essential semantic feature content of the learning image. Therefore, the trained image comparison learning model is more accurate in discrimination and higher in efficiency on the basis of a semantic content two-image discrimination task.
Owner:HEFEI IFLYTEK TOYCLOUD TECH