Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

784results about "Metadata still image retrieval" patented technology

Multi-modal agricultural question and answer method and system for generating RAG (Retrieval Enhanced Generation) based on retrieval

The invention discloses a multi-modal agricultural question and answer method and system for generating RAG based on retrieval enhancement, and the method comprises the steps: collecting and constructing crop disease image data containing a farmland complex background and corresponding text description, and forming an agricultural image-text knowledge database; based on the agricultural image-text knowledge database, multi-modal features of an input image and a query text are extracted, image-text joint similarity retrieval is carried out, and knowledge image-text candidates are obtained; inputting the knowledge image-text candidates into an image-text rearrangement module for rearrangement; taking the reordered image-text knowledge as condition input, accessing a large language model, and generating diagnosis description and prevention and treatment suggestions for the current crop diseases; and outputting a multi-modal question and answer result according to the image and question input by the user. According to the method, the agricultural image-text database is constructed, and multi-modal retrieval, rearrangement and large language model generation technologies are combined, so that accurate and efficient crop disease multi-modal question answering is realized, and the reliability of diagnosis and prevention suggestions is improved.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Generating descriptive tags for images that characterize a condition of utility assets

A condition generator analyzes a set of utility asset images of a particular utility asset using the ML (machine learning) model identify a type and condition of the particular utility asset depicted in the set of utility asset images. The identified condition is assigned a confidence score, and the set of utility asset images includes at least two images of the particular utility asset captured at different angles. The condition generator generates a descriptive tag for the set of utility asset images based on the identified type and condition. The descriptive tag characterizes an operational status of the particular utility asset. The condition generator stores the set of utility asset images and the generated descriptive tag in a utility asset database. The utility asset database stores images of utility assets.
Owner:FLORIDA POWER & LIGHT CO

Artificial intelligence-based image search refinement

Systems and methods for image search result filtering can include obtaining a search query, determining a plurality of candidate image search results, processing the search query with a generative model to determine a plurality of search result criteria, and refining the plurality of candidate image search results based on determining whether the candidate results satisfy the plurality of search results criteria. The systems and methods can perform a plurality of determinations based on the output of the generative model.
Owner:GDM HOLDING LLC

Section-free image service system and method based on dynamic projection and real-time mosaic

The invention belongs to the technical field of geographic information technology and remote sensing image processing and service, and discloses a slice-free image service system and method based on dynamic projection and real-time mosaic, which abandons the traditional pre-slicing mode, directly provides online service based on original image data, and improves the service efficiency. The problem of redundant storage caused by pre-generation and storage of massive tiles is fundamentally eliminated, and storage space is saved by up to 90%. Meanwhile, when the original data is updated, the system does not need to carry out a time-consuming re-slicing process, real-time updating and publishing of services can be realized, and the core pain points of long data updating period and high delay in the traditional technology are thoroughly solved. Through the integrated dynamic projection engine and the real-time mosaic module, the on-demand service request of the client for any coordinate system, any spatial range and any resolution can be responded.
Owner:JINGZHOU INSTITUTE OF SURVEYING & MAPPING (JINGZHOU INSTITUTE OF LAND & SPACE PLANNING JINGZHOU NATURAL RESOURCES SATELLITE APPLICATION TECHNOLOGY CENTER)

Automatic material identification and matching method and system for three-dimensional model and medium

The invention provides an automatic material identification and matching method and system for a three-dimensional model and a medium, and relates to the technical field of material management, and the method comprises the steps: constructing a standardized material database, and extracting and standardizing material attribute information; analyzing first material description information in the three-dimensional software model by utilizing an SQL (Structured Query Language), and matching the first material description information with the material attribute information to obtain a first material matching result; extracting second material description information in the project material document by using VBA programming, and matching the second material description information with the material attribute information to obtain a second material matching result; and data integration is carried out through PowerBI, and a project material state visual billboard is constructed. The technical problems that in the prior art, a unified standardized material database is lacked, material description formats in different sources are inconsistent, expression is not standard, a construction material list cannot be accurately generated, and informatization and intellectualization of project material management are limited are solved.
Owner:ZHONGHAI FULU HEAVY IND CO LTD

Method and system for quickly retrieving and matching inspection images of power distribution network

The invention relates to the technical field of power grid image retrieval, and discloses a power distribution network inspection image rapid retrieval matching method and system, and the method comprises the steps: obtaining a to-be-retrieved inspection image of power distribution network equipment, and extracting the equipment structure features of the inspection image through a hierarchical convolutional network; quantifying a surface texture attenuation index of the power distribution network equipment through fractal dimension based on the equipment structure characteristics; performing mapping relation coupling on the surface texture attenuation index and a space coordinate of an equipment connecting piece to generate a dynamic feature coding sequence containing an equipment structure topological relation; performing time sequence consistency matching on the dynamic feature coding sequence and a pre-constructed reference image library, and aligning an equipment aging track through a dynamic time warping algorithm to generate a similarity sorting result; according to the method, the problems that effective features cannot be extracted during retrieval matching and the retrieval precision is low are solved.
Owner:安徽明生恒卓科技有限公司 +1

Image secret cloud storage method and device based on feature block encryption

The invention discloses an image secret type cloud storage method and device based on feature block encryption, and the method comprises the steps: a cloud server receives and stores an encrypted image, an encrypted image feature set and a feature inner product, and forms an image database; the cloud server receives a feature ciphertext query feature uploaded by the authorized user when the authorized user needs to query the image; the cloud server performs retrieval according to the received ciphertext query features, separately calculates the similarity between color and texture sub-features in the ciphertext query features and features in each encrypted image feature set in the image database by using a hierarchical earth moving distance algorithm, and calculates the comprehensive similarity corresponding to each encrypted image; and screening out a plurality of encrypted images according to the comprehensive similarity and returning the encrypted images to the authorized user. According to the method, safe and efficient search of the image can be realized, and the image retrieval performance is improved.
Owner:CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY

Training method and application method of image recognition retrieval model, equipment and medium

The invention discloses a training method and an application method of an image recognition retrieval model, equipment and a medium, and belongs to the field of image recognition. The method comprises the following steps: extracting visual modal features of a known category; obtaining semantic modal information of the seen and unseen categories, wherein the semantic modal information comprises an attribute modal, a category name modal and a semantic description modal; fusing the semantic modal information to obtain relevance semantic representation; training a generative adversarial network based on the visual modal features of the known categories and the semantic representation, and generating pseudo-visual modal features of the unseen categories; and training at least one of a classification model and a retrieval model based on the pseudo-visual modal features corresponding to the unseen categories, so that the classification model at least identifies the image samples of the unseen categories, and / or the retrieval model at least retrieves retrieval results matched with the unseen categories. According to the method, a category separation optimization mechanism is introduced, so that the generalization ability and the recognition accuracy of the model under the zero sample condition are improved.
Owner:GRG BANKING EQUIPMENT CO LTD

Efficient data storage and retrieval method for point cloud digital twinning

The invention relates to the technical field of data storage, and particularly provides a point cloud digital twinning-oriented efficient data storage and retrieval method, which comprises the following steps of: converting an original disordered point set into a dynamic voxel unit with probability density by analyzing a spatial distribution entropy value of a point cloud; generating a hierarchical voxel network with entropy weight marks; based on a hierarchical voxel network, extracting a voxel state transition rule across a time sequence, and encoding dynamic change features into a directed weighted graph; map nodes of the directed weighted map represent voxel stability, and a topological map with a space-time coupling characteristic is formed; training an implicit neural network to generate a continuous semantic vector field by using the space-time coupling characteristic of the topological graph; during retrieval, a target feature region is directly positioned through vector field gradient descent and then is reversely mapped to an original voxel level, and discrete retrieval is converted into vector navigation in a continuous space. According to the method, the contradiction between storage compression and retrieval efficiency is unified.
Owner:HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER +1

Wafer defect retrieval method and device and storage medium

The invention provides a wafer defect retrieval method and device and a storage medium, and the method comprises the steps: obtaining a to-be-retrieved wafer image and related text information, the to-be-retrieved wafer image comprises defect features, and the related text information of the to-be-retrieved wafer image comprises an image description text; inputting the to-be-retrieved wafer image and the related text information into a pre-trained multi-modal feature extraction model to obtain multi-modal features of the to-be-retrieved wafer image; and retrieving a preset wafer defect database by using the multi-modal features of the to-be-retrieved wafer image to obtain similar cases of the to-be-retrieved wafer image. By means of the scheme, the accuracy and distinction degree of the detection result can be improved, and the wafer defect retrieval efficiency is improved.
Owner:SKYVERSE TECH CO LTD

Picture sharing method and device based on AI platform, electronic equipment and storage medium

The invention relates to the technical field of picture sharing, and discloses a picture sharing method and device based on an AI platform, electronic equipment and a storage medium, and the method comprises the steps: obtaining a generated data stream of a picture uploaded by a first user, converting the generated data stream into a structured JSON file, and enabling a built-in cross-platform compatible tag to support multi-terminal sharing; based on the second user creation preference portraits and the key features, non-key parameters are screened through an AI semantic matching model, an adaptive priority adjustable component is generated, and a difference thermodynamic diagram and a new picture are generated after user adjustment; the enhanced copyright metadata containing original author identification, content fingerprints and the like are embedded into the JSON parameter stream; and calculating a hierarchical hash value and carrying out block chain evidence storage during each parameter iteration, calculating contribution degree distribution earnings based on a multi-dimensional index, generating a corresponding NFT copyright certificate, and supporting on-chain reverse tracing. According to the method and the device, the problem of low efficiency in secondary creation parameter adjustment caused by lack of standardized structured packaging and dependency relationship labeling in parameter generation can be solved.
Owner:URBAN PLANNING & DESIGN INST OF SHENZHEN UPDIS

Shared geographic information system for creation, change management and delivery of 3D structure models

A new generation 3D modeling system is described for the management of the 3D characteristics of man-made structures in a shared geographic information system (GIS) platform. The system accommodates the conversion of 2D GIS topology to 3D real-world topology, originate and perpetually maintain 3D models and non-graphical attribute data through time managed by a community of users based on a series of permissions and user roles. The system integrates 3D model data with attribute data and may be supplemented with published map or attribute data and / or data made available by users of the system.
Owner:GEOSPAN CORP

Precipitation risk assessment and early warning method and system based on InSar technology

The invention relates to the technical field of meteorological monitoring, in particular to a rainfall risk assessment and early warning method and system based on the InSar technology, and the method comprises the following steps: obtaining a deformation graph and rainfall data, aligning image time and space, screening a settlement region after rainfall, extracting a trend consistent segment and an abnormal mapsheet, and judging a response unit through combining direction turning and rate mutation. And extending a tracking settlement process, and polymerizing to form a risk area range. According to the method, the time-space corresponding relation of rainfall starting and deformation response is screened, a settlement area caused under the rainfall effect is screened, interference response is eliminated by combining curve direction change and trend consistency, an abnormal unit is positioned by means of rate sudden change and direction turning, and the tracking response process is extended along the settlement trend. And the mapsheet fragments are aggregated to form a space coverage range, so that the judgment capability of a precipitation triggered settlement process is enhanced, the capture depth of a local settlement dynamic state is improved, and the construction continuity and response granularity of a risk region boundary are perfected.
Owner:HENGFENG INFORMATION TECH CO LTD

Intelligent quality inspection method and system based on cooperation of multi-modal large model and industrial knowledge base

The invention discloses an intelligent quality inspection method and system based on cooperation of a multi-modal large model and an industrial knowledge base, and belongs to the technical field of artificial intelligence. According to the method, the cross-modal semantic alignment capability of a multi-modal large model and static knowledge of an industrial knowledge base are fused through the steps of industrial quality inspection data acquisition, preprocessing, semantic guidance abnormal region positioning, accurate defect segmentation, industrial knowledge base construction and retrieval, multi-modal large model semantic reasoning and decision making, feedback closed loop execution and the like; and the upgrade from'defect identification 'to'intelligent diagnosis' is realized. The system is provided with seven modules corresponding to the steps of the method to cooperatively complete the whole process of quality inspection. According to the method, the problems of weak generalization ability, semantic understanding deficiency, decision isolation and the like of traditional quality inspection are effectively solved, the quality inspection accuracy and efficiency are improved, and the industrial production quality control cost is reduced.
Owner:GUANGXI ACAD OF SCI

Distributed medical image retrieval and hierarchical storage management system and method

The invention discloses a distributed medical image retrieval and hierarchical storage management system and method. The system comprises an interface processing module, a distributed image storage cluster, a routing forwarding module, a metadata management module and a storage management module. The interface processing module receives the image retrieval request, generates a routing identifier for positioning the storage position of a target medical image file and sends the routing identifier to the routing forwarding module; and the routing forwarding module queries a mapping relationship between the medical image file and the storage node and / or the storage hierarchy based on the routing identifier, determines a first storage node and forwards the request, so that the first storage node returns a target medical image file. And the storage management module obtains the access statistical information, determines a target storage level and / or a target storage node, migrates the medical image file, and triggers updating of the mapping relation after migration is completed. Therefore, cross-node accurate positioning and unified retrieval are realized, positioning consistency is kept after file migration, and hierarchical storage management based on access conditions is supported.
Owner:安徽影联云享医疗科技有限公司

Embedding based retrieval for image search

Methods, systems, and apparatus including computer programs encoded on a computer storage medium, for retrieving image search results using embedding neural network models. In one aspect, an image search query is received. A respective pair numeric embedding for each of a plurality of image-landing page pairs is determined. Each pair numeric embedding is a numeric representation in an embedding space. An image search query embedding neural network processes features of the image search query and generates a query numeric embedding. The query numeric embedding is a numeric representation of the image search query in the same embedding space. A subset of the image-landing page pairs having pair numeric embeddings that are closest to the query numeric embedding of the image search query in the embedding space are identified as first candidate image search results.
Owner:GOOGLE LLC

Image retrieval method and device

The invention discloses an image retrieval method and device, and relates to the technical field of knowledge distillation, and the method comprises the steps: taking a query text input by a user as a retrieval, and screening and outputting a target image consistent with text content description from a preset image library through text semantic comprehension and an image feature matching algorithm, the technical problems that in the related technology, it is difficult to deeply mine core knowledge such as fine-grained probability distribution and coarse-grained structural features of an attention mechanism in the image retrieval process, a distillation process cannot cover full-dimensional knowledge of an attention level, and the integrity of knowledge migration is insufficient are solved. The technical effects that through deep correlation matching of the text and the image, the user can be helped to rapidly position the target image through simple text description, the image search operation process is effectively simplified, and efficient and convenient image retrieval and browsing experience are provided for the user are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Three-dimensional model retrieval method based on double viewpoint sampling and modal constraint

The invention relates to the technical field of three-dimensional model retrieval, in particular to a three-dimensional model retrieval method based on double viewpoint sampling and modal constraint, which comprises the following steps of: constructing a double-branch image processing network model; inputting the rendering view into an instance-level view sampling module for processing to obtain a sampling view; inputting the sampling view into a hierarchical self-attention view sampling module for processing to obtain a final descriptor of the three-dimensional model and a category center of the three-dimensional model; inputting a sketch in the sketch modal data set into a sketch feature learning module to obtain sketch features; the modal constraint module designs modal constraint loss, and performs weighted fusion by using the modal constraint loss, the alignment loss and the classification loss to obtain a total loss function. According to the method, the compact and efficient three-dimensional model descriptor can be extracted, the domain difference between the three-dimensional model and the sketch can be effectively reduced, and the effectiveness of three-dimensional model retrieval is improved.
Owner:JIANGXI NORMAL UNIV

Deep hash image retrieval method based on diffusion model for power grid defect maintenance

The invention relates to the field of power grid defect retrieval, in particular to a diffusion model-based deep hash image retrieval method for power grid defect maintenance, which comprises the following steps of: 1, performing fusion coding by inputting text data and image data, and constructing an initial hash code generation model; 2, using a Pair-wise loss function to optimize the distribution of sample pairs in a hash space, introducing a quantization loss function, generating an efficient binary hash code, and generating a high-quality binary hash code; 3, constructing a Hash code-image latent diffusion model, performing diffusion generation by encoding and decoding the Hash code / image to a continuous latent space, enabling the Hash code to correspond to the image in a generative manner, and directly fitting spatial distribution; and 4, defining a loss function of the Hash code-image diffusion model, generating a high-quality Hash code and an image, and obtaining a power grid defect type in a mode of searching images by images. Auxiliary training is carried out through fusion of text features and image features, so that the semantic features understand the images more deeply.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO JIMO POWER SUPPLY CO

Automatic examination method and system for three-dimensional model of power grid

The invention discloses an automatic review method and system for a three-dimensional model of a power grid, and belongs to the technical field of power grid engineering design and intelligent review, and the system comprises the steps: analyzing the three-dimensional model of the power grid into a JSON file containing model attribute information; based on the model attribute information in the JSON file, retrieving from the OWL ontology knowledge base through an RAG retrieval enhancement module to obtain related standard knowledge; based on the model attribute information and related standard knowledge in the JSON file, executing model attribute automatic verification through a review rule script, and obtaining a model attribute verification result; and based on the model attribute information, the related standard knowledge and the model attribute verification result in the JSON file, guiding a power grid bright large model reasoning unit to generate a review result through a cue word template. According to the method, the problems of low efficiency, non-uniform standard, high error rate and the like of existing manual review are solved, and automatic and intelligent review of the three-dimensional model of the power grid is realized.
Owner:STATE GRID SHANGHAI ELECTRIC POWER DESIGN

Method and system for text retrieval of picture archives based on cross-modal feature alignment

The invention discloses a method and a system for retrieving a picture file through a text based on cross-modal feature alignment, and the method comprises the steps: extracting text and picture features, and guaranteeing that a high-value mode contributes to a higher weight based on multi-modal attention weighted fusion; calculating a dynamic temperature coefficient through initial semantic similarity of positive and negative samples to construct a loss function for multi-modal contrast learning, and mapping original features of a text and an image to a unified space to obtain alignment features; according to the method, the picture archives are retrieved through texts, the image-text similarity, the time sequence weight and the core area proportion weight are comprehensively considered, optimization sorting of time sequence perception is carried out, and the most matched picture archives are obtained. According to the method, a text-image cross-modal semantic gap is solved, so that semantic alignment of two types of features in a unified space is realized; according to the method, the sample difficulty is dynamically adapted to improve the feature distinction degree; weights of texts and images are distributed according to needs so as to retain core information; according to the method, retrieval result sorting is optimized in combination with time attributes.
Owner:ZHEJIANG UNIV OF FINANCE & ECONOMICS

Retrieval method for open-vocabulary 3D target, device and storage medium

A retrieval method for an open-vocabulary 3D target, a device and a storage medium are provided. The method includes: inputting text description information of a target object and an image sequence of a real scene into an open-vocabulary 3D target retrieval model. The retrieval model includes an LLM, an open-vocabulary 2D detection and segmentation model, an object filtering module and a 3D projection module. The LLM is used to enhance the text description information to obtain names of a plurality of candidate objects of the target object. The open-vocabulary 2D detection and segmentation model is used to detect and segment according to the names of the candidate objects, to obtain 2D segmentation masks of the plurality of candidate objects. The object filtering module is used to determine the target object from the candidate objects according to the text description information and a 2D segmentation mask of each candidate object.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Computer-based systems for determining a look-alike domain names in webpages and methods of use thereof

A method includes instructing a computing device of a user to determine when a webpage is loaded via a web browser; where the web browser includes at least one of: a plug-in or an extension associated with the provider server associated with a first entity; where the plug-in or the extension is configured to request and receive a unique schema-specific identifier from the provider server; instructing the computing device to determine that the webpage is a particular webpage based on a determination that the webpage includes at least one field to be populated with user information; determining that the particular webpage is associated with a look-alike domain name based on: comparing a domain name of the particular webpage with at least one domain name of at least one second entity associated with at least one previously used unique schema-specific identifier associated with the user or at least one other user.
Owner:CAPITAL ONE SERVICES LLC

Landscape greening maintenance planning method and system based on big data

The invention relates to the technical field of maintenance planning, and particularly discloses a landscape greening maintenance planning method and system based on big data, and the method comprises the steps: constructing a maintenance region model, and updating the maintenance region model based on a shot image; rasterizing the maintenance area model, and determining greening parameters in each grid area; and clustering the grid areas based on the greening parameters, synchronously determining the demand precision of each type of grid areas, obtaining the real-time information of each grid area, feeding back the real-time information to the management end, and receiving a maintenance scheme which is fed back by the management end and contains a maintenance target. According to the method, the landscape model is constructed based on the BIM model, the landscape model is updated based on a large amount of image data, the real-time performance of the landscape model is extremely high, software analysis is performed on the landscape model, a maintenance target can be determined, a more targeted maintenance task can be generated, and the labor workload is greatly reduced under the condition that the maintenance effect is almost unchanged.
Owner:AI LU ENG CONSTR (SHANGHAI) CO LTD

Common error automatic screening method based on BIM model

The invention discloses a method for automatically screening common errors in a BIM model in engineering. The method comprises the following steps: opening the BIM model, and opening an automatic screening device for the common errors of the BIM model; selecting the BIM model to be subjected to error screening from the BIM models; selecting a function module of error type screening to be executed in the BIM model common error automatic screening device, and operating the device; the device circularly traverses the selected BIM model and sequentially compares BIM model data according to a set error screening rule, if the BIM model data accords with the error screening rule, error prompt information is generated, and if the BIM model data does not accord with the error screening rule, the step is quitted; the BIM model checking efficiency is greatly improved, and meanwhile the problems that manual checking is low in efficiency, high in omission ratio and high in misjudgment rate are solved.
Owner:SHANGHAI BAOYE GRP CORP +1

Aesthetic image retrieval system and method

A method of retrieving visual content includes receiving user input defining an initial search query from a client application. The initial search query and a meta prompt are then delivered to a refined query generating model which is trained to analyze the initial search query to determine user intent and to generate a refined search query based on the initial search query and the meta prompt. The refined search query is delivered to a visual content retrieval model which retrieves aesthetic visual content with reference to a visual content index. Retrieved aesthetic visual content is returned to the client application.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

User gallery multi-modal retrieval method and system based on AI intention recognition

The invention relates to the field of information retrieval, and provides a user gallery multi-modal retrieval method and system based on AI intention recognition. The method comprises the following steps: performing intention analysis on a query statement input by a user through a multi-modal semantic understanding model to obtain an intention classification result and structured query representation; according to the intention classification result, distributing a retrieval weight of multi-source gallery data, and obtaining a weight configuration scheme for the current query; inputting the structured query representation and the weight configuration scheme into a hybrid retrieval engine, and performing collaborative retrieval processing on image data, text data and metadata in a user image library to obtain a candidate retrieval result set; and according to a multi-modal similarity calculation method, performing comprehensive scoring on a plurality of items in the candidate retrieval result set to obtain a sorted final retrieval result. According to the method, the complex query intention can be accurately identified, the retrieval weight is dynamically adjusted, and accurate multi-modal collaborative retrieval is realized on the premise of protecting privacy.
Owner:E-SURFING DIGITAL LIFE TECH CO LTD

Visual question and answer method based on language and visual fine-grained semantic alignment

The invention provides a visual question and answer method based on language and visual fine-grained semantic alignment, which comprises the following steps: describing group images and texts thereof, extracting image features by using an image encoder of a contrast language-image pre-training model, namely a CLIP model, and extracting text features by using a text encoder of a pre-training language model BERT based on a Transform architecture; obtaining an image lexical element sequence and a text lexical element sequence; correspondingly generating a multi-granularity global image feature set and a multi-granularity global text feature set; obtaining a joint similarity; inputting the joint similarity scf, the multi-granularity global image feature set and the multi-granularity global text feature set into a decoder of a Transform architecture to generate a final answer; according to the method, a fine-grained information mining process is introduced, fine differences between images and texts can be more accurately captured, more accurate image and text association can be established, and the accuracy and correlation of answer generation can be improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Cross-view geographic positioning acceleration method based on semantic description

The invention relates to the technical field of computer vision and geographic information, and particularly discloses a cross-view geographic positioning acceleration method based on semantic description, which comprises the following steps of: performing instruction fine tuning on an input query image and a database image by adopting a visual language model to generate natural language description; coding the natural language description into a low-dimensional semantic vector by utilizing a word embedding model; constructing a semantic keyword association graph based on a graph embedding technology, and fusing a general semantic knowledge base and a geographic domain ontology to construct a surface feature special semantic tree; screening candidate subsets according to a preset ground feature priority and a saliency detection result of the query image; extracting global features and local key point descriptors of the images in the candidate subsets by adopting a lightweight model; constructing a visual similarity calculation model based on the state space model, and modeling a feature sequence dependency relationship; and fusing the semantic matching score and the visual similarity, outputting a final matching result by adopting a weighted sorting strategy, and cooperatively accelerating according to the matching result in combination with a lightweight model.
Owner:PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV