Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

582results about "Still image data clustering/classification" patented technology

Multimodal ai-based search for digital assets

Embodiments of the present disclosure relate to multimodal AI-based search for digital assets via an indexing and / or search pipeline. With respect to the indexing pipeline, some embodiments obtain first data and second data associated with a first digital asset. Such data represents different data types or modalities of the same digital asset. After obtaining the first and second data, some embodiments then generate a composite index. After the composite index is built such index can then be used to execute a query via the search pipeline. To execute the query some embodiments compute a relevance score for each digital asset, of multiple digital assets, based at least in part on a measure in which each digital asset satisfies one or more parameters or conditions for two or more data types of the query. Various embodiments then rank each digital asset and present one or more associated indicators.
Owner:NVIDIA CORP

Building engineering crack detection method and system based on image recognition

The embodiment of the invention discloses a building engineering crack detection method and system based on image recognition, and the method comprises the steps: obtaining a building surface image, carrying out the preprocessing of the image, obtaining a standardized image, and carrying out the multi-scale decomposition extraction and integration of various features, and forming a multi-dimensional feature descriptor set; after feature importance is evaluated, a compact feature vector is generated through dimension reduction, quantization coding and compression, and then a multi-level feature index mechanism for optimized compression is constructed. A query feature vector is extracted from a newly collected image, searching and screening are completed by means of an index mechanism and a tolerance threshold, and a crack matching result is obtained; and based on the result, positioning cracks, classifying types, measuring parameters and evaluating severity, and generating a crack state report. The crack trend is analyzed in combination with the historical data time sequence, a multi-stage early warning mechanism is designed, maintenance suggestions are provided, and a real-time monitoring and early warning system is formed. According to the embodiment of the invention, the technical problems of high storage pressure and low real-time detection efficiency in the prior art can be effectively solved.
Owner:内江市住房保障和房地产事务中心

Photo content clustering for digital picture frame display and automated frame storytelling

A method and system for automated routing of pictures taken on mobile electronic devices to a digital picture frame including a camera, microphone, and speaker integrated with the frame, and a network connection module allowing the frame for direct contact and upload of photos from electronic devices or from photo collections of community members. Clustering photos by content is used to improve display and to respond to photo viewer desires. Trends or patterns can be detected from the photo collections and that information used for various purposes beyond photo display. The frame includes a conversational intelligence that provides a verbal communication with a viewer, such as for determining an identity or preferences of the frame viewer, determining photos to display for the viewer, discussing displayed photos with the viewer, or telling stories or life histories to the viewer based upon photo content.
Owner:PUSHD INC

Data processing method and device based on smart city

The invention provides a data processing method and device based on a smart city, and relates to the technical field of data processing.The method comprises the steps that a city-level spatial data base is constructed through oblique photography of an unmanned aerial vehicle, and a unified and real spatial foundation is formed by generating a live-action three-dimensional model, a digital orthoimage and a digital elevation model; on this basis, a building base and a building white model are extracted, spatial logic verification is carried out on multi-source business basic data, a standard address system stably associated with a building three-dimensional entity is further constructed, precise spatial anchoring of governance objects such as population, houses and units is achieved, and the method is suitable for mass production. And finally, the associated data is uniformly converged to a city information model platform, and reliable and updatable spatial data support is provided for smart city application. According to the invention, the timeliness of smart city data processing in a dynamic earth surface deformation scene can be improved.
Owner:湖北省国土测绘院

Object and action recognition via text-based classification models

Example implementations include a method, apparatus and computer-readable medium of object / action recognition using a text-based classification model, comprising generating an image vector configured to represent one or more features of the first image. Additionally, the implementations further include computing a vector distance between the image vector and each of a first text vector and a second text vector, wherein the first text vector is configured to represent a first text, and wherein the second text vector is configured to represent a second text. Additionally, the implementations further include classifying the first image according to the first text or the second text based on which computed vector distance indicates a highest similarity between the image vector and either the first text vector or the second text vector relative to the other of the first text vector or the second text vector.
Owner:TYCO FIRE & SECURITY GMBH

Deep hash image retrieval method based on diffusion model for power grid defect maintenance

The invention relates to the field of power grid defect retrieval, in particular to a diffusion model-based deep hash image retrieval method for power grid defect maintenance, which comprises the following steps of: 1, performing fusion coding by inputting text data and image data, and constructing an initial hash code generation model; 2, using a Pair-wise loss function to optimize the distribution of sample pairs in a hash space, introducing a quantization loss function, generating an efficient binary hash code, and generating a high-quality binary hash code; 3, constructing a Hash code-image latent diffusion model, performing diffusion generation by encoding and decoding the Hash code / image to a continuous latent space, enabling the Hash code to correspond to the image in a generative manner, and directly fitting spatial distribution; and 4, defining a loss function of the Hash code-image diffusion model, generating a high-quality Hash code and an image, and obtaining a power grid defect type in a mode of searching images by images. Auxiliary training is carried out through fusion of text features and image features, so that the semantic features understand the images more deeply.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO JIMO POWER SUPPLY CO

Parameter efficient prompt tuning for efficient models at scale

Systems and methods for natural language processing can leverage trained prompts to condition a large pre-trained machine-learned model to generate an output for a specific task. For example, a subset of parameters may be trained for the particular task to then be input with a set of input data into the pre-trained machine-learned model to generate the task-specific output. During the training of the prompt, the parameters of the pre-trained machine-learned model can be frozen, which can reduce the computational resources used during training while still leveraging the previously learned data from the pre-trained machine-learned model.
Owner:GOOGLE LLC

Searcher model training method, and Burt-Hooger-Durob syndrome recognition method and system based on retrieval enhancement generation

The invention discloses a searcher model training method, and a Burt-Hooge-Durob syndrome recognition method and system based on retrieval enhancement generation, and relates to the field of rare disease recognition. The training method comprises the steps of obtaining a positive sample pair and a negative sample pair; the positive sample pair and the negative sample pair are input into an initial searcher model, a loss function of the initial searcher model is calculated, and the loss function adjusts the size of an angle margin in real time according to a measurement variance adaptive mechanism; and optimizing the initial retrieval model according to a loss function calculation result. According to the method, the angle interval between the BHD and the non-BHD is forcibly expanded through the angle margin of the loss function, and the angle margin is dynamically adjusted according to the statistical variance of the cosine similarity between all positive sample pairs in the current training batch by using a measurement variance adaptive mechanism; the problems of weak DCLDs image difference and fuzzy category decision boundary caused by very similar iconography features of various types of rare diseases of DCLDs are solved, and the recognition precision of a large model on query information is improved.
Owner:UNIV OF SCI & TECH OF CHINA

Automatically monitoring retail products based on captured images

A system for acquiring images of products in a retail store is disclosed. The system may include at least one first housing configured for location on a retail shelving unit, and at least one image capture device included in the at least one first housing and configured relative to the at least one first housing such that an optical axis of the at least one image capture device is directed toward an opposing retail shelving unit when the at least one first housing is fixedly mounted on the retail shelving unit. The system may further include a second housing configured for location on the retail shelving unit separate from the at least one first housing, the second housing may contain at least one processor configured to control the at least one image capture device and also to control a network interface for communicating with a remote server. The system may also include at least one data conduit extending between the at least one first housing and the second housing, the at least one data conduit being configured to enable transfer of control signals from the at least one processor to the at least one image capture device and to enable collection of image data acquired by the at least one image capture device for transmission by the network interface.
Owner:TRAX TECH SOLUTIONS

Land parcel improvements with machine learning image classification

ActiveUS12555361B2FinanceProduct appraisalLand improvementData set
A method for identification of land improvements in a given parcel of land that includes generating a first training data set by clipping a large image of a parcel of land into individual parcel images each including images of improvements, and each improvement being labeled with an improvement type, providing the individual parcel images to a first classification model, training the first classification model based on the individual parcel images to identify unlabeled improvements in a parcel image and to obtain a multi-label classifier, generating a second training data set of images, and training a second, semantic segmentation model based on the second training data set.
Owner:MEDICI LAND GOVERNANCE

Parameter Efficient Prompt Tuning for Efficient Models at Scale

Systems and methods for natural language processing can leverage trained prompts to condition a large pre-trained machine-learned model to generate an output for a specific task. For example, a subset of parameters may be trained for the particular task to then be input with a set of input data into the pre-trained machine-learned model to generate the task-specific output. During the training of the prompt, the parameters of the pre-trained machine-learned model can be frozen, which can reduce the computational resources used during training while still leveraging the previously learned data from the pre-trained machine-learned model.
Owner:GOOGLE LLC

Agricultural disease hash retrieval method based on DV stabilization and adaptive feature enhancement

The invention discloses an agricultural disease hash retrieval method based on DV stabilization and adaptive feature enhancement, and belongs to the technical field of agricultural disease image retrieval and deep hash, and the method comprises the following steps: 1, obtaining an original disease image as the input of a backbone network; step 2, extracting disease image features by using a backbone network; step 3, based on a self-adaptive feature enhancement module, enhancing the disease image features to obtain self-adaptive image features; step 4, optimizing network parameters through a total loss function containing DV stabilization; step 5, obtaining a binary hash code through the hash layer; and step 6, carrying out Hamming distance sorting on the obtained Hamming codes and the Hamming codes of all the images in the retrieval set obtained by the same method, and returning a plurality of disease images with the Hamming distance smaller than a preset Hamming distance. The agricultural disease image retrieval efficiency and precision are effectively improved.
Owner:SHANDONG UNIV OF SCI & TECH

Construction equipment safety early warning and prediction method based on dynamic perception of unmanned aerial vehicle

The invention provides a construction equipment safety early warning and prediction method based on unmanned aerial vehicle dynamic perception. The method comprises the following steps that real-time video data of a construction site and high-precision pose data of an unmanned aerial vehicle are collected through an inspection unmanned aerial vehicle; processing the video data, identifying construction equipment, key components and constructors, and obtaining geographic coordinates of the construction equipment, the key components and the constructors; generating a dynamic electronic fence changing along with the operation state based on the key component; predicting a future movement track of the construction equipment and analyzing track abnormity in combination with historical data and task types of the equipment; calculating a collision risk level between the construction equipment, and optimizing risk judgment in combination with an out-of-range overlapping degree; according to different construction equipment types, dynamically planning an unmanned aerial vehicle inspection path through a differential inspection strategy; and risk analysis is performed in combination with the dynamic electronic fence and the prediction result, construction equipment related safety early warning is generated, and the safety management level of the construction site can be improved through the method.
Owner:HUBEI HIGHWAY ENG CONSULTANTS SUPERVISION CENT

User gallery multi-modal retrieval method and system based on AI intention recognition

The invention relates to the field of information retrieval, and provides a user gallery multi-modal retrieval method and system based on AI intention recognition. The method comprises the following steps: performing intention analysis on a query statement input by a user through a multi-modal semantic understanding model to obtain an intention classification result and structured query representation; according to the intention classification result, distributing a retrieval weight of multi-source gallery data, and obtaining a weight configuration scheme for the current query; inputting the structured query representation and the weight configuration scheme into a hybrid retrieval engine, and performing collaborative retrieval processing on image data, text data and metadata in a user image library to obtain a candidate retrieval result set; and according to a multi-modal similarity calculation method, performing comprehensive scoring on a plurality of items in the candidate retrieval result set to obtain a sorted final retrieval result. According to the method, the complex query intention can be accurately identified, the retrieval weight is dynamically adjusted, and accurate multi-modal collaborative retrieval is realized on the premise of protecting privacy.
Owner:E-SURFING DIGITAL LIFE TECH CO LTD

Label inheritance for soft label generation in information processing system

Label inheritance techniques are disclosed for soft label generation in an information processing system that uses machine learning. For example, a method generates at least one label for a given data instance from a training data set useable to train a machine learning-based model. The at least one label is generated by assigning one or more labels associated with one or more ancestors of the data instance such that the data instance inherits the one or more labels associated with the one or more ancestors as the at least one label.
Owner:DELL PROD LP

Zero-shot reasoning in vision-language models

Disclosed are examples of training-free systems, methods and apparatuses, rooted in Chainof-Thought (CoT) reasoning, used to enhance the zero-shot performance of vision language models (VLMs) such as CLIP on a variety of downstream tasks. Hierarchical questions reflecting human visual cognition can be used with a pre-trained visual question answering model to extract the context of a query image from a global to local perspective through strategic questioning. Those CoT-based question-answer (QA) pairs, in conjunction with predefined class names, can serve as input to a language encoder, resulting in multi-level textual embeddings that emphasize various aspects of the image to improve existing VLM performance without additional training or labelled data.
Owner:NATIONAL UNIVERSITY OF SINGAPORE

Method, computer device, and computer-readable recording medium to provide message summary and associated image

A method of providing a message summary and an associated image may include requesting an image search in relation to a message summary created based on a message in a chatroom; receiving an image bundle that includes at least one image in response to an image search request; and displaying the image bundle in association with the message summary.
Owner:LINE PLUS

Optimization method for improving knowledge base image data caching processing

The invention discloses an optimization method for improving knowledge base image data caching processing, a platform can dynamically adjust a caching strategy according to a data access mode and access time change by utilizing a self-adaptive caching strategy combined with access time change and through replacement between two levels of caching strategies, so that the hit rate of caching is improved, and the caching efficiency is improved. According to the method, the access to a back-end database is reduced, the effective utilization of cache resources is ensured, and the platform can respond to the change of a user access mode in real time by dynamically adjusting a cache strategy, so that smoother and more efficient online transaction experience is provided; wherein the first-level cache stores frequently-used and frequently-accessed data, it is ensured that the data can be rapidly accessed, the time for obtaining the data from a rear-end database is shortened, the response speed of the system is improved, meanwhile, the second-level cache serves as a supplement of the first-level cache, more kinds of data are stored, and the coverage rate and efficiency of the cache are further improved.
Owner:WAN SHI ZHI DA KE JI (SU ZHOU) YOU XIAN GONG SI

Attribute-based neighborhood relation guided combined image retrieval method and system

The invention relates to a combined image retrieval method and system guided by a neighborhood relation based on attributes, and the method comprises the steps: reading training set data in batches, and carrying out the extraction of global features and local features of the training set data; splicing the local features and the global features to form original attribute features of the vision and the text; extracting attribute prototype features; constructing coherent prototype semantics; combining the attribute prototype characteristics of each element in the multi-modal query; for the combined attribute prototype features, modeling is carried out in combination with double relationships, and modeling is carried out on a pairwise relationship and a neighborhood relationship through attribute similarity, so that the measurement learning process is optimized; the COMBINER model generates a corresponding combination feature; similarity scores are solved, and descending order arrangement is carried out; and according to the actual demand, selecting the first K target images as a formal result set so as to complete the combined image retrieval. According to the method, the precision and robustness of combined image retrieval are improved.
Owner:SHANDONG UNIV +1

Relighting of outdoor images using machine learning

A media application provides, as input to a diffusion model, an initial image and a request to change a lighting in the initial image, wherein the initial image includes a subject and a sky. The media application outputs, with the diffusion model, an output image that satisfies the request. The media application determines, from the initial image, a sky segment and a subject segment. The media application generates a sky mask that corresponds to the sky segment and a subject mask that corresponds to the subject segment. The media application modifies a coloring of the initial image to match a coloring of the output image. The media application blends the modified initial image with the output image to form a blended image while using the subject mask to prevent modification to the subject from the modified initial image and the sky mask to prevent modification to the sky from the output image during the blending.
Owner:GOOGLE LLC

Mongolian intangible cultural heritage display system based on ai and virtual reality

The Mongolian intangible cultural heritage display system based on AI and virtual reality is provided, comprising a user intention analysis module, a behavior generation module, a pattern generation module, an exhibition content fusion module and a virtual reality display module.A unified expression and scheduling mechanism from semantic driving to exhibition execution is constructed, the key problems of disconnection between behavior and pattern generation, cultural mismatch of display timing, and difficulty in real-time feedback of user input in the existing system are solved, and the expression accuracy, generation flexibility and interactive response capability of intangible cultural heritage in a virtual environment are significantly improved.
Owner:INNER MONGOLIA UNIV OF TECH

Hashing image retrieval method based on anti-confusion factors

ActiveCN116910295BImage retrieval is accurateImage retrieval results are accurateStill image data indexingStill image data clustering/classificationAlgorithmTheoretical computer science
The present disclosure relates to an anti-confusion factor based hash image retrieval method, which comprises: obtaining a query image, the query image comprising a reason factor and a confusion factor; inputting the query image into a trained hash network to obtain a hash code expressing the reason factor in the query image; calculating the similarity between the hash code of the query image and hash codes of a plurality of historical images, and taking the historical image corresponding to the hash code with the highest similarity as the image retrieval result. The anti-confusion factor based hash image retrieval method provided by the present disclosure trains a hash network, so that the hash network ignores the confusion factor and focuses on expressing the reason factor in the image in the process of generating a hash code, thereby generating an accurate hash code. Then, by calculating the similarity between the hash codes, an accurate image retrieval result is retrieved.
Owner:CITY UNIV OF HONG KONG (DONGGUAN) (PREPARATORY)

Picture display method and device, equipment, storage medium and program product

The invention discloses a picture display method and device, equipment, a storage medium and a program product, and the method comprises the steps: carrying out the grouping of the picture information of a to-be-processed picture, obtaining an attribute grouping result, determining an attribute grouping value corresponding to each group, carrying out the matching of the attribute grouping value with a preset color mode, generating a corresponding color feature, and carrying out the processing of the to-be-processed picture. And finally, rendering the frame of the picture to be processed by using the color features, and displaying the picture after the frame is rendered. According to the method, the semantic information of the picture is converted into the visual color features, so that accurate association of the picture semantics and the color identification is realized, and the frame is not a decoration any more, but becomes an information carrier for bearing picture attributes. Besides, the picture list browsed by the user is accurately grouped, and the classification result is more intuitively presented in a frame rendering mode, so that the user is helped to quickly screen out the target picture.
Owner:CHINA MOBILE INTERNET CO LTD +1

A CAD drawing loading method and system combined with cache optimization

The application discloses a CAD drawing loading method and system combined with cache optimization, and relates to the technical field of graphic cache optimization.The method comprises the following steps: reading a target CAD drawing for identification; dynamically distributing tile data to multiple cache areas; performing real-time monitoring to generate a first monitoring data set, performing cache adjustment to generate a first cache parameter; performing data unloading to the disk cache area to generate a second cache parameter; performing interactive operation on the drawing to load the drawing in the memory cache area, and triggering the second cache parameter to perform asynchronous loading in the disk cache area according to the loading result.The application solves the technical problems of low cache management efficiency, long loading waiting time, unsmooth drawing display, and poor user interactive experience in the prior art CAD drawing loading, achieves efficient cache management of the CAD drawing loading, effectively reduces the loading waiting time, and improves the smoothness of the drawing display and the user interactive experience.
Owner:BEIJING GUANGLIANDA YUNTU DREAM TECH CO LTD

Computer architecture for interactive image based database searching

A computer architecture for interactive image based database searching includes a first server that acquires category information indicating a category of a vehicle owned by a user. The first server also acquires part information for identifying a vehicle part located at a position specified by a user operation within an image of a vehicle displayed on a terminal device of the user. A second server is in communication with the first server and searches for products by accessing a storage storing thereon product information on vehicle part products and by using the acquired category information and the acquired part information as search criteria. The system transmits search result information indicating results of the search to the terminal device.
Owner:RAKUTEN GROUP INC

File structured information extraction method, device, equipment, medium and product

The application discloses a kind of extraction methods, device, equipment, medium and product of file structured information, it is related to data processing technical field, comprising: determining the file content type of to-be-processed file;In the case where it is determined that file content type is image content file, to-be-processed file is carried out text recognition, and the text area coordinates of to-be-processed text contained in to-be-processed file and to-be-processed text in to-be-processed file are determined;To-be-processed text is carried out structured content entity identification, and the content entity coordinates of each structured content entity in text area coordinates respectively corresponding to structured content entity contained in to-be-processed text are determined;According to each content entity coordinates, the content entity relationship data between each structured content entity is constructed, and according to content entity relationship data, to-be-processed text is carried out structured information extraction, and the target structured information contained in to-be-processed file is obtained.The application can improve the accuracy and integrity of structured information extraction.
Owner:HANGZHOU HAOLINK INTELLIGENT TECHNOLOGY CO LTD +1

Natural disaster damage analysis

Systems and methods for natural disaster damage analysis. Aerial images are converted (710) into a set of property images. The property images are then clustered (720) using vector embedding. Clustering the property images includes classifying the property images into separate categories, determining a symmetric similarity between the property images in each category, and clustering the property images based on the similarity. A first artificial intelligence agent extracts (730) damage information from a set of representative members of the clustered property images, where the damage information includes damage level information and damage reasoning information. The damage information is then stored (740) in a datastore with at least two databases, where at least one of the databases stores damage reasoning information and at least one other one of the databases stores damage level information.
Owner:NEC LABORATORIES AMERICA INC

A method for combining an item and related equipment.

This application discloses a method and related equipment for matching items. This method applies artificial intelligence technology to the field of item search. The method includes: acquiring an image input by a user, the image containing a background and at least two items; based on the feature information of the image and the feature information of the at least two items, obtaining a target category that matches the image through a first neural network; and displaying the target items of the target category. This method not only allows searching for desired items by providing an image, but also, even when the input image is complex, it can still obtain items of a target category that matches the entire image, greatly expanding the application scenarios of this solution and improving user engagement.
Owner:HUAWEI TECH CO LTD

Sequence extraction using screenshot images

A system and method for sequence extraction using screenshot images to generate a robotic process automation workflow is disclosed. The system and method include capturing a plurality of screenshots of steps performed by a user on an application using a processor, storing the screenshots in memory, determining action clusters from the captured screenshots by randomly clustering actions into an arbitrary predefined number of clusters, wherein screenshots of different variations of a same action is labeled in the clusters, extracting a sequence from the clusters, and discarding consequent events on the screen from the clusters, and generating an automated workflow based on the extracted sequences.
Owner:UIPATH INC

Image processing apparatus and operating method thereof

An image processing apparatus for processing an image by using one neural network, includes: a memory storing one instruction; and one processor configured to execute the one instruction to: obtain first feature data, based on a first image, obtain pieces of second feature data corresponding to first areas of the first image by performing first image processing on the first feature data, the first areas comprising a first number of pixels, obtain third feature data, based on the first image, obtain pieces of fourth feature data corresponding to second areas of the first image, by performing second image processing on the third feature data, the second areas comprising a second number of pixels that is greater than the first number, and generate a second image, based on the pieces of second feature data and the pieces of fourth feature data.
Owner:SAMSUNG ELECTRONICS CO LTD