Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1028results about "Still image data clustering/classification" patented technology

Knowledge distillation method, electronic equipment and computer readable storage medium

The invention discloses a knowledge distillation method, electronic equipment and a computer readable storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the grouping of an attention map matrix of a teacher network based on the number of attention heads in an initial network structure of a student network, a high-dimensional attention space of a teacher network can be divided into subspaces with the same dimension as a student network, and attention map matrixes in groups are spliced, so that each group of spliced matrixes is aligned with an attention head of an initial network structure, a relatively regular corresponding relationship between the attention maps of the teacher network and the student network is realized, and the attention map splicing efficiency is improved. Therefore, the distillation loss can be calculated, and the knowledge of the complex teacher network can be accurately transmitted to the lightweight student network. Therefore, the problem that in a multi-head attention mechanism, the number difference between the teacher network and the student network hinders the student network to fully learn the teacher network knowledge can be solved, and the technical effects of improving the knowledge distillation effect and enhancing the student network performance are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Soft label-based noise robust text-to-image pedestrian retrieval method and device

The invention provides a noise robust text-to-image pedestrian retrieval method and device based on a soft label. The method comprises the following steps: calculating cosine similarity based on image global features and text global features, and normalizing the cosine similarity to generate a soft label representing image-text pairing confidence; distributing a sample weight for each training sample according to the soft label, and obtaining a joint weight for current iteration in combination with a dynamic weight factor progressively increased along with the training process; respectively constructing cross-modal contrast learning loss and similarity distribution matching loss by using the joint weight, and carrying out weighted summation on the cross-modal contrast learning loss and the similarity distribution matching loss to obtain a total loss function; and updating parameters of the image encoder and the text encoder by using the total loss function until training convergence, and obtaining a cross-modal alignment model. According to the method, robust cross-modal alignment can be realized, and the pedestrian retrieval accuracy in a noise scene is improved.
Owner:北京衔远有限公司 +1

Image processing model

A method comprising generating image descriptions of images in an original training set of images; determining, using at least one LLM, at least one domain and / or class which is under-represented in the original training set; generating, using a second LLM and based on the determination of the at least one domain and / or class, at least one instruction for a third LLM to generate at least one text prompt; generating, using the third LLM and based on the at least one instruction, the at least one text prompt for a text-to-image model; generating, using the text-to-image model and based on the at least one text prompt, at least one synthetic image; and generating an enhanced training set of images for use in training an image processing machine learning, ML, model, the enhanced training set of images comprising the original training set of images and the at least one synthetic image.
Owner:FUJITSU LTD +1

Multi-modal power grid fault diagnosis method and system based on causal event atlas

The invention discloses a multi-modal power grid fault diagnosis method and system based on a causal event atlas, and belongs to the technical field of intelligent operation and maintenance of power systems. The method comprises the following steps: preprocessing historical fault case data of power grid equipment, and constructing a causal event atlas database; when a fault diagnosis request is received, analyzing the fault diagnosis request by the planning agent, generating an initial fault hypothesis set in combination with power field knowledge, endowing a corresponding credibility score to the fault hypothesis, retrieving the cause subgraph as evidence in an iterative loop, and updating the credibility score of the fault hypothesis by the reasoning agent; and generating a multi-modal diagnosis report after the termination condition is met. According to the method, the problems of insufficient causal modeling and poor interpretability of a traditional method are solved, and the diagnosis accuracy, efficiency and user credibility are remarkably improved.
Owner:STATE GRID JIANGXI ELECTRIC POWER CO LTD RES INST +1

Three-area three-line database automatic quality inspection evaluation method based on deep learning

The invention discloses a three-area three-line database automatic quality inspection evaluation method based on deep learning. The method comprises the following steps of S1, obtaining three-area three-line database data and performing preprocessing; s2, checking a directory structure and a file name of the original data set; s3, extracting spatial layer data from the hierarchical data set; s4, performing segmentation and attribute information extraction on various pattern spots in the layer data after spatial verification; s5, inputting the pattern spot relation characteristic data into an artificial fish swarm algorithm; s6, automatically comparing the pattern spot area and distribution data with a quality inspection standard preset in a three-area three-line database; and S7, performing parallel processing on an automatic comparison result through a multi-thread thread pool. According to the method, an artificial fish swarm algorithm, a relational attention network and a task allocation and thread pool management technology are combined, and three-area three-line database automatic quality inspection evaluation based on deep learning is realized.
Owner:HEILONGJIANG PROVINCIAL LAND & SPACE PLANNING RES INST

Multimodal ai-based search for digital assets

Embodiments of the present disclosure relate to multimodal AI-based search for digital assets via an indexing and / or search pipeline. With respect to the indexing pipeline, some embodiments obtain first data and second data associated with a first digital asset. Such data represents different data types or modalities of the same digital asset. After obtaining the first and second data, some embodiments then generate a composite index. After the composite index is built such index can then be used to execute a query via the search pipeline. To execute the query some embodiments compute a relevance score for each digital asset, of multiple digital assets, based at least in part on a measure in which each digital asset satisfies one or more parameters or conditions for two or more data types of the query. Various embodiments then rank each digital asset and present one or more associated indicators.
Owner:NVIDIA CORP

Multi-modal picture data processing method and device, equipment and storage medium

The invention relates to the technical field of computer vision and multi-modal data processing, and discloses a multi-modal picture data processing method and device, equipment and a storage medium, and the method comprises the steps: carrying out the multi-dimensional feature extraction of a pre-obtained standardized picture set, and obtaining composite feature data; inputting the composite feature data and a preset cue word into a multi-modal large model to generate a natural language description text, and extracting key semantic information from the natural language description text to form a structured semantic set; and performing associative storage on the standardized picture set, the composite feature data and the structured semantic set, and constructing a multi-modal index based on associative storage data. According to the method, deep understanding of multi-modal picture data is realized based on multi-dimensional feature extraction and semantic understanding of a multi-modal large model.
Owner:SHENZHEN MATRIX ORIGIN TECH CO LTD

Multi-mode image searching method based on images

The invention discloses a multi-modal image searching method, and solves the problems that in the prior art, effective utilization of document image text information is lacked, and the multi-modal retrieval requirement in a complex scene is difficult to meet. According to the invention, the input image is classified by using the deep learning model, and the image classification result is obtained; the image classification result is a text image or a non-text image; generating a multi-dimensional composite feature vector for the text image; for a non-text image, generating a visual feature vector; retrieving a text image based on the multi-dimensional composite feature vector, calculating a comprehensive score by fusing the text semantic similarity, the abstract semantic similarity and the format visual similarity, and returning a retrieval result according to the comprehensive score; and for a non-text image, performing retrieval based on the cosine similarity of the visual feature vector, and returning a retrieval result.
Owner:BEIJING YINGYAN CHUANGXIN TECH DEV CO LTD

Method and system performing pattern clustering

A method of clustering patterns of an integrated circuit includes; providing a pattern image and numeric data, as input data corresponding to a first pattern to a first model, wherein the first model is trained by a plurality of sample images and a plurality of sample values, obtaining a content latent variable using the first model, and grouping a plurality of content latent variables corresponding to a plurality of patterns into a plurality of clusters based on a Euclidean distance, wherein the numeric data represents at least one attribute of the first pattern.
Owner:SAMSUNG ELECTRONICS CO LTD

Building engineering crack detection method and system based on image recognition

The embodiment of the invention discloses a building engineering crack detection method and system based on image recognition, and the method comprises the steps: obtaining a building surface image, carrying out the preprocessing of the image, obtaining a standardized image, and carrying out the multi-scale decomposition extraction and integration of various features, and forming a multi-dimensional feature descriptor set; after feature importance is evaluated, a compact feature vector is generated through dimension reduction, quantization coding and compression, and then a multi-level feature index mechanism for optimized compression is constructed. A query feature vector is extracted from a newly collected image, searching and screening are completed by means of an index mechanism and a tolerance threshold, and a crack matching result is obtained; and based on the result, positioning cracks, classifying types, measuring parameters and evaluating severity, and generating a crack state report. The crack trend is analyzed in combination with the historical data time sequence, a multi-stage early warning mechanism is designed, maintenance suggestions are provided, and a real-time monitoring and early warning system is formed. According to the embodiment of the invention, the technical problems of high storage pressure and low real-time detection efficiency in the prior art can be effectively solved.
Owner:内江市住房保障和房地产事务中心

Photo content clustering for digital picture frame display and automated frame storytelling

A method and system for automated routing of pictures taken on mobile electronic devices to a digital picture frame including a camera, microphone, and speaker integrated with the frame, and a network connection module allowing the frame for direct contact and upload of photos from electronic devices or from photo collections of community members. Clustering photos by content is used to improve display and to respond to photo viewer desires. Trends or patterns can be detected from the photo collections and that information used for various purposes beyond photo display. The frame includes a conversational intelligence that provides a verbal communication with a viewer, such as for determining an identity or preferences of the frame viewer, determining photos to display for the viewer, discussing displayed photos with the viewer, or telling stories or life histories to the viewer based upon photo content.
Owner:PUSHD INC

Crystal pixel lookup table generation method and device, computer equipment and storage medium

The invention relates to a crystal pixel lookup table generation method and device, computer equipment and a storage medium. Determining a preset number of first pixel peak points from the first crystal pixel lookup table, performing data enhancement processing on the first pixel peak points in the first crystal pixel lookup table to obtain a second crystal pixel lookup table, and performing data enhancement processing on the first pixel peak points in the first crystal pixel lookup table to obtain a second crystal pixel lookup table, and obtaining a target crystal pixel lookup table based on the first crystal pixel lookup table and the second crystal pixel lookup table. According to the embodiment of the invention, the second crystal pixel lookup table is obtained by performing data enhancement processing on the first pixel peak point, and the phenomenon that the crystal of the medical equipment disappears along with the use time is simulated, so that the target crystal pixel lookup table is obtained based on the first crystal pixel lookup table and the second crystal pixel lookup table; the number and diversity of crystal pixel lookup table data sets are improved, and the complexity of crystal pixel lookup table generation is reduced.
Owner:SHANGHAI UNITED IMAGING HEALTHCARE

Cultural symbol transmission method and system based on Sichuan opera facial makeup art patterns

The invention provides a culture symbol transmission method and system based on Sichuan opera facial makeup art patterns, and particularly relates to the field of facial makeup digital transmission. According to the method, multi-level modeling of a Sichuan opera facial makeup pattern from a pixel layer to a component layer and then to a semantic layer is realized by fusing image processing, component contour matching, image structure modeling and an image neural network technology; according to the method, main and auxiliary color block areas and component information in an image can be automatically identified and structurally stored, a spatial semantic relationship of the image is learned through a graph neural network, a structural semantic code is output, and finally a multi-dimensional cultural label index capable of being filed, clustered and classified is generated; the method effectively overcomes the limitation that in the prior art, only stays at an image level and is lack of structure and semantic modeling, enables the Sichuan opera facial makeup pattern to have machine-readable component structure representation and semantic feature representation, and lays a foundation for cultural inheritance and intelligent propagation.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Binocular vision-based AI intelligent shooting system for cultural relic exhibition hall

The invention discloses a cultural relic exhibition hall AI intelligent shooting system based on binocular vision, and relates to the technical field of cultural relic digitization and interaction, and the system comprises a binocular vision image collection module which is used for shooting and storing a multi-view high-definition image of a cultural relic; the AI image splicing fusion module preprocesses the image and generates a panorama; the cultural relic three-dimensional reconstruction module constructs a three-dimensional model with fine textures; real-time fusion of the intelligent shooting module fusion model and a visitor image is carried out to generate an interactive picture; the interactive output and management module provides related functions and manages data; by integrating binocular vision image acquisition, AI image splicing and fusion, cultural relic three-dimensional reconstruction and real-time fusion intelligent shooting technology modules, intelligent upgrading of cultural relic display and visitor interaction is achieved, the system adopts a binocular camera shooting mechanism triggered by multiple strategies, timing and induction triggering modes are combined, and the intelligent upgrading of cultural relic display and visitor interaction is achieved. The method can flexibly adapt to the demands of different exhibition scenes, guarantees the high efficiency of cultural relic image collection, and avoids the waste of resources.
Owner:WEIMAI TECH CO LTD

Picture display method, electronic equipment and computer readable storage medium

The embodiment of the invention discloses a picture display method, electronic equipment and a computer readable storage medium, in the method, the electronic equipment responds to a target operation used for uploading a picture, and target information related to interface content of an application program is obtained; according to the first picture number and the target information, a target candidate picture set of each target element is obtained from the picture set, the target elements are determined according to the target information, the target candidate picture set comprises at least one target candidate picture, and the total number of the target candidate pictures of each target element is smaller than or equal to the first picture number; the first picture number is the number of pictures capable of being displayed on a first screen interface of the picture selection interface; and displaying the target candidate picture of each target element on the first screen interface. Thus, the user can obtain the pictures of the multiple target elements on the first screen interface, the pictures of the target elements do not need to be obtained through page turning, and the convenience of selecting the pictures by the user is improved.
Owner:HUAWEI TECH CO LTD

Museum exhibit identification method and device, electronic equipment and storage medium

The invention discloses a museum exhibit identification method and device, electronic equipment and a storage medium. The method comprises the steps that a target image is acquired, and the target image comprises an exhibit to be recognized; single-subject detection is carried out on the target image, a subject area image in the target image is obtained, and the subject area image is the part, containing the to-be-recognized exhibit, in the target image; performing feature extraction on the target image to obtain global features, and performing feature extraction on the main region image to obtain local features; according to the global features and the local features, matching retrieval is carried out in an exhibit feature library to obtain exhibit information corresponding to the target image, the exhibit feature library comprises feature vectors of a plurality of exhibits, and the feature vectors are determined according to the global features and the local features of the exhibits. The museum exhibit identification method solves the technical problem of poor identification precision of a museum exhibit identification method in related technologies.
Owner:CHINA TELECOM CORP LTD +1

Image processing model training method and device, image recommendation method and device, equipment and medium

The invention provides an image processing model training method and device, an image recommendation method and device, equipment and a medium, relates to the technical field of artificial intelligence, in particular to the technical fields of natural language processing, deep learning and the like, and can be used for application scenes such as generative retrieval, document intelligent editing, intelligent assistants, virtual assistants, intelligent e-commerce and the like. The training method comprises the steps of obtaining a first sample image with annotation data and a plurality of unannotated second sample images, wherein the annotation data comprises a classification result and an interpretation text for a plurality of preset evaluation dimensions; performing fine tuning training on the multi-modal visual understanding large model by using the first sample image and the corresponding annotation data; classifying the plurality of second sample images by using the multi-modal visual understanding large model subjected to fine tuning training; constructing a plurality of image groups, wherein each image group comprises at least two images with different categories in the same evaluation dimension; and training the lightweight image processing model based on the plurality of image groups.
Owner:BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD

Training method and application method of image recognition retrieval model, equipment and medium

The invention discloses a training method and an application method of an image recognition retrieval model, equipment and a medium, and belongs to the field of image recognition. The method comprises the following steps: extracting visual modal features of a known category; obtaining semantic modal information of the seen and unseen categories, wherein the semantic modal information comprises an attribute modal, a category name modal and a semantic description modal; fusing the semantic modal information to obtain relevance semantic representation; training a generative adversarial network based on the visual modal features of the known categories and the semantic representation, and generating pseudo-visual modal features of the unseen categories; and training at least one of a classification model and a retrieval model based on the pseudo-visual modal features corresponding to the unseen categories, so that the classification model at least identifies the image samples of the unseen categories, and / or the retrieval model at least retrieves retrieval results matched with the unseen categories. According to the method, a category separation optimization mechanism is introduced, so that the generalization ability and the recognition accuracy of the model under the zero sample condition are improved.
Owner:GRG BANKING EQUIPMENT CO LTD

Classification using multi-modal large language models

Methods, systems, and devices for classification. In one aspect, a method includes receiving an input and a request to classify the input into one of a plurality of categories, processing the input using a multi-modal model to generate (i) a description of the input and (ii) a category prediction, a description of the input and the category prediction are processed using a text encoder embedding neural network to generate (i) a text description feature embedding and (ii) a predictive feature embedding, a query feature embedding representing the input being generated from at least the description feature embedding and the predictive feature embedding, and classifying the input into one of a plurality of categories using the query embedding.
Owner:GOOGLE LLC

Text-to-image pedestrian retrieval method and device based on mask denoising and medium

The invention provides a text-to-image pedestrian retrieval method and device based on mask denoising and a medium. The method comprises the following steps: respectively executing mask and similar word random replacement on entity words and attribute words according to a set probability, and generating a training text replaced by a mask; inputting the text feature vector and the image feature vector into a cross-modal interaction encoder to obtain fusion feature representation; predicting the original words at the masked positions based on the fused feature representation, and calculating mask prediction loss; calculating image-text contrast learning loss based on a similarity relationship between the text feature vector and the image feature vector; and extracting features of the query text and the to-be-retrieved pedestrian image library by using the pedestrian retrieval model, calculating the similarity between the features of the query text and the features of the to-be-retrieved pedestrian images, and generating a sorting result so as to output a target pedestrian image matched with the query text. According to the method, the robustness of visual semantic alignment in a noise scene can be improved, and the pedestrian retrieval accuracy is remarkably improved.
Owner:北京衔远有限公司 +1

Celestial body data retrieval method and device, storage medium and electronic equipment

The invention discloses a celestial body data retrieval method and device, a storage medium and electronic equipment, and the method comprises the steps: matching the coordinate information of each celestial body image in a celestial body image header file with the coordinate information contained in each celestial body star catalogue record in a star catalogue database of each different data source according to the coordinate similarity; and determining each celestial body data pair formed by the celestial body image and the celestial body star catalogue record which have a matching relationship, and for each celestial body data pair, carrying out feature coding on the celestial body image and the celestial body star catalogue record in the celestial body data pair to obtain the multi-modal features of the celestial body data pair. A multi-modal feature is constructed by associating a celestial body image in a celestial body image header file with celestial body star catalog records in a star catalog database. And when celestial body data retrieval needs to be carried out, obtaining a retrieval result according to the multi-modal features. According to the method, the similarity of data features in the celestial body image and the celestial body star catalogue record is considered at the same time, and more accurate retrieval precision can be achieved.
Owner:ZHEJIANG LAB

Fruit and vegetable surface state recognition method based on computer vision

The invention discloses a fruit and vegetable surface state recognition method based on computer vision, belongs to the technical field of fruit and vegetable storage state monitoring, and aims at solving the problems that in traditional fruit and vegetable surface state recognition, subjectivity is high, efficiency is low, and abnormal changes are difficult to find in time. The method comprises the steps that fruit and vegetable surface images are collected in a triggering mode and subjected to standardized preprocessing; extracting foreground and background images of fruits and vegetables according to background features of the storage space, and dividing independent fruit and vegetable areas by combining fruit and vegetable types; feature points are extracted to compare and match independent fruit and vegetable areas in the continuous images, newly added, updated and matched areas are identified, and fruit and vegetable change and maintenance areas are divided; establishing a classification evaluation database for the fruit and vegetable change region, and constructing a region index set; judging the state of the fruit and vegetable maintaining region through similarity, and updating or continuing to use the region index set; and performing abnormal trend analysis and state early warning according to the regional index set. Automatic and precise monitoring and early warning of the surface states of the fruits and vegetables are achieved.
Owner:杭州道秾科技有限公司

Data processing method and device based on smart city

The invention provides a data processing method and device based on a smart city, and relates to the technical field of data processing.The method comprises the steps that a city-level spatial data base is constructed through oblique photography of an unmanned aerial vehicle, and a unified and real spatial foundation is formed by generating a live-action three-dimensional model, a digital orthoimage and a digital elevation model; on this basis, a building base and a building white model are extracted, spatial logic verification is carried out on multi-source business basic data, a standard address system stably associated with a building three-dimensional entity is further constructed, precise spatial anchoring of governance objects such as population, houses and units is achieved, and the method is suitable for mass production. And finally, the associated data is uniformly converged to a city information model platform, and reliable and updatable spatial data support is provided for smart city application. According to the invention, the timeliness of smart city data processing in a dynamic earth surface deformation scene can be improved.
Owner:湖北省国土测绘院

Image retrieval method and apparatus, and storage medium and electronic device

The present application relates to the technical field of computers. Disclosed are an image retrieval method and apparatus, and a storage medium and an electronic device. The method comprises: using a preset deep hashing model to process an image to be retrieved, so as to obtain an image hash code, wherein the preset deep hashing model is obtained by means of performing incremental learning training on the basis of a plurality of preset class centers and query data; and on the basis of a preset image corresponding to a preset hash code matching the image hash code, obtaining a matched image. The present application effectively improves the accuracy of image retrieval.
Owner:SHENZHEN TCL NEW-TECH CO LTD

Classification using multimodal large language models

Methods, systems, and apparatus for classification. In one aspect, a method includes receiving an input and a request to classify the input into one of a plurality of classes, processing the input using a multimodal model to generate (i) a description of the input and (ii) a class prediction, processing the description of the input and the class prediction using a text encoder embedding neural network to generate a (i) text description feature embedding and (ii) a prediction feature embedding, generating, from at least the description feature embedding and the prediction feature embedding, a query feature embedding representing the input, and classifying the input into one of the plurality of classes using the query embedding.
Owner:GOOGLE LLC

Object and action recognition via text-based classification models

Example implementations include a method, apparatus and computer-readable medium of object / action recognition using a text-based classification model, comprising generating an image vector configured to represent one or more features of the first image. Additionally, the implementations further include computing a vector distance between the image vector and each of a first text vector and a second text vector, wherein the first text vector is configured to represent a first text, and wherein the second text vector is configured to represent a second text. Additionally, the implementations further include classifying the first image according to the first text or the second text based on which computed vector distance indicates a highest similarity between the image vector and either the first text vector or the second text vector relative to the other of the first text vector or the second text vector.
Owner:TYCO FIRE & SECURITY GMBH

Image processing method and apparatus, device, and medium

In an image processing method, a reference library and a query library are obtained; a reference image in the reference library and a prompt are inputted into a diffusion model to obtain estimated noise; the estimated noise is merged to obtain a reference noise feature; a plurality of query noise features corresponding to a query image are determined; and a target label corresponding to the query image is determined based on feature similarities between the plurality of query noise features and the reference noise features.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Image retrieval method and device based on feature fusion, equipment and medium

The invention relates to the technical field of image processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an image retrieval method, device and equipment based on feature fusion and a medium, and the method comprises the following steps: receiving a to-be-processed image, and extracting basic features through a convolutional neural network; respectively processing the basic features through a global feature branch module and a local feature branch module, and generating an image overall representation vector and an image region detail vector; fusing the two feature vectors to generate a fused feature vector; and generating an image retrieval identifier based on the fused feature vector, establishing a corresponding relationship between the image retrieval identifier and the to-be-processed image, and when a retrieval request is received, querying the corresponding image based on querying the image retrieval identifier and the corresponding relationship. According to the method, the image features are efficiently extracted through the single-stage network architecture, the unique index is generated, high calculation overhead and low efficiency of a traditional double-stage feature extraction method are avoided, and the image retrieval speed and accuracy are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Big model-based document content customized extraction method and system

The invention relates to a document content customized extraction method and system based on a large model. The method comprises the following steps: inputting a document, and dividing the document into a text document or an image document; formulating a structured cue word according to the document and user requirements; calling a large model to process the document, and performing content extraction based on the structured cue word to generate a preliminary extraction result; comparing the preliminary extraction result with the extraction target, updating the structured cue word for iterative optimization, and recalling the large model until the extraction result meets a precision threshold value; and confirming that the extraction result meeting the precision threshold meets an extraction target and storing the extraction result as a final extraction result, and storing the final extraction result as an output format. According to the method, efficient extraction of multi-modal document content information is achieved through large model calling, customized extraction is achieved through structured cue words, the output precision is improved through an interactive process, a large amount of time cost is saved for a user, and remarkable achievement benefits are brought.
Owner:INNOVATION ACAD FOR MICROSATELLITES OF CAS +1