Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

328results about "Multimedia data indexing" patented technology

Multi-layered, multi-pathed apparatus, system, and method of using cognoscible computing engine (CCE) for automatic decisioning on sensitive, confidential and personal data

A computer-implemented apparatus, system, and method is disclosed for protecting sensitive data. A cognoscible computing engine is multi-layered and multi-pathed. It includes features for handling different data formats, including structured, semi-structured, and unstructured data. Features are included to support near real-time processing at scale with high accuracy. Applications include redacting or masking sensitive data to comply with data privacy and security standards.
Owner:DATA SAFEGUARD INC

Meeting content intelligent generation processing method and system based on multi-modal large model

The invention discloses a conference content intelligent generation processing method and system based on a multi-modal large model, and the method comprises the steps: collecting the original data of a conference, and completing the standardization preprocessing; inputting a multi-modal large model, extracting multi-modal features and carrying out semantic alignment; executing cross-modal hash coding, generating binary codes and establishing an index database; performing hash retrieval on the related fragments, and constructing a conference content directed graph; based on a conference content directed graph structure, searching an optimized path by adopting a Monte Carlo tree; and generating structured conference content, and outputting a summary, an abstract and an action item. According to the method, efficient extraction, accurate retrieval and structured intelligent generation of the conference content are realized by fusing a multi-modal large model, cross-modal Hash coding and Monte Carlo tree search.
Owner:NANJING WEITEXI NETWORK SCI & TECH

Multi-modality-based minority non-abandoned pattern knowledge graph construction method and multi-modality-based minority non-abandoned pattern knowledge graph construction system

The invention discloses a multi-modality-based minority non-abandoned pattern knowledge graph construction method and system, and relates to the technical field of cultural heritage digital protection. Deep association of multi-modality knowledge is realized through a double-path entity relationship extraction mechanism, context semantics are coded by a text path by utilizing a pre-training language model, and a multi-modality knowledge graph is constructed; the method comprises the following steps: accurately extracting entities such as a pattern and an inheritor, a semantic relationship and a visual path, analyzing a pattern topological structure through a graph convolutional network, converting visual features such as lines and contours into structured relationship data, calculating cosine similarity of a text and a visual feature vector through comparative learning, establishing cross-modal mapping of the visual features and culture description, and obtaining a visual feature model; according to the mechanism, the knowledge graph simultaneously contains semantic logic and visual feature association, and construction of a complete knowledge chain from a pattern form to cultural connotation is realized.
Owner:NORTHEAST FORESTRY UNIV

Lake and Hunan woodcarving image generation method, device and equipment based on LoRA model and storage medium

The invention discloses a Lake and Hunan wood carving image generation method, device and equipment based on a LoRA model and a storage medium, and relates to the technical field of process digitization and image generation, and the method comprises the steps: constructing a Lake and Hunan wood carving manufacturing process feature library based on a material object scanning graph, wood texture data and a manufacturing process video of the Lake and Hunan wood carving; according to the feature library, performing hierarchical training on the initial LoRA model according to a texture layer-cutter layer-pattern layer hierarchical logic to obtain a hierarchical LoRA model, and performing hierarchical fusion on the hierarchical LoRA model and a potential diffusion model to form a lake and Hunan woodcarving image generation model; and analyzing text cue words input by a user by using a keyword system in the feature library, extracting wood carving types, timber texture parameters, folk pattern requirements and scene adaptation features, and inputting the features into the lake and Hunan wood carving image generation model to obtain a target lake and Hunan wood carving image. According to the method, the lake and Hunan woodcarving image with the process reduction degree and the scene adaptability can be generated.
Owner:HUNAN VOCATIONAL COLLEGE OF SCI & TECH

Network model training method, data processing method, and apparatus

The present disclosure provides a network model training method, a data processing method, and an apparatus. The network model training method comprises: acquiring target sample data, wherein the target sample data comprises text sample data and image sample data; inputting the target sample data into a network model to be trained to obtain a sample recognition result; and adjusting a parameter of a text encoder on the basis of a text recognition result and first supervision data corresponding to the text recognition result, adjusting a parameter of an image encoder on the basis of an image recognition result and second supervision data corresponding to the image recognition result, and a hybrid image-text recognition result and third supervision data corresponding to the hybrid image-text recognition result, and adjusting a parameter of a hybrid encoder on the basis of the hybrid image-text recognition result and the third supervision data corresponding to the hybrid image-text recognition result to obtain the trained network model formed by the text encoder, the image encoder, and the hybrid encoder.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Multi-dimensional knowledge base construction and query method based on multi-modal large model

The invention relates to a multi-dimensional knowledge base construction and query method based on a multi-modal large model, and belongs to the technical field of AI large model application. The method comprises the following steps of: analyzing a text file containing text, table and picture contents; selecting a pre-trained large language model and a multi-modal large model, generating text abstracts and picture abstracts by utilizing the models, and enabling storage paths of the picture abstracts to correspond to storage paths of pictures one by one; selecting a text embedding model to construct a vector database and a retriever; and constructing a retrieval chain based on the multi-modal large model. According to the method, construction from a single text knowledge base to a multi-dimensional knowledge base is achieved, a multi-dimensional information retrieval method is provided, the problem of construction of the multi-dimensional knowledge base based on the vectorization technology is solved to a certain extent, and the large model output quality based on the retrieval enhancement generation technology is remarkably improved.
Owner:HONGHE POWER SUPPLY BUREAU OF YUNNAN POWER GRID

Digital media content element accurate screening method based on artificial intelligence image recognition

The invention discloses a digital media content element accurate screening method based on artificial intelligence image recognition, and relates to the technical field of digital media, and the method comprises the steps: building a distributed capture network to form a dynamic content pool, and building a metadata index database; calling a multi-modal image perception engine to generate a double-layer characteristic spectrum containing dominant and recessive elements; constructing a distributed recognition model cluster based on federated learning; converting the user demand into a screening parameter set and generating a decision tree; screening and secondarily verifying an output result through a double-path matching mechanism; and constructing a reinforcement learning reward function based on user behaviors, and driving the model and the decision tree to co-evolve. Multi-source heterogeneous content full-dimension analysis is achieved, the recognition comprehensiveness and depth are improved, knowledge barriers and privacy risks are solved, the screening accuracy and flexibility are improved, the system is endowed with the continuous optimization capacity, and the method is suitable for efficient and accurate digital media content screening scenes.
Owner:XIAMEN HUAXIA UNIV

Aggregation, Organization, Branding, Stake and Mining of Image, Video and Digital Rights

Systems and immersive methods are provided for aggregation and organization of image, video and digital rights data. In an exemplary embodiment, a system and method for aggregation and organization of media data acquired from a plurality of sources is provided. The system and method include a data collection element, a brand commercial (POPmercial) placement for sponsorship element, a stake and mining rewards (Futures In Popular) against placed media element and a user interface that allows publishing of a unique curated media stream as a first-time-ever NFT BROADCAST, (POPcast).
Owner:REY REY +1

Multi-modal data verification system and method in medical scientific research

The invention belongs to the field of medical data processing. The invention provides a multi-modal data verification system and method in medical scientific research. The method comprises the following steps: extracting multi-modal data from different storage media; respectively analyzing and processing the extracted multi-modal data, respectively establishing a feature data model of each modal data and establishing a feature index; mapping the feature data models of the modal data to a unified vector space, performing data fusion according to the relationship of feature indexes of the feature data models, and generating and storing a wide model with multi-dimensional indexes; and according to a verification rule in a data processing process, slicing the wide model according to the feature index, and extracting a data feature information part corresponding to the verification rule to complete data verification. The method has the beneficial effects that all data of different modalities are subjected to homogenized fusion, so that verification for data of the same modality is not needed any more, and the verification result is more accurate.
Owner:成都华唯科技股份有限公司

Cross-modal conference information association retrieval method and system and medium

The invention discloses a cross-modal conference information association retrieval method and system and a medium, and relates to the technical field of artificial intelligence, and the method comprises the following steps: carrying out feature extraction on obtained multi-source heterogeneous data to obtain multi-modal features, and uniformly mapping the multi-modal features to a first feature space of a preset dimension; in the first feature space, cross-modal deep fusion processing is performed on the multi-modal features, and a joint embedding space with consistent semantics is constructed according to the cross-modal deep fusion processing; constructing a vector index database based on a multi-modal feature vector in the joint embedding space, receiving a natural language query and mapping the natural language query to the joint embedding space, executing two-stage retrieval, and then obtaining a semantic fusion score based on calculated semantic fusion scores; multiplying a time sequence reward value based on the query time deviation and a dynamic reward value based on the core word matching degree to obtain a dynamic fusion score, and performing fusion sorting on the candidate set to obtain a final sorting result and an associated retrieval result; according to the method, refined sorting of the retrieval results is realized.
Owner:UNIV OF SCI & TECH OF CHINA

Task processing method and intelligent device with body

The invention discloses a task processing method and an intelligent device, and relates to the technical field of intelligent devices, and the method comprises the steps: obtaining multi-modal data which comprises task information of a to-be-processed task in a real-time interaction scene of a physical entity and an environment; performing feature extraction and feature alignment processing on different modal data in the multi-modal data to obtain a first feature after feature alignment; performing cross-modal data retrieval on a knowledge base based on the first feature to obtain target retrieval data; performing feature fusion on different modal data features in the first features to obtain second features; and making a decision based on the target retrieval data and the second feature by using a decision model to obtain an action instruction sequence, and executing the to-be-processed task based on the action instruction sequence.
Owner:LENOVO (BEIJING) LTD

Digital asset intelligent analysis platform based on block chain intelligent contract technology

The invention relates to the technical field of digital asset management analysis, in particular to a digital asset intelligent analysis platform based on a block chain intelligent contract technology. The computing power resource state coupling module is used for obtaining the saturation degree of computing resources by capturing information of a task control block and combining the load rate of a processor, and embedding and collecting mapping data streams to generate a management and control data packet; the abnormal data screening module is used for executing local serial processing when the saturation degree of the computing resources is lower than an I / O throughput threshold value, and otherwise, hardware acceleration is carried out, and a to-be-verified digital sequence is constructed; generating an abrupt change collection feature vector based on the to-be-verified digital sequence; and the component resume anchoring module is used for carrying out data binding operation on the abrupt change collection feature vector, generating a single-piece digital traceability certificate, carrying out aggregation compression, obtaining a batch verification root value and outputting a quality signal digital right. According to the method, a load awareness and distributed storage partitioning strategy is constructed through indexes, and the I / O frequency and transmission bandwidth occupation of whole library retrieval are reduced.
Owner:ZHEJIANG CULTURAL PROPERTY RIGHTS EXCHANGE CO LTD

Intelligent processing method and system for new media data

The application relates to the field of information technology, in particular to an intelligent processing method and system for new media data. The method comprises the following steps: reading historical new media material data; analyzing the historical new media material data, calculating an availability estimation value of the historical new media material, and sorting and screening the historical new media resource according to the availability estimation value; storing the analyzed historical new media material data in a classified manner, and establishing a multidimensional index of the historical new media material data; receiving a new media content generation demand, converting the new media content generation demand into a new media material matching condition; screening and sorting new media materials based on the new media material matching condition, and selecting the first N new media material data as basic new media material data; and fusing a theme content based on the basic new media material, generating multi-form new media content, and realizing intelligent processing of high-efficiency, high-quality and self-optimizable new media content.
Owner:HANGZHOU XIAOLU CORGI NETWORK TECHNOLOGY CO LTD

Privacy Controls for Sharing Embeddings for Searching and Indexing Media Content

This document describes techniques and systems that enable privacy controls for sharing embeddings for searching and indexing media content. A set of images of a user's face are obtained and a machine-learned model is applied to the set of images to generate a user-specific dataset of face embeddings for the user. Media content stored in a media storage is indexed by applying the machine-learned model to the media content to provide indexed media information identifying one or more faces shown in the media content. Access to the indexed media information by another user querying the media content for images or videos depicting the user is controlled based on a digital key shared by the user with the other user, where the digital key is associated with the user-specific dataset and the user-specific dataset is usable to identify the images or videos depicting the user.
Owner:GOOGLE LLC

Method and System for Real-Time Collaboration, Task Linking, and Code Design and Maintenance in Software Development

A method for automated document processing and task assignment including receiving an input document from a document source, performing a content analysis on the input document, extracting metadata from the input document, identifying an identified task type to be performed, generating standardized metadata by converting the metadata into a standardized JSON format, storing the standardized metadata in a database, determining a user assignment for the identified task type, the including an assigned user, generating an action item including the identified task type, the user assignment, and the standardized metadata, and adding the action item to a task management system for processing by the assigned user.
Owner:MADISETTI VIJAY

Visual compression and retrieval method and device of document, equipment and storage medium

The invention discloses a document visual compression and retrieval method and device, equipment and a storage medium, and relates to the technical field of computers, the method comprises the following steps: obtaining a to-be-processed document page image, segmenting the image into a plurality of image blocks, and determining the structure category and the structure importance score of each image block; obtaining a plurality of structure regions based on structure category and spatial position aggregation, and distributing a preset number of compression tokens for each region in combination with structure category weights and importance scores; generating structure anchor point tokens corresponding to the regions by the compressed tokens to form a set; receiving a query request, converting the query request into a query vector, and performing retrieval in the anchor point token set to obtain a target structure region; and performing local decoding reconstruction based on the compressed token of the target region, and outputting a region image or a structure mask. According to the method, through structure-guided self-adaptive compression and fine-grained retrieval, the long document processing efficiency is greatly improved, and the compression effect and the retrieval accuracy are both considered.
Owner:BEIJING DIGITAL CHINA CLOUD COMPUTING CO LTD

Multi-source heterogeneous data and big language driven process intelligent design system and method

The invention relates to the technical field of artificial intelligence and intelligent manufacturing, and discloses a multi-source heterogeneous data and large language driven process intelligent design system and method, and the system comprises a multi-modal data processing module which is used for achieving the structural information extraction of multi-source heterogeneous engineering data; the multi-source heterogeneous mixed database construction module is used for realizing storage of multi-source process knowledge and joint query of a database; the large model construction and optimization module is used for constructing a large model to realize user intention understanding, process knowledge fusion and process scheme generation and optimization; the intelligent processing technology collaborative design system provides a multi-mode interaction interface for a user to realize full-process automatic technology design from demand input to scheme output; according to the system and the method, the full-process intelligence from part information extraction, process knowledge retrieval, process scheme generation to simulation verification optimization is realized, and the system and the method can be widely applied to machining process design and optimization scenes in the fields of industrial mother machines, robots, high-end equipment manufacturing and the like.
Owner:NINGXIA UNIVERSITY

Flight accident scene-oriented intelligent civil aircraft separated emergency flight data storage system

The invention provides an intelligent civil aircraft separated emergency flight data storage system facing flight accident scenes. The system collects multi-source data such as flight parameters, audios, videos and data links, carries out preprocessing and semantic modeling on the multi-source data, and extracts multi-modal feature information; multi-modal features are fused based on a modal perception attention mechanism and a diagnosis feedback readjustment mechanism, and intelligent diagnosis of a current flight state is realized in combination with a semantic consistency discrimination mechanism. When the flight state is judged to be abnormal or accident, starting a satellite communication channel to carry out remote emergency transmission on key data; and when the flight state is judged to be normal, performing modular hierarchical storage on the acquired data according to the data type. And the separated emergency data transmission subsystem is immediately started under the extreme conditions of power failure of the main system and the like. According to the method, strategies such as multi-mode intelligent diagnosis and multi-level storage are fused, and the flight data acquisition efficiency and the flight safety guarantee capability are remarkably improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Video image analysis system and analysis method thereof

The invention belongs to the technical field of computer vision and intelligent analysis, and particularly relates to a video image analysis system and an analysis method thereof.The method includes the steps that firstly, through the multi-modal processing thought of scene partition compensation, dynamic noise distinguishing and detail enhancement and self-adaptive interference coping, high-quality images are output for follow-up analysis; the target identification basic precision is guaranteed from the source; then fusing the appearance, motion and semantic multi-dimensional features of the target to realize accurate recognition, establishing a unique target ID in combination with a prediction matching algorithm, and matching with a shielding trajectory prediction and recovery mechanism to ensure that the target is continuously and stably tracked in a shielding and rapid moving scene and avoid interruption of a tracking link; on the basis, the method can adaptively cope with complex environments such as backlighting and rainy days, the image quality is improved through targeted preprocessing, a reliable foundation is laid for target recognition, and the problem of low recognition accuracy in complex scenes is effectively solved.
Owner:BEIJING DONGYU HONGDA TECH CO LTD

Document analysis and retrieval enhancement generation method fused with multi-modal embedding model

The invention relates to a document analysis and retrieval enhancement generation method fused with a multi-modal embedding model, and belongs to the technical field of AI large model application. The method comprises the following steps of: analyzing a file and extracting contents of texts, tables and pictures; performing joint semantic embedding on the extracted texts, tables and pictures by using a multi-modal embedding model to generate uniform vector representation, and constructing a vector database supporting multi-modal data; and constructing a multi-modal retrieval chain based on a vector retriever. According to the method, total element analysis and joint modeling of complex documents are achieved, the problem that multi-modal information is difficult to process in a unified mode through a traditional vectorization technology is solved to a certain degree, and the information coverage capacity and the answer accuracy of a retrieval enhancement generation system in a real scene are remarkably improved.
Owner:HONGHE POWER SUPPLY BUREAU OF YUNNAN POWER GRID

Systems, methods, and user interface for navigating media playback using scrollable text

A mobile computing device can be configured, with an improved user interface, to synchronously play audio (or video) and text associated thereto, such as text stored in a synchronization index. Using the synchronization index, the device can periodically compare the current track time with that time associated to a word or range of words, such as a line (or segment) in a plurality of lines of text (or segments of text). Improved navigability of content using an improved mobile computing device and user interface is provided, because a user of the device can scroll through the lines of text associated with the audio or video to find a target word or range of words. If the user selects a particular word or range of words, by making a gesture on the mobile computing device, the device can identify a start time for the selected text. The device can then play the audio or video at the identified start time of the selected text. The improved user interface is a practical application for navigating audio (or video) and text associated thereto on a mobile computing device, providing bimodal reading on mobile computing devices and ease of navigability.
Owner:EVANS CURTIS

Data analysis method

The invention provides a data analysis method. The data analysis method comprises the steps of obtaining query information for data; according to the method, the query information is analyzed, the query requirement of the user is determined, the pre-constructed data warehouse is searched based on the query requirement, the target query content is determined, and the pre-constructed data warehouse is associated with databases storing different types, so that cross-database query can be performed based on the query requirement of the user, and the query efficiency is improved. Therefore, various types of data query of texts, images, videos and the like can be realized, and multi-dimensional and complex-dimensional data analysis can be realized. According to the query information, the generation template of the data is determined, and the analysis report is generated according to the generation template based on the target query content, so that automatic report generation is realized, the generated report conforms to the identity of the user, and the working efficiency of the user is improved.
Owner:中国卫通集团股份有限公司

Multi-modal corpus duplicate removal method and device based on AI large model, equipment and medium

The invention discloses a multi-modal corpus deduplication method and device based on an AI large model, equipment and a medium, and relates to the technical field of data processing, the method comprises the following steps: performing type identification on each pre-training corpus contained in an obtained pre-training corpus set, and determining a modal type; preprocessing each pre-training corpus based on the modal type to obtain a current pre-training corpus set; performing feature extraction on each current pre-training corpus based on the modal type to obtain a semantic vector and a structure vector corresponding to each current pre-training corpus; based on the semantic vector and the structure vector, determining a structure sensing semantic fingerprint corresponding to each current pre-training corpus; determining a redundant cluster based on the structure perception semantic fingerprint; determining redundant corpora from the redundant clusters; and deleting redundant corpora in the current pre-training corpus set to obtain a de-duplicated pre-training corpus set, and the multi-modal corpus de-duplication accuracy and efficiency based on the AI large model can be improved.
Owner:广东知业科技有限公司

Intelligent exhibition hall multi-mode interactive digital human system and implementation method

The invention relates to the technical field of computer data processing, and discloses an intelligent exhibition hall multi-modal interaction digital human system and an implementation method, which are used for solving the problem that multi-source data of multi-modal interaction lacks a unified clock and verifiable timestamp alignment mechanism in a traditional method. According to the method, a unified time domain and a time version are established on an edge side, terminal access is restrained, and an acquisition time mark, an access time mark and a serial number are written in data or a state; performing gating shunting according to a time version, performing de-duplication and out-of-order rearrangement based on a serial number and double time marks, and generating and solidifying a session window evidence index; locking a transaction window boundary according to the evidence index, generating a participation source list and a transaction number, establishing a fragment reference relationship and generating an alignment voucher; in the linkage stage, phase division issuing is carried out according to a preparation phase, an execution phase and a confirmation phase, backward reading verification is carried out according to an action sequence number, a transaction log is archived, and alignment, rechecking and playback of an interaction link are achieved.
Owner:SUZHOU CHUANGJIE MEDIA EXHIBITION CO LTD

Multimedia file operation method and device, electronic equipment and storage medium

The invention provides a multimedia file operation method and device, electronic equipment and a storage medium, and relates to the technical field of multimedia. The file system is used for storing multimedia files; and then mounting the target snapshot to a target directory of the service end, and associating the target snapshot with a required file format of the service end. According to the method, the snapshot of the file system is introduced, and a channel is provided for the access of the multimedia file in a snapshot mounting manner, so that a service end can quickly access the target file in the required file format, and the file access efficiency is improved. According to the method, the format of the required file can be automatically converted without an additional video-on-demand server, and a streaming playing mode is not needed. When the access is triggered, format conversion is performed on the to-be-accessed file in the target snapshot in real time, so that the to-be-accessed file in the target snapshot is applied by the service end in the required file format in time, the application efficiency of the service end is improved, and the hardware cost is reduced.
Owner:ZHEJIANG UNIVIEW TECH CO LTD

Film series intelligent dynamic collection method and device based on multi-modal large model

The invention discloses a film series intelligent dynamic collection method and device based on a multi-modal large model, and belongs to the technical field of film and television content processing, and the method comprises the steps: extracting text information, audio information and visual information of film and television content; the text information is analyzed to extract key semantic features, the audio information is analyzed to extract emotion tone features, the visual information is analyzed to extract image hue, shot language and scene composition visual style features, and multi-modal feature vectors are obtained; inputting the multi-modal feature vectors into a pre-trained artificial intelligence large model, and inferring an association relationship between the film and television contents and a series episode category to which the film and television contents belong in combination with a preset series episode structure rule base; and carrying out dynamic mapping comparison on the inferred incidence relation between the film and television contents and the category of the series with the stock series, and storing a comparison result. The collection efficiency and accuracy are improved, the labor cost is reduced, and the experience of watching the series of the user is improved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

Method and system for generating three-dimensional object based on semantic analysis

The invention relates to a method and system for generating a three-dimensional object based on semantic analysis. More specifically, the present invention relates to a method and system for automatically generating a three-dimensional object corresponding to the meaning of a sentence by analyzing the sentence in a text form.
Owner:NATIONA INC

Index recovery method and device for streaming media data and storage medium

The invention discloses an index recovery method and device for streaming media data and a storage medium, and the method comprises the steps: persistently storing a data frame to a data storage area, caching an index to an index cache area, and sending the index to a slave storage device for backup storage; if the preset index flashing condition is met, persistently storing the cached index into an index storage area; after the main storage device is abnormal, the main storage device returns to normal, a backup index in the slave storage device is obtained, and the backup index is persistently stored in an index storage area; and if the index corresponding to the streaming media in the index storage area is still not completely recovered, reading the index of the data frame in the data storage area, and persistently storing the read index in the index storage area. The index recovery is carried out according to the backup index in the slave storage device, the index recovery speed can be increased, the information of the recovered index is more complete, then the index of the data frame in the data storage area is read for index recovery, and the index integrity of each data frame is ensured.
Owner:ZHEJIANG DAHUA TECH CO LTD

A Zero-Sample Cross-Modal Retrieval Method Based on Adaptive Class-Related Discrete Hashing

This invention discloses a zero-shot cross-modal retrieval method based on adaptive class-related discrete hashing. A novel cross-modal zero-shot hashing method is proposed to effectively transfer class attribute knowledge. This method constructs a semantically enhanced embedding by fusing label information with class attribute information, which can solve the problem of class attribute correspondence for multi-label instances. By learning the semantically enhanced embedding, more semantic information is embedded into the feature representation, thereby balancing the retrieval results between images and text. This method fully considers the correlation between class attributes and adaptively embeds more class attribute semantic information into the hash code. Simultaneously, the hash code can effectively capture the relationship between visible and invisible classes, thus transferring attribute knowledge from visible classes to invisible classes. Finally, pairwise similarity is embedded in the hash code learning process to enhance the semantic information in the hash code. This invention improves retrieval accuracy in zero-shot cross-modal retrieval scenarios.
Owner:KUNMING UNIV OF SCI & TECH

Multi-modal data processing and retrieval method

The invention particularly relates to a multi-modal data processing and retrieving method. The multi-modal data processing and retrieval method comprises the following steps: respectively carrying out depth feature extraction on image data and text data to generate an image vector and a text vector; carrying out interaction on the image vector and the text vector, and capturing semantic association information among multiple modes; mapping the vectors after interaction to a unified semantic space to realize semantic alignment among different modes; dividing the unified semantic space vector into fragments, storing the fragments in distributed nodes, and establishing a vector index; and generating a query vector by using an image or text queried by a user, carrying out parallel calculation on the similarity with a storage vector, and returning a retrieval result according to the similarity. According to the multi-modal data processing and retrieval method, the problems that the semantic difference between different modal data such as images and texts is large, the retrieval efficiency is low and storage is difficult to expand are solved, the accuracy and efficiency of multi-modal data retrieval are remarkably improved, and the method has remarkable technical advantages and wide application scenes.
Owner:INSPUR QILU SOFTWARE IND