Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

272 results about "Semantic gap" patented technology

The semantic gap characterizes the difference between two descriptions of an object by different linguistic representations, for instance languages or symbols. According to Hein, the semantic gap can be defined as "the difference in meaning between constructs formed within different representation systems". In computer science, the concept is relevant whenever ordinary human activities, observations, and tasks are transferred into a computational representation.

Accurate micro-crack segmentation method integrating feature fusion and convolution attention

The invention provides a microcrack precise segmentation method integrating feature fusion and convolution attention, and belongs to the field of image processing. According to the method, a crack segmentation network based on an encoder-decoder architecture is constructed, a convolution block attention module is introduced at an encoder end, background noise is adaptively suppressed and obvious characteristics of cracks are enhanced through a channel and space dual attention mechanism, and the method is suitable for the adaptive segmentation of the cracks on the premise of almost not increasing the calculation overhead. The sensitivity of the model to microcracks is improved; a feature fusion module is introduced at a decoder end, and cooperation of low-layer details and high-layer semantics is realized through cross-layer fusion, so that a semantic gap is effectively bridged, detail loss caused by traditional convolution stacking is avoided, and continuity and a complete topological structure of a long and narrow crack are ensured. According to the method, through collaborative optimization of multi-scale feature extraction and an attention mechanism, accurate capture of the saliency features of the crack and effective suppression of complex background interference are realized, and the detection sensitivity and overall segmentation consistency of the micro-crack are remarkably improved.
Owner:DALIAN UNIV OF TECH

Multi-modal content understanding method and system based on knowledge graph

The invention discloses a multi-modal content understanding method and system based on a knowledge graph, and belongs to the technical field of multi-modal content understanding, the method comprises the steps of obtaining multi-modal input data and conducting feature decoupling, semantic information in the multi-modal data and noise and redundant information peculiar to modals can be effectively separated through the feature decoupling technology, and the multi-modal content understanding efficiency is improved. The method comprises the following steps of: establishing a space-time perception graph attention network, improving the purity and semantic expression capability of features, bridging semantic gaps among different modals through the space-time perception graph attention network, realizing cross-modal semantic alignment, enhancing the generalization capability of a model, deeply mining space-time modes, relationships and anomalies in data through space-time association reasoning, and improving the accuracy of the data. The method provides support for multi-modal data analysis in a complex scene, and combines a graph neural network and a recurrent neural network to capture semantic association and spatial-temporal dynamics and optimize the performance of an inference model.
Owner:HUNAN UNIV OF SCI & TECH

Enterprise multi-modal data intelligent processing system fusing RAG technology and intelligent processing method of enterprise multi-modal data intelligent processing system

The invention discloses an enterprise multi-modal data intelligent processing system fused with an RAG technology and an intelligent processing method of the enterprise multi-modal data intelligent processing system, and relates to the technical field of enterprise-level multi-modal data intelligent processing. And the data processing module is configured to respectively process the structured data and the unstructured data through the dynamic heterogeneous encoder and output unified semantic representation by adopting a cross-modal adversarial alignment mechanism. According to the enterprise multi-modal data intelligent processing system fused with the RAG technology, the problem of enterprise multi-modal data splitting is solved through dynamic adversarial semantic alignment and a stepped fusion mechanism. Semantic gaps are eliminated through self-adaptive convergence of cross-modal features in a hidden space, deep association of heterogeneous data is achieved based on concept mapping and credibility arbitration of an ontology network, key information of unstructured data is accurately extracted and converted into structured knowledge, and the accuracy of cross-modal association analysis and decision reliability are improved.
Owner:SHANGHAI WICRESOFT

Intelligent personalized learning path recommendation system

The invention discloses an intelligent personalized learning path recommendation system, and relates to the technical field of learning systems, and the system comprises a multi-source data collection module which is used for synchronizing behavior data of students in a cross-platform learning scene in real time; the cognitive feature analysis engine comprises a style recognition sub-engine and a demand prediction sub-engine to predict the potential learning demand intensity of students for unmastered knowledge points and generate a priority list comprising knowledge gaps; the personalized path generator generates a three-dimensional path plan; the proportion of guided questioning, example demonstration and autonomous exploration is automatically configured according to knowledge difficulty; the dynamic self-adaptive adjustment unit is used for designing a reward function including short-term progress speed and long-term ability growth potential by taking real-time performance of students as a state space and taking path adjustment action as a decision space based on a reinforcement learning framework; and a double-loop feedback mechanism is realized. According to the method, the dimension limitation of traditional learning analysis is broken through, the difficulty of stiffness of a static course template is broken through, and meanwhile, the semantic gap of subject cognition is broken through.
Owner:SUZHOU HAOYI LIGHTING TECHNOLOGY CO LTD

Multi-modal data fusion method and system based on large model

The invention discloses a multi-modal data fusion method and system based on a large model, and relates to the technical field of data fusion, and the method comprises the steps: receiving multi-modal data, and carrying out the noise layering filtering and time-space alignment; mapping data to a large model hidden space through each modal lightweight encoder, extracting initial features by using a single-modal pre-training model, and converting the initial features through an adapter to generate a hidden space vector; each modal hidden space vector and a task cue word are received, a fusion feature matrix is dynamically aggregated and generated through a self-attention mechanism and a cross attention mechanism of a large model, and semantic integration of multi-modal information is realized; taking the fusion feature matrix as a soft label, and learning a cross-modal semantic mapping capability through knowledge distillation; the trained model is deployed to the edge through dynamic quantization and pruning optimization, an optimized feature matrix is input into a task-customized lightweight head network, and a final task result is output in combination with multi-modal context information. The method can break through the semantic gap between modals, and improves the fusion efficiency.
Owner:浪潮智慧城市科技有限公司 +1

Semantic-based migration and consistency verification method, system and equipment and medium

The invention provides a semantic-based migration and consistency verification method, system and device and a medium, and relates to the technical field of databases. The method comprises the following steps: performing metadata topology scanning analysis by obtaining metadata, including generating an abstract syntax tree and constructing a semantic graph; according to a metadata topology scanning analysis result, performing data mapping and conversion, including loading a YAML rule and generating corresponding type mapping and constraint conversion; data migration is carried out through primary key fragmentation parallel migration, batch writing optimization and real-time double-writing verification; through structure-data-business three-layer verification, simulation business SQL comparison and automatic difference repair, the integrity of migrated data and business compliance are ensured, the semantic gap problem in heterogeneous database migration is solved, high-precision and high-efficiency database migration is realized, the migration efficiency is improved, and the data migration efficiency is improved. The method is suitable for scenes with strict requirements on data consistency and migration efficiency, such as financial, government affair and enterprise-level data centers.
Owner:CHINA YANGTZE POWER

Association modeling method and system based on knowledge-data dual drive

The invention relates to a knowledge-data dual-drive-based association modeling method and system, and the method comprises the steps: firstly obtaining video monitoring data, sensor data and domain safety specification knowledge, and carrying out the frame extraction and feature extraction, time sequence standardization processing and structured representation and coding; respectively extracting visual features, time sequence features and semantic features to obtain safety knowledge features; modeling the multi-modal data and the high-order association of the multi-modal data and the safety knowledge by using a hypergraph neural network; multi-modal features and security knowledge features are input into the hypergraph neural network, deep fusion of multi-modal data and security knowledge is realized through hypergraph convolution operation, and a unified representation space is generated; and constructing a potential safety hazard identification model, and performing accurate identification and real-time monitoring on the potential safety hazard in the complex scene through the potential safety hazard identification model. The problems of heterogeneity of multi-modal data and semantic gaps between modals are solved, and the accuracy and real-time performance of potential safety hazard recognition are improved.
Owner:SHANDONG HI SPEED CONSTRUCTION MANAGEMENT GROUP CO LTD +1

TransUNet-based medical image segmentation method

The invention discloses a medical image segmentation method based on TransUNet, and belongs to the technical field of medical image segmentation. The method comprises the steps of firstly performing data preprocessing on an original image to obtain preprocessed data; and a DCA attention module is used at a jump joint, so that the problem that a semantic gap exists between characteristics of an encoder and a decoder due to the fact that a simple jump connection scheme is difficult to capture a multi-scale context is solved. The semantic difference leads to redundancy between low-level and high-level features, and finally the segmentation performance is limited. Secondly, a multi-scale boundary sensing module is added to the top layer of the encoder, so that the neural network can better segment the boundary of the target image in the training process; and inputting the preprocessed data into the improved TransUNet model to train the medical image, and outputting an image segmentation result.
Owner:BEIJING UNIV OF TECH

Remote sensing image multi-source heterogeneous data fusion processing method and system

The invention discloses a remote sensing image multi-source heterogeneous data fusion processing method and system, and the method comprises the steps: extracting global information from remote sensing image data through employing a convolutional neural network, extracting image local features through cutting operation, and capturing local feature information in the remote sensing image data; constructing a cross-time-domain attention mechanism for the local feature information through a cyclic matrix to extract mutual information among different modal variables, and screening highly-associated cross-time-domain key information; establishing a cross-time-domain sensing hierarchical aggregation module for the cross-time-domain key information, and obtaining detail information and edge information of the remote sensing image; acquiring global fusion data by adopting an asymptotic fusion strategy, and introducing a loss function to reduce a semantic gap; according to the method, missing information is repaired by adopting an interactive network model of mixed contrast learning, complete real-time remote sensing image multi-source heterogeneous fusion data is obtained, and efficient and high-precision fusion processing of the multi-source remote sensing data is realized by constructing a multi-collaborative deep fusion framework.
Owner:CHINA AERO GEOPHYSICAL SURVEY & REMOTE SENSING CENT FOR LAND & RESOURCES

Robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention

The invention relates to a robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention, and belongs to the technical field of image processing. Aiming at the problems of small target feature loss, semantic gap, background noise interference and the like caused by a fixed convolution kernel scale, one-way feature fusion and a static attention mechanism in an existing unmanned aerial vehicle aerial image target detection method, the method comprises the following steps: constructing a detection model comprising a backbone network, a neck network and a detection head network; a feature rearrangement and extraction module is designed in the backbone network to enhance feature learning, an enhanced double-flow feature fusion pyramid is designed in the neck network to optimize multi-scale feature fusion, and a dynamic multi-scale context attention mechanism is designed in the detection head network to suppress irrelevant background noise. The method effectively improves the accuracy and robustness of small target detection, and achieves a clearer and more stable detection effect in a complex environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Emotion analysis method, system and equipment based on multi-modal information and large language model and medium

The invention discloses an emotion analysis method, system and device based on multi-modal information and a large language model and a medium, belongs to the technical field of natural language processing and multi-modal calculation, and aims at solving the technical problem of how to overcome the defects of data distribution difference, modal credibility deviation and low generated data quality in the prior art. According to the technical scheme, the method comprises the following steps: collecting multi-modal data; multi-modal data preprocessing and feature extraction: preprocessing the text data, the image data and the audio data respectively and extracting corresponding features; feature alignment: mapping features of different modes of texts, images and audios to a unified potential space by adopting a cross-modal alignment algorithm, and eliminating semantic gaps among the modes; weight distribution: adopting a dynamic weight distribution mechanism to adaptively and dynamically distribute weights of different modes; performing cross-domain attribute-level sentiment analysis; evaluating the credibility; and calibrating the confidence coefficient.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

Ship user behavior self-learning recommendation system based on large model

The invention relates to a ship user behavior self-learning recommendation system based on a large model, and relates to the technical field of ship informatization. According to the system, through collection and fusion of multi-source heterogeneous ship user behavior data, a large-scale pre-training language model (large model) is utilized to carry out deep understanding and semantic mining on massive ship field text information and user behavior sequences, and a ship field knowledge graph or semantic vector space is constructed. The large model can identify and predict potential demands, behavior patterns and preference changes of ship users, and generates highly personalized, accurate and prospective ship service, product, route or information recommendations in combination with real-time operation data and external environment factors. Besides, a user feedback self-learning mechanism is introduced into the system, recommendation strategies and model parameters are continuously optimized according to interaction behaviors and explicit evaluation of the users in modes of reinforcement learning or continuous learning and the like, and intelligent iteration of the system and continuous improvement of the recommendation effect are achieved. According to the method, the challenges of a traditional recommendation system in the aspects of data complexity, semantic gaps and dynamic demand adaptability in the ship field are effectively solved, and the ship operation efficiency and the user satisfaction degree are remarkably improved.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Remote nursing monitoring method and system based on multi-modal image fusion

The invention provides a remote nursing monitoring method and system based on multi-modal image fusion, and the method comprises the steps: obtaining multi-modal data, carrying out the multi-modal feature fusion processing of the multi-modal data, and generating cross-modal semantic information, the multi-modal data comprising at least two of a visible light image, an infrared thermal imaging and a point cloud; the cross-modal semantic information is subjected to dynamic optimization retrieval processing based on a query intention, a retrieval result set matched with the query intention is generated, and the query intention is input through a natural language or obtained through a preset template instruction; and performing abnormal feature detection processing on the retrieval result set to generate a real-time monitoring result. By adopting the method, the cross-modal semantic gap problem can be effectively solved, the precision and real-time performance of nursing monitoring are improved, and reliable technical support is provided for remote nursing monitoring.
Owner:GENERAL HOSPITAL OF THE NORTHERN WAR ZONE OF THE CHINESE PEOPLES LIBERATION ARMY

Motion track construction method based on visual large model

The invention discloses a motion track construction method based on a visual large model, and relates to the technical field of computer vision, and the method comprises the steps: employing a target tracking algorithm based on Kalman filtering, and generating a smooth time sequence synchronization track from an original laser radar point cloud and an original video stream; dividing the trajectory into a series of candidate kinematic segments which are complete in kinematics and semantics by applying a hybrid method combining a minimum description length principle and visual language model semantic verification; performing preliminary semantic annotation on the segments by utilizing a visual language model to generate an initial annotation track; performing logic consistency refining on the initial labeling track until the initial labeling track is converged into a refined labeling track; and formatting the refined trajectory into a standard structured semantic trajectory representation character string. According to the method, a semantic gap between low-dimensional physical observation and high-dimensional driving intention is bridged, and ideal input is provided for understanding and prediction tasks of downstream complex scenes.
Owner:ANHUI GUOZHI DATA TECH CO LTD

Multi-type context-aware dialogue recommendation method based on hybrid expert model

The invention discloses a multi-type context-aware dialogue recommendation method based on a hybrid expert model. The aim of the invention is fulfilled by a hybrid expert framework comprising a plurality of expert module coordination systems. The knowledge graph expert model is responsible for extracting and coding structured entity relationship information, the dialogue expert model is used for deeply mining key contents and user intentions in dialogues, and the comment experts are used for analyzing user emotion and preference information in commodity comments. In this way, the system can make full use of semantic information of the structured data and the unstructured data, and the problem of semantic gaps caused by data heterogeneity in the prior art is solved. On the basis, a coordination system is introduced to serve as a core module for coordination and integration. The coordination system can dynamically allocate the output of the expert module according to the weights and contributions of different context information to generate a final recommendation result, so that efficient fusion of multiple types of context information is realized, the problems of semantic alignment and inconsistency between different types of data are solved, and the performance of the dialogue recommendation method is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Attribute decoupling multi-subspace proxy learning text-image pedestrian re-identification method

The invention discloses an attribute decoupling multi-subspace proxy learning text-image pedestrian re-recognition method. The method comprises the following steps: analyzing original text description into a plurality of attribute-level sentences through a fine-grained text reconstruction module; decomposing the global visual features into a plurality of semantic subspaces by adopting an image subspace projection mechanism; a unified multi-granularity subspace loss function is used for training optimization, and the loss integrates global comparison loss, attribute-level comparison loss and subspace diversity regularization loss. According to the method, attribute-level semantics are explicitly decoupled, and fine-grained alignment is realized in a plurality of agent subspaces, so that a semantic gap between a text and an image is effectively bridged, the precision and robustness of cross-modal matching are improved, competitive performance is obtained on standard indexes such as Rank-1 and the like, and the method can be widely applied to the field of text and image processing. And meanwhile, higher convergence speed and better interpretability are shown.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

Electric power operation target detection method based on multi-mode large model knowledge distillation

The invention relates to the field of target detection, and particularly discloses an electric power work target detection method based on multi-modal large model knowledge distillation, which utilizes a vision-language multi-modal large model as a teacher model, and improves the target detection efficiency by expanding prompt word guidance. A high-quality pseudo label and a region-text pair are generated for an unlabeled electric power work image as a supervision signal, and on this basis, through joint optimization of detection loss, feature distillation loss, logic distillation loss and multi-modal contrast learning loss, a lightweight YOLO student model is guided to learn positioning and classification knowledge and to learn a multi-modal contrast learning loss. And deep alignment with the open vocabulary understanding ability of the teacher model is carried out on the feature space and semantic level, so that a semantic gap between closed category detection and open world perception is effectively bridged. Through the mode, the detection precision and generalization ability of the student model on common, rare and even unseen targets in the electric power work scene are remarkably improved.
Owner:MARKETING SERVICE CENT OF STATE GRID HENAN ELECTRIC POWER CO

Intelligent analysis method and system based on multi-source data

The invention relates to the technical field of encrypted communication, and discloses an intelligent analysis method and system based on multi-source data, and the method comprises the steps: obtaining multi-dimensional metadata of an encrypted communication flow, and generating a feature vector set; calculating a state transition probability of a communication session by using a Markov chain model, dividing the communication session according to a time window and mapping the communication session to a discrete state, and generating a session state evolution sequence; analyzing a session state evolution sequence based on a hierarchical topic model, mapping the state sequence to a behavior topic through potential Dirichlet allocation, mapping the behavior topic to an intention category through a hierarchical Dirichlet process, and outputting a communication intention category and a confidence score; according to the method, the limitation that static feature analysis cannot reflect the complete life cycle of the session is overcome, the problem of semantic gaps caused by limited metadata information dimensions is solved, and the accuracy and robustness of intention recognition are improved.
Owner:TANGREN COMM TECH CO LTD

Fault diagnosis method and system based on multi-modal feature fusion and lightweight fine tuning

The invention relates to the technical field of industrial intelligent diagnosis and large language model crossing, and provides a fault diagnosis method and system based on multi-modal feature fusion and lightweight fine tuning. According to the method, multi-source heterogeneous data such as a time sequence, an image and a text of industrial equipment are collected, after feature extraction and cross-modal fusion are conducted, fusion features are adapted to a large language model input space through a projection embedding layer, and then intelligent diagnosis is achieved in combination with a large language model subjected to low-rank self-adaptive fine adjustment. According to the method, the problems of semantic gaps and dimension mismatch of multi-modal industrial data are effectively solved, the diagnosis precision and the model generalization ability are remarkably improved, meanwhile, the requirement for computing resources is greatly reduced, the interpretability of diagnosis results is enhanced through structured output, and an innovative technical path is provided for intelligent operation and maintenance of industrial equipment.
Owner:武汉中云康崇科技有限公司

Image geographic positioning system and method based on geographic feature extraction

The invention relates to the field of image content understanding, and discloses an image geographic positioning system and method based on geographic feature extraction, and the system comprises a feature extraction module which is used for extracting a visual feature vector, a GPS feature vector, a text position description feature vector and a text scene description feature vector from multi-modal input; the comparative learning module is used for realizing multi-modal feature alignment through comparative learning of the visual features, the GPS features, the text position description features and the text scene description features; and the data set construction module is used for fusing the multi-modal features to generate geographic feature vectors and constructing a retrieval vector data set. By adopting the technical scheme of multi-modal feature fusion and cross-modal comparative learning, the technical effect of improving the geographic positioning precision and generalization ability is achieved. Compared with a scheme depending on a single mode or simple feature splicing in the prior art, the problem of semantic gaps caused by mode information splitting in a traditional method is solved.
Owner:JUSAFE (BEIJING) TECHNOLOGY CO LTD

Multi-modal feature coding and cross-modal adaptive fusion method

The invention provides a multi-modal feature coding and cross-modal adaptive fusion method, which comprises the following steps of: respectively extracting radar features from radar sequence data, extracting visual features from a radar target image, respectively extracting text data and knowledge graph data from text data, and fusing the features of different modals to obtain a multi-modal feature coding and cross-modal adaptive fusion model; specifically, distribution differences among modals are eliminated through statistical normalization, then the modals are projected to a shared feature space to realize dimension unification, and finally efficient cooperation of radar, vision and text modals is realized in combination with semantic alignment guided by a knowledge graph and dynamic weighting of input self-adaption. According to the method, the problems of feature isomerism and semantic gaps in heterogeneous data fusion are effectively solved.
Owner:CHINA SHIPBUILDING LINGJIU HIGH TECH (WUHAN) CO LTD +1

Knowledge graph-based long video key frame retrieval method and device

The invention relates to the technical field of multi-mode intelligent video understanding, and provides a long video key frame retrieval method and device based on a knowledge graph. Through the processes of frame-level subtitle generation, frame-level knowledge graph construction, similarity video segmentation and fragment and abstract generation, a long video knowledge graph construction assembly line is constructed, and structured modeling of long video semantic content is realized. By setting a two-stage retrieval mechanism, higher retrieval precision can be obtained while the efficiency is ensured. Vector matching and multi-hop neighbor extension are carried out on a unified knowledge graph, and a node set strongly related to the problem is positioned, so that the semantic gap between the natural language problem and a structured graph is reduced, and the accuracy of key frame selection is improved. By setting an iterative retrieval mechanism, the retrieved key frame can be used as a basis for answering a question text to the maximum extent.
Owner:NAT UNIV OF DEFENSE TECH

Engineering quantity list intelligent matching method

The invention discloses an intelligent engineering quantity list matching method, which comprises the following steps of: firstly, performing deep analysis on project description to form a key feature dictionary and form a basis for subsequent accurate matching; then, adopting a two-stage process of coarse screening and fine arrangement: in the coarse screening stage, combining project codes pointing to professional classification with key features capable of quickly locking core elements, performing efficient filtering on a huge standard quota library, and quickly generating a candidate matching list which is controllable in scale and highly related; and in a fine ranking stage, turning to more detailed semantic and feature dual verification, comprehensively utilizing the deep semantic similarity of a complete item description text and an accurate comparison result of a key feature dictionary, and performing fusion calculation to obtain a final score so as to realize accurate ranking of candidate lists. In this way, the semantic gap caused by synonyms, inverted word order, non-standard expression and the like in traditional text matching can be effectively overcome, and the matching accuracy is greatly improved.
Owner:浙江省建筑科学设计研究院建筑设计所

Network structure combining residual jump enhancement and multi-dimensional interactive attention

The invention belongs to the field of pulmonary embolism detection and segmentation, and discloses a residual error jump enhancement and multi-dimensional interactive attention combined network structure, which comprises an encoder, a decoder, a jump connection fusion module and a bottleneck module which are respectively responsible for feature extraction, spatial restoration, multilayer semantic fusion and modeling of a long-distance dependency relationship. The AGRS module is fused with a residual jump enhancement structure and a multi-dimensional interaction quadruple attention module; the CDB-ASPP module realizes multi-scale context information efficient fusion through five parallel convolution branches with different expansion rates, combines a global-local enhancement strategy, and considers both a global receptive field and a local detail perception capability; and a dual attention Transform module is used for enhancing the global feature modeling capability. According to the method, an effective solution is provided for solving key problems such as semantic gaps, context information missing and small target detection sensitivity, and important reference and enlightenment are provided for deep learning-based pulmonary embolism auxiliary diagnosis research.
Owner:INNER MONGOLIA UNIV OF SCI & TECH +1

Remote sensing image-text retrieval method based on knowledge enhancement and asymmetric structure

The invention relates to a remote sensing image-text retrieval method based on knowledge enhancement and an asymmetric structure, and belongs to the technical field of remote sensing image-text cross-modal retrieval. Comprising the following steps: inputting a to-be-retrieved remote sensing image and text into a trained vision-language basic model with knowledge enhancement and an asymmetric structure, and realizing cross-modal retrieval through model processing; and obtaining a retrieval result, obtaining matching retrieval output of the remote sensing image and the text, and completing a cross-modal retrieval task. According to the method, the problem that the retrieval performance is limited due to the fact that significant information asymmetry exists between the remote sensing image and the text modality at present is solved. According to the method, the cross-modal asymmetric Kolmogorov-Arnod adapter fine tuning method is designed, so that efficient modal fine-grained shared feature learning is realized; meanwhile, knowledge is extracted from ConceptNet and a remote sensing knowledge graph, and a knowledge enhancement sentence is generated to enrich text semantics, so that a semantic gap between modals is bridged, and remote sensing image-text retrieval performance is improved.
Owner:YUNNAN NORMAL UNIV

Construction process interactive visual disclosure system driven by multi-modal fusion

The invention relates to a multi-modal fusion driven construction process interactive visualization disclosure system, which comprises a multi-modal data analysis and fusion module, a process flow modeling and verification module, an animation generation and physical simulation module, an interactive visualization and cooperation module and a quality verification and feedback module, visual viewing and interactive operation of the construction disclosure process are achieved through multi-item collaboration, and the problems that in a three-dimensional disclosure mode in the prior art, semantic gaps exist in construction documents and visual content, the production efficiency of disclosure animation materials is low, and the disclosure process lacks interactivity and standard verification capacity are solved. According to the multi-modal fusion driven interactive visual disclosure system disclosed by the invention, natural language processing, knowledge graph, physical simulation and distributed rendering technologies are integrated, automatic analysis, dynamic visualization and interactive disclosure of a construction process are realized, the construction disclosure efficiency is remarkably improved, the human error rate is reduced, and the construction efficiency is improved. The method is suitable for complex engineering scenes such as large housing construction and long-piled wharfs.
Owner:CCCC TIANJIN ECO ENVIRONMENTAL PROTECTION DESIGN & RES INST CO LTD

Luggage material identification method and system based on multi-modal fusion knowledge distillation

The invention relates to a luggage material identification method and system based on multi-modal fusion knowledge distillation. The method comprises the steps of obtaining image data and point cloud data of a to-be-detected target surface; constructing a multi-modal teacher model to obtain image texture features and point cloud geometric features; obtaining a teacher object query; outputting a material category score and a bounding box position coordinate; constructing a lightweight student model, and generating student object query, material category prediction and bounding box position prediction; characteristic distillation loss is designed for characteristic distillation, and a total loss joint training lightweight student model is constructed; actually operating the trained lightweight student model in a luggage detection scene of an airport luggage turntable; calculating a stacking score and mapping the stacking score into a stacking label; according to the method, the semantic gap between the perception recognition module and the downstream planning strategy module is effectively linked, and the contradiction between the insufficient precision of traditional single-mode perception and the high cost of multi-mode deployment is solved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Multi-modal data semantic alignment method and system based on knowledge graph embedding

The invention discloses a multi-modal data semantic alignment method and system based on knowledge graph embedding, relates to the technical field of inspection and detection, and solves the problems of low data accuracy and interpretability after multi-modal data fusion. According to the embodiment of the invention, through a knowledge graph embedding and semantic alignment mechanism, the semantic gap problem of cross-modal data is solved, and accurate association of multi-modal data in an inspection and detection scene is realized; the optimized fusion vector obtained by fusion enhances the interpretability and consistency of the data, and improves the detection accuracy and robustness of a subsequent detection model.
Owner:GUILIN UNIV OF ELECTRONIC TECH

High-standard farmland scene recognition method based on deep neural network

The invention discloses a high-standard farmland scene recognition method based on a deep neural network, and relates to the technical field of remote sensing image processing and agricultural information, and the method comprises the steps: 1, constructing a high-standard farmland scene sample library with prior significance, the high-standard farmland scene sample library comprises high-standard farmland sample features and high-standard farmland sample scale quantities; 2, constructing a high-standard farmland scene recognition model based on multi-source data fusion; 3, predicting a result, and outputting a scene category graph; according to the high-standard farmland scene recognition method based on the deep neural network provided by the invention, the problem that the prior art stays at a pixel level or an object level, depends on a single threshold value or shallow learning and has a semantic gap is solved.
Owner:CHINA AGRI UNIV

Method and system for text retrieval of picture archives based on cross-modal feature alignment

The invention discloses a method and a system for retrieving a picture file through a text based on cross-modal feature alignment, and the method comprises the steps: extracting text and picture features, and guaranteeing that a high-value mode contributes to a higher weight based on multi-modal attention weighted fusion; calculating a dynamic temperature coefficient through initial semantic similarity of positive and negative samples to construct a loss function for multi-modal contrast learning, and mapping original features of a text and an image to a unified space to obtain alignment features; according to the method, the picture archives are retrieved through texts, the image-text similarity, the time sequence weight and the core area proportion weight are comprehensively considered, optimization sorting of time sequence perception is carried out, and the most matched picture archives are obtained. According to the method, a text-image cross-modal semantic gap is solved, so that semantic alignment of two types of features in a unified space is realized; according to the method, the sample difficulty is dynamically adapted to improve the feature distinction degree; weights of texts and images are distributed according to needs so as to retain core information; according to the method, retrieval result sorting is optimized in combination with time attributes.
Owner:ZHEJIANG UNIV OF FINANCE & ECONOMICS