Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

216 results about "Semantic gap" patented technology

The semantic gap characterizes the difference between two descriptions of an object by different linguistic representations, for instance languages or symbols. According to Hein, the semantic gap can be defined as "the difference in meaning between constructs formed within different representation systems". In computer science, the concept is relevant whenever ordinary human activities, observations, and tasks are transferred into a computational representation.

Accurate micro-crack segmentation method integrating feature fusion and convolution attention

The invention provides a microcrack precise segmentation method integrating feature fusion and convolution attention, and belongs to the field of image processing. According to the method, a crack segmentation network based on an encoder-decoder architecture is constructed, a convolution block attention module is introduced at an encoder end, background noise is adaptively suppressed and obvious characteristics of cracks are enhanced through a channel and space dual attention mechanism, and the method is suitable for the adaptive segmentation of the cracks on the premise of almost not increasing the calculation overhead. The sensitivity of the model to microcracks is improved; a feature fusion module is introduced at a decoder end, and cooperation of low-layer details and high-layer semantics is realized through cross-layer fusion, so that a semantic gap is effectively bridged, detail loss caused by traditional convolution stacking is avoided, and continuity and a complete topological structure of a long and narrow crack are ensured. According to the method, through collaborative optimization of multi-scale feature extraction and an attention mechanism, accurate capture of the saliency features of the crack and effective suppression of complex background interference are realized, and the detection sensitivity and overall segmentation consistency of the micro-crack are remarkably improved.
Owner:DALIAN UNIV OF TECH

Multi-modal data fusion method and system based on large model

The invention discloses a multi-modal data fusion method and system based on a large model, and relates to the technical field of data fusion, and the method comprises the steps: receiving multi-modal data, and carrying out the noise layering filtering and time-space alignment; mapping data to a large model hidden space through each modal lightweight encoder, extracting initial features by using a single-modal pre-training model, and converting the initial features through an adapter to generate a hidden space vector; each modal hidden space vector and a task cue word are received, a fusion feature matrix is dynamically aggregated and generated through a self-attention mechanism and a cross attention mechanism of a large model, and semantic integration of multi-modal information is realized; taking the fusion feature matrix as a soft label, and learning a cross-modal semantic mapping capability through knowledge distillation; the trained model is deployed to the edge through dynamic quantization and pruning optimization, an optimized feature matrix is input into a task-customized lightweight head network, and a final task result is output in combination with multi-modal context information. The method can break through the semantic gap between modals, and improves the fusion efficiency.
Owner:浪潮智慧城市科技有限公司 +1

TransUNet-based medical image segmentation method

The invention discloses a medical image segmentation method based on TransUNet, and belongs to the technical field of medical image segmentation. The method comprises the steps of firstly performing data preprocessing on an original image to obtain preprocessed data; and a DCA attention module is used at a jump joint, so that the problem that a semantic gap exists between characteristics of an encoder and a decoder due to the fact that a simple jump connection scheme is difficult to capture a multi-scale context is solved. The semantic difference leads to redundancy between low-level and high-level features, and finally the segmentation performance is limited. Secondly, a multi-scale boundary sensing module is added to the top layer of the encoder, so that the neural network can better segment the boundary of the target image in the training process; and inputting the preprocessed data into the improved TransUNet model to train the medical image, and outputting an image segmentation result.
Owner:BEIJING UNIV OF TECH

Remote sensing image multi-source heterogeneous data fusion processing method and system

The invention discloses a remote sensing image multi-source heterogeneous data fusion processing method and system, and the method comprises the steps: extracting global information from remote sensing image data through employing a convolutional neural network, extracting image local features through cutting operation, and capturing local feature information in the remote sensing image data; constructing a cross-time-domain attention mechanism for the local feature information through a cyclic matrix to extract mutual information among different modal variables, and screening highly-associated cross-time-domain key information; establishing a cross-time-domain sensing hierarchical aggregation module for the cross-time-domain key information, and obtaining detail information and edge information of the remote sensing image; acquiring global fusion data by adopting an asymptotic fusion strategy, and introducing a loss function to reduce a semantic gap; according to the method, missing information is repaired by adopting an interactive network model of mixed contrast learning, complete real-time remote sensing image multi-source heterogeneous fusion data is obtained, and efficient and high-precision fusion processing of the multi-source remote sensing data is realized by constructing a multi-collaborative deep fusion framework.
Owner:CHINA AERO GEOPHYSICAL SURVEY & REMOTE SENSING CENT FOR LAND & RESOURCES

Robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention

The invention relates to a robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention, and belongs to the technical field of image processing. Aiming at the problems of small target feature loss, semantic gap, background noise interference and the like caused by a fixed convolution kernel scale, one-way feature fusion and a static attention mechanism in an existing unmanned aerial vehicle aerial image target detection method, the method comprises the following steps: constructing a detection model comprising a backbone network, a neck network and a detection head network; a feature rearrangement and extraction module is designed in the backbone network to enhance feature learning, an enhanced double-flow feature fusion pyramid is designed in the neck network to optimize multi-scale feature fusion, and a dynamic multi-scale context attention mechanism is designed in the detection head network to suppress irrelevant background noise. The method effectively improves the accuracy and robustness of small target detection, and achieves a clearer and more stable detection effect in a complex environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Ship user behavior self-learning recommendation system based on large model

The invention relates to a ship user behavior self-learning recommendation system based on a large model, and relates to the technical field of ship informatization. According to the system, through collection and fusion of multi-source heterogeneous ship user behavior data, a large-scale pre-training language model (large model) is utilized to carry out deep understanding and semantic mining on massive ship field text information and user behavior sequences, and a ship field knowledge graph or semantic vector space is constructed. The large model can identify and predict potential demands, behavior patterns and preference changes of ship users, and generates highly personalized, accurate and prospective ship service, product, route or information recommendations in combination with real-time operation data and external environment factors. Besides, a user feedback self-learning mechanism is introduced into the system, recommendation strategies and model parameters are continuously optimized according to interaction behaviors and explicit evaluation of the users in modes of reinforcement learning or continuous learning and the like, and intelligent iteration of the system and continuous improvement of the recommendation effect are achieved. According to the method, the challenges of a traditional recommendation system in the aspects of data complexity, semantic gaps and dynamic demand adaptability in the ship field are effectively solved, and the ship operation efficiency and the user satisfaction degree are remarkably improved.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Attribute decoupling multi-subspace proxy learning text-image pedestrian re-identification method

The invention discloses an attribute decoupling multi-subspace proxy learning text-image pedestrian re-recognition method. The method comprises the following steps: analyzing original text description into a plurality of attribute-level sentences through a fine-grained text reconstruction module; decomposing the global visual features into a plurality of semantic subspaces by adopting an image subspace projection mechanism; a unified multi-granularity subspace loss function is used for training optimization, and the loss integrates global comparison loss, attribute-level comparison loss and subspace diversity regularization loss. According to the method, attribute-level semantics are explicitly decoupled, and fine-grained alignment is realized in a plurality of agent subspaces, so that a semantic gap between a text and an image is effectively bridged, the precision and robustness of cross-modal matching are improved, competitive performance is obtained on standard indexes such as Rank-1 and the like, and the method can be widely applied to the field of text and image processing. And meanwhile, higher convergence speed and better interpretability are shown.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

Electric power operation target detection method based on multi-mode large model knowledge distillation

The invention relates to the field of target detection, and particularly discloses an electric power work target detection method based on multi-modal large model knowledge distillation, which utilizes a vision-language multi-modal large model as a teacher model, and improves the target detection efficiency by expanding prompt word guidance. A high-quality pseudo label and a region-text pair are generated for an unlabeled electric power work image as a supervision signal, and on this basis, through joint optimization of detection loss, feature distillation loss, logic distillation loss and multi-modal contrast learning loss, a lightweight YOLO student model is guided to learn positioning and classification knowledge and to learn a multi-modal contrast learning loss. And deep alignment with the open vocabulary understanding ability of the teacher model is carried out on the feature space and semantic level, so that a semantic gap between closed category detection and open world perception is effectively bridged. Through the mode, the detection precision and generalization ability of the student model on common, rare and even unseen targets in the electric power work scene are remarkably improved.
Owner:MARKETING SERVICE CENT OF STATE GRID HENAN ELECTRIC POWER CO

Intelligent analysis method and system based on multi-source data

The invention relates to the technical field of encrypted communication, and discloses an intelligent analysis method and system based on multi-source data, and the method comprises the steps: obtaining multi-dimensional metadata of an encrypted communication flow, and generating a feature vector set; calculating a state transition probability of a communication session by using a Markov chain model, dividing the communication session according to a time window and mapping the communication session to a discrete state, and generating a session state evolution sequence; analyzing a session state evolution sequence based on a hierarchical topic model, mapping the state sequence to a behavior topic through potential Dirichlet allocation, mapping the behavior topic to an intention category through a hierarchical Dirichlet process, and outputting a communication intention category and a confidence score; according to the method, the limitation that static feature analysis cannot reflect the complete life cycle of the session is overcome, the problem of semantic gaps caused by limited metadata information dimensions is solved, and the accuracy and robustness of intention recognition are improved.
Owner:TANGREN COMM TECH CO LTD

Fault diagnosis method and system based on multi-modal feature fusion and lightweight fine tuning

The invention relates to the technical field of industrial intelligent diagnosis and large language model crossing, and provides a fault diagnosis method and system based on multi-modal feature fusion and lightweight fine tuning. According to the method, multi-source heterogeneous data such as a time sequence, an image and a text of industrial equipment are collected, after feature extraction and cross-modal fusion are conducted, fusion features are adapted to a large language model input space through a projection embedding layer, and then intelligent diagnosis is achieved in combination with a large language model subjected to low-rank self-adaptive fine adjustment. According to the method, the problems of semantic gaps and dimension mismatch of multi-modal industrial data are effectively solved, the diagnosis precision and the model generalization ability are remarkably improved, meanwhile, the requirement for computing resources is greatly reduced, the interpretability of diagnosis results is enhanced through structured output, and an innovative technical path is provided for intelligent operation and maintenance of industrial equipment.
Owner:武汉中云康崇科技有限公司

Multi-modal feature coding and cross-modal adaptive fusion method

The invention provides a multi-modal feature coding and cross-modal adaptive fusion method, which comprises the following steps of: respectively extracting radar features from radar sequence data, extracting visual features from a radar target image, respectively extracting text data and knowledge graph data from text data, and fusing the features of different modals to obtain a multi-modal feature coding and cross-modal adaptive fusion model; specifically, distribution differences among modals are eliminated through statistical normalization, then the modals are projected to a shared feature space to realize dimension unification, and finally efficient cooperation of radar, vision and text modals is realized in combination with semantic alignment guided by a knowledge graph and dynamic weighting of input self-adaption. According to the method, the problems of feature isomerism and semantic gaps in heterogeneous data fusion are effectively solved.
Owner:CHINA SHIPBUILDING LINGJIU HIGH TECH (WUHAN) CO LTD +1

Knowledge graph-based long video key frame retrieval method and device

The invention relates to the technical field of multi-mode intelligent video understanding, and provides a long video key frame retrieval method and device based on a knowledge graph. Through the processes of frame-level subtitle generation, frame-level knowledge graph construction, similarity video segmentation and fragment and abstract generation, a long video knowledge graph construction assembly line is constructed, and structured modeling of long video semantic content is realized. By setting a two-stage retrieval mechanism, higher retrieval precision can be obtained while the efficiency is ensured. Vector matching and multi-hop neighbor extension are carried out on a unified knowledge graph, and a node set strongly related to the problem is positioned, so that the semantic gap between the natural language problem and a structured graph is reduced, and the accuracy of key frame selection is improved. By setting an iterative retrieval mechanism, the retrieved key frame can be used as a basis for answering a question text to the maximum extent.
Owner:NAT UNIV OF DEFENSE TECH

Network structure combining residual jump enhancement and multi-dimensional interactive attention

The invention belongs to the field of pulmonary embolism detection and segmentation, and discloses a residual error jump enhancement and multi-dimensional interactive attention combined network structure, which comprises an encoder, a decoder, a jump connection fusion module and a bottleneck module which are respectively responsible for feature extraction, spatial restoration, multilayer semantic fusion and modeling of a long-distance dependency relationship. The AGRS module is fused with a residual jump enhancement structure and a multi-dimensional interaction quadruple attention module; the CDB-ASPP module realizes multi-scale context information efficient fusion through five parallel convolution branches with different expansion rates, combines a global-local enhancement strategy, and considers both a global receptive field and a local detail perception capability; and a dual attention Transform module is used for enhancing the global feature modeling capability. According to the method, an effective solution is provided for solving key problems such as semantic gaps, context information missing and small target detection sensitivity, and important reference and enlightenment are provided for deep learning-based pulmonary embolism auxiliary diagnosis research.
Owner:INNER MONGOLIA UNIV OF SCI & TECH +1

Remote sensing image-text retrieval method based on knowledge enhancement and asymmetric structure

The invention relates to a remote sensing image-text retrieval method based on knowledge enhancement and an asymmetric structure, and belongs to the technical field of remote sensing image-text cross-modal retrieval. Comprising the following steps: inputting a to-be-retrieved remote sensing image and text into a trained vision-language basic model with knowledge enhancement and an asymmetric structure, and realizing cross-modal retrieval through model processing; and obtaining a retrieval result, obtaining matching retrieval output of the remote sensing image and the text, and completing a cross-modal retrieval task. According to the method, the problem that the retrieval performance is limited due to the fact that significant information asymmetry exists between the remote sensing image and the text modality at present is solved. According to the method, the cross-modal asymmetric Kolmogorov-Arnod adapter fine tuning method is designed, so that efficient modal fine-grained shared feature learning is realized; meanwhile, knowledge is extracted from ConceptNet and a remote sensing knowledge graph, and a knowledge enhancement sentence is generated to enrich text semantics, so that a semantic gap between modals is bridged, and remote sensing image-text retrieval performance is improved.
Owner:YUNNAN NORMAL UNIV

Luggage material identification method and system based on multi-modal fusion knowledge distillation

The invention relates to a luggage material identification method and system based on multi-modal fusion knowledge distillation. The method comprises the steps of obtaining image data and point cloud data of a to-be-detected target surface; constructing a multi-modal teacher model to obtain image texture features and point cloud geometric features; obtaining a teacher object query; outputting a material category score and a bounding box position coordinate; constructing a lightweight student model, and generating student object query, material category prediction and bounding box position prediction; characteristic distillation loss is designed for characteristic distillation, and a total loss joint training lightweight student model is constructed; actually operating the trained lightweight student model in a luggage detection scene of an airport luggage turntable; calculating a stacking score and mapping the stacking score into a stacking label; according to the method, the semantic gap between the perception recognition module and the downstream planning strategy module is effectively linked, and the contradiction between the insufficient precision of traditional single-mode perception and the high cost of multi-mode deployment is solved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Multi-modal data semantic alignment method and system based on knowledge graph embedding

The invention discloses a multi-modal data semantic alignment method and system based on knowledge graph embedding, relates to the technical field of inspection and detection, and solves the problems of low data accuracy and interpretability after multi-modal data fusion. According to the embodiment of the invention, through a knowledge graph embedding and semantic alignment mechanism, the semantic gap problem of cross-modal data is solved, and accurate association of multi-modal data in an inspection and detection scene is realized; the optimized fusion vector obtained by fusion enhances the interpretability and consistency of the data, and improves the detection accuracy and robustness of a subsequent detection model.
Owner:GUILIN UNIV OF ELECTRONIC TECH

High-standard farmland scene recognition method based on deep neural network

The invention discloses a high-standard farmland scene recognition method based on a deep neural network, and relates to the technical field of remote sensing image processing and agricultural information, and the method comprises the steps: 1, constructing a high-standard farmland scene sample library with prior significance, the high-standard farmland scene sample library comprises high-standard farmland sample features and high-standard farmland sample scale quantities; 2, constructing a high-standard farmland scene recognition model based on multi-source data fusion; 3, predicting a result, and outputting a scene category graph; according to the high-standard farmland scene recognition method based on the deep neural network provided by the invention, the problem that the prior art stays at a pixel level or an object level, depends on a single threshold value or shallow learning and has a semantic gap is solved.
Owner:CHINA AGRI UNIV

Method and system for text retrieval of picture archives based on cross-modal feature alignment

The invention discloses a method and a system for retrieving a picture file through a text based on cross-modal feature alignment, and the method comprises the steps: extracting text and picture features, and guaranteeing that a high-value mode contributes to a higher weight based on multi-modal attention weighted fusion; calculating a dynamic temperature coefficient through initial semantic similarity of positive and negative samples to construct a loss function for multi-modal contrast learning, and mapping original features of a text and an image to a unified space to obtain alignment features; according to the method, the picture archives are retrieved through texts, the image-text similarity, the time sequence weight and the core area proportion weight are comprehensively considered, optimization sorting of time sequence perception is carried out, and the most matched picture archives are obtained. According to the method, a text-image cross-modal semantic gap is solved, so that semantic alignment of two types of features in a unified space is realized; according to the method, the sample difficulty is dynamically adapted to improve the feature distinction degree; weights of texts and images are distributed according to needs so as to retain core information; according to the method, retrieval result sorting is optimized in combination with time attributes.
Owner:ZHEJIANG UNIV OF FINANCE & ECONOMICS

Wind turbine blade defect detection method based on improved RT-DETR

The invention belongs to the technical field of target detection, and discloses a wind turbine blade defect detection method based on improved RT-DETR, and the method introduces an SWRepBlock module in a backbone network, thereby effectively solving the problem that a conventional network fixed receptive field is difficult to adaptively extract different scale target features. In addition, according to the method, an AIFI module in the encoder is replaced by a DyT-AIFI module, so that the semantic understanding effect of long-distance feature interaction is improved. In addition, according to the method, a CAA-HSFPN module is further introduced into an encoder, and the problem that a traditional feature pyramid semantic gap and feature importance distinguishing is insufficient is effectively solved. According to the wind turbine blade defect detection method provided by the invention, the detection capability of tiny damage under the background of complex surface texture of the wind turbine blade is remarkably improved, and multiple types of defect characteristics such as cracks, erosion and oil leakage can be effectively identified.
Owner:SHANDONG UNIV OF SCI & TECH

Method for accessing standard knowledge graph to PLM platform

The invention discloses a method for accessing a standard knowledge graph to a PLM platform, and the method comprises the following steps: S1, environment initialization: deploying an AP I gateway in the PLM platform, and setting an address and a login credential of a standard knowledge graph service in the AP I gateway; s2, metadata mapping: establishing a mapping relationship between the PLM data model and a standard knowledge graph ontology; s3, service registration: registering a PLM service event trigger service through an AP I interface; and S4, calling service: calling reasoning service based on the PLM data and the marker knowledge graph based on an event-driven mechanism. The method has the advantages that seamless connection between the standard knowledge graph and the PLM platform is achieved by designing an interface, semantic association between a structured data model in the PLM platform and a standard knowledge graph body is established by adopting dynamic body modeling and a hybrid (rule + artificial intelligence) entity mapping technology, and the semantic gap problem between heterogeneous data is solved.
Owner:TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI

Sheet metal part manufacturability reasoning method based on space-semantic map alignment

The invention discloses a sheet metal part manufacturability reasoning method based on space-semantic map alignment, and relates to the field of manufacturing-oriented design evaluation and industrial knowledge reasoning, and the method comprises the following steps: carrying out geometric analysis on a CAD geometric model of a sheet metal part to be evaluated; abstracting the geometric features and the topological / metric spatial relationship thereof into a computable spatial semantic graph; performing semantic analysis on the process specification described by a natural language, converting the process rule into formalized logic check expression by using a large language model through context learning, and generating an executable domain-specific language check script; executing the script on a spatial semantic graph, realizing deterministic reasoning through graph matching and attribute verification, and completing accurate mapping and violation detection of text rules and geometric features; and outputting an interpretable diagnosis result containing violation feature positioning, triggering rules and numerical evidence. In order to solve the problems that a process rule'natural language-geometric model 'has a semantic gap, a traditional rule system is poor in adaptability, and an end-to-end learning method is high in data dependence and cannot be explained, a new rule can be quickly adapted under the condition that a large amount of data does not need to be labeled and a model does not need to be retrained; the method can accurately identify the violation of the micro-size and spatial relationship, and has reasoning preciseness, interpretability and engineering availability.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Image emotion prediction method based on double attention and diversified knowledge distillation

The invention provides an image emotion prediction method based on double attention and diversified knowledge distillation, and the method comprises the steps: constructing a channel attention module based on a channel attention mechanism, constructing a feature fusion module based on a space attention mechanism, obtaining an image emotion data set, and carrying out the image emotion prediction. The emotion image in the image emotion data set is preprocessed to obtain a preprocessed emotion image, and the preprocessed emotion image is processed through ConvNeXt, a channel attention module and a feature fusion module in sequence to obtain fused multi-scale features; and processing the fused multi-scale features through an improved Transform encoder, average pooling, a full connection layer and a classifier in sequence to obtain probability distribution output by the classifier. According to the method, the characterization capability of the multi-scale features is further enhanced by increasing channel attention and space attention, semantic gaps among the multi-scale features are reduced, key local information is reserved, and multi-scale context information is fully mined.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Image editing method based on characteristic frequency domain fusion, program product and medium

The invention discloses an image editing method based on characteristic frequency domain fusion, a program product and a medium, and belongs to the field of image editing. Step-by-step inversion is carried out on the encoded submerged space features to obtain a series of submerged space features of which the noise-containing degree is gradually increased; iterative denoising is carried out to obtain a final denoising feature; in the previous part of iterative denoising, fusing the high-frequency features of the reconstructed stream and the low-frequency features of the generated stream, introducing randomness, and meanwhile, adopting a residual structure based on the features of the generated stream; on the basis, a certain degree of noise is injected into query features and key features of the self-attention mechanism and the cross-attention mechanism. According to the method, effective fusion of features between the reconstructed stream and the generated stream can be realized, feature semantic gaps of different branches are reduced, content editing and background keeping are better performed, and the problems of poor content editing semantic conformity degree and insufficient background keeping degree in a non-rigid editing task in an existing method are effectively solved.
Owner:HUAZHONG UNIV OF SCI & TECH

Intention recognition system based on multimode agent and data annotation

PendingCN121959435ASemantic gapMemory bank
The invention relates to an intention recognition system based on a multi-mode agent and data annotation, in particular to the field of artificial intelligence, according to the scheme, a semantic gap between different modes is effectively bridged by constructing a unified semantic projection space, the reliability of information of each mode is quantitatively and dynamically evaluated by using uncertainty, weighted fusion is performed on the basis, and the intention recognition accuracy is improved. A cross-modal semantic memory network and an online learning mechanism are innovatively introduced, so that the scheme not only can make more accurate intention inference by integrating historical experience and real-time perception, but also can continuously optimize own confidence assessment ability and a memory bank in an interaction process; therefore, the method has higher adaptability and robustness in the environment with noise, missing information and dynamic change.
Owner:GUANGDONG BAOGU TECH CO LTD

Multi-modal feature alignment semantic fusion method based on deep learning

The invention relates to the technical field of data processing, in particular to a multi-modal feature alignment semantic fusion method based on deep learning, which comprises the following steps of: acquiring a multi-modal data fragment through an event triggering acquisition mechanism, generating a cross-modal time sequence association confidence coefficient matrix by adopting a multi-granularity time sequence modeling network, and performing semantic fusion on the cross-modal time sequence association confidence coefficient matrix; a bidirectional iteration alignment module is used for realizing fine alignment of feature sequences, adaptive semantic fusion is performed through a dynamic cross-modal Transform architecture, emotional state probability distribution is output in combination with an online emotion classification and delay prediction mechanism, and finally a system parameter optimization closed loop is constructed. According to the method, the complex interaction problem of the multi-modal data in the aspects of feature distribution heterogeneity, time sequence asynchronism and semantic gap is effectively solved, the emotion calculation precision is improved, and the response delay of the mental health service is reduced.
Owner:LUSHAN COLLEGE OF GUANGXI UNIV OF SCI & TECH

Medical image treatment method and system for multi-source heterogeneous data

The invention discloses a medical image treatment method and system for multi-source heterogeneous data, and relates to the field of medical image treatment, and the method comprises the steps: firstly, respectively extracting an image embedding vector and a text embedding vector from original medical image data and a description text thereof through a deep learning model; then, the vectors of the two different modes are fused, and a unified semantic embedding vector is formed; based on the unified vector, through semantic similarity calculation with a standard term library, candidate mapping can be automatically generated, and dependence on rigid artificial rules is eliminated. More importantly, a closed-loop mechanism of manual auditing-feedback learning is introduced in the scheme, high-confidence mapping is adopted automatically, low-confidence mapping is audited by experts, and an auditing result is absorbed into a mapping knowledge base, so that the system has continuous learning and self-evolution capabilities, and a semantic gap of cross-mechanism data can be eliminated more intelligently and more accurately.
Owner:ZHEJIANG FEITU IMAGING TECH CO LTD

A State-Space Model-Based Medical Image Segmentation Method for Abdominal Multi-Organs

PendingCN122089756Aresolve integritySolve the problem of mutual invasion between organsImage analysisBiological modelsComputation complexityFeed forward network
This invention belongs to the field of medical image processing technology, specifically relating to a method for abdominal multi-organ medical image segmentation based on a state-space model. Addressing the shortcomings of existing technologies such as the difficulty of CNNs in modeling long-distance dependencies, the high computational complexity of Transformers, and the large semantic gaps, incomplete segmentation contours, and easy organ encroachment issues in VM-UNet skip connections, this solution makes key improvements: It constructs an improved VM-UNet-Skip architecture, placing skip connections before downsampling to reduce the semantic gap between the small decoder features; it introduces a multi-scale global-local information aggregation module, capturing global anatomical dependencies through multi-head Mamba units and enhancing local detail representations with a convolutional gated feedforward network; and it relies on the VMamba encoder-decoder to achieve efficient feature extraction and reconstruction. The method flow includes: constructing initial feature representations through patch embedding layers, extracting multi-scale hierarchical features through the encoder, enhancing features through an information fusion module, and fusing and mapping the enhanced features to obtain the segmentation result through the decoder. This invention effectively improves the segmentation accuracy of abdominal multi-organs, solves the problems of insufficient contour integrity and organ encroachment, while reducing the number of model parameters and computational complexity, providing reliable technical support for clinical diagnosis and surgical planning.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A method for generating a cross-modal image of sugar metabolism based on a WARPNet framework

PendingCN122368266AData setSemantic gap
The present application relates to the technical field of medical image computing and artificial intelligence, in particular to a sugar metabolism cross-modal image generation method based on a WARPNet framework, based on the WARPNet framework, through three core steps of multi-scale feature coding, weighted attention cross-modal reasoning and context perception generation, precise synthesis from MRI to PET is realized. The innovation lies in introducing a learnable sugar metabolism semantic prior weight in cross-modal reasoning, modulating feature attention, thereby explicitly guiding the model to focus on key metabolic areas, effectively solving the semantic gap problem between anatomical and functional modalities. The method improves the quality, physiological reasonableness and cross-dataset generalization ability of the synthesized image, and has application potential in auxiliary diagnosis, treatment planning and medical research.
Owner:XUANWU HOSPITAL OF CAPITAL UNIV OF MEDICAL SCI

Infrared and visible image fusion method based on semantic guidance and scene graph reasoning

PendingCN122656871AEnhance deep understandingimprove consistencyPattern recognitionData set
The present application belongs to the field of image fusion and multi-modal image processing, and in particular to an infrared and visible light image fusion method based on semantic guidance and scene graph reasoning, comprising: S1, preparing a data set: preparing three visible light and infrared image fusion data sets, data set one and data set two are used for network training, fine tuning, and data set three is used for testing. The present application constructs a visible light and infrared image fusion model based on the infrared and visible light image fusion method based on semantic guidance and scene graph reasoning, designs a scene graph construction module based on self-supervised fine tuning, a semantic guidance deformable alignment module and a dynamic calculation and fusion module. Through experimental data, it is proved that the present application can effectively solve the problems of geometric misalignment and semantic gap, gradually improve the registration accuracy of key semantic regions, and alleviate the phenomenon of insufficient representation of pre-trained models on cross-modal data. The present application shows good performance in the qualitative and quantitative evaluation of visible light and infrared image fusion tasks.
Owner:CHINESE PEOPLES LIBERATION ARMY UNIT 63869

Neural symbol hybrid reasoning-based multi-modal clinical scientific research data processing method and system

The invention discloses a multi-modal clinical scientific research data processing method and system based on neural symbol hybrid reasoning, and belongs to the field of medical informatization and artificial intelligence. The method comprises the steps of receiving a natural language analysis instruction; analyzing the instruction into a structured formal problem description data object through a language understanding unit by utilizing a nerve-symbol hybrid inference engine, and selecting a data analysis algorithm according to an expert rule base through a symbol inference unit; then, generating a directed acyclic graph analysis flow based on the selected algorithm; and finally, according to the directed acyclic graph analysis flow, processing the multi-modal clinical scientific research data stored in the database, and generating an analysis report containing quantitative indexes. The technical problems that in the prior art, a clinical scientific research data analysis process is split, efficiency is low, and a semantic gap exists in man-machine interaction are solved, automation and intelligence of scientific research analysis are achieved, and the efficiency, depth and scientificity of data processing are remarkably improved.
Owner:GUANGDONG HOSPITAL OF TRADITIONAL CHINESE MEDICINE