Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

29results about How to "Bridging the Semantic Gap" patented technology

Medical image segmentation method, system and equipment based on multi-attention and multi-scale fusion

The invention discloses a medical image segmentation method, system and device based on multi-attention and multi-scale fusion, and relates to the technical field of image segmentation, and the method comprises the steps: constructing an MAMF-Net model which comprises an encoder and a decoder which are in multi-layer jump connection; the encoder adopts a hybrid architecture of convolution and Transform, and is integrated with a self-adaptive expansion convolution method; the decoder integrates dual-channel attention gating and a multi-scale global channel feature enhancement method to enhance features transmitted by jump connection, and combines features extracted by the encoder to fuse and reconstruct a segmentation result; training the MAMF-Net model by adopting the historical medical image sample set to obtain a medical image segmentation model; and obtaining any medical image to be identified and inputting the medical image to the medical image segmentation model, and determining a corresponding segmentation result. The problem of insufficient fusion of global semantics and local details in medical image segmentation is solved.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Medical image treatment method and system for multi-source heterogeneous data

The invention discloses a medical image treatment method and system for multi-source heterogeneous data, and relates to the field of medical image treatment, and the method comprises the steps: firstly, respectively extracting an image embedding vector and a text embedding vector from original medical image data and a description text thereof through a deep learning model; then, the vectors of the two different modes are fused, and a unified semantic embedding vector is formed; based on the unified vector, through semantic similarity calculation with a standard term library, candidate mapping can be automatically generated, and dependence on rigid artificial rules is eliminated. More importantly, a closed-loop mechanism of manual auditing-feedback learning is introduced in the scheme, high-confidence mapping is adopted automatically, low-confidence mapping is audited by experts, and an auditing result is absorbed into a mapping knowledge base, so that the system has continuous learning and self-evolution capabilities, and a semantic gap of cross-mechanism data can be eliminated more intelligently and more accurately.
Owner:ZHEJIANG FEITU IMAGING TECH CO LTD

An unmanned aerial vehicle target detection method based on feature fusion DINO

The application discloses a kind of unmanned aerial vehicle target detection methods based on feature fusion DINO, belong to target detection technical field. Including the following steps: obtaining original unmanned aerial vehicle image and pre-processing;The data after pre-processing is sent into feature extraction network Backbone and image feature extraction is carried out, output multi-scale feature map, and it is input into integrated collection feature Neck module to strengthen spatial feature and semantic feature fusion, output enhanced feature map;Fusion enhanced feature map is sent into encoder, and global feature information is highlighted through the multi-scale receptive field of hollow convolution path, finally, the class and boundary box prediction of unmanned aerial vehicle target are generated after decoder generation.The present application is based on DINO architecture, and a high-precision unmanned aerial vehicle target detection model is constructed by a multi-scale hollow convolution fusion method.The detection accuracy and robustness of small-size unmanned aerial vehicle targets are improved by effectively extracting multi-scale features and recursively fusing.
Owner:SICHUAN JIUQIANG COMM TECH CO LTD

Multi-source fusion log compression method and device for anomaly detection

ActiveCN119420534BBridging the Semantic Gapreduce dependenceSecuring communicationDomain nameAlgorithm
This application discloses a multi-source fusion log compression method and apparatus for anomaly detection, belonging to the field of anomaly detection technology. The multi-source fusion log compression method for anomaly detection includes: generating an audit origination graph corresponding to the system audit log, an application origination graph corresponding to the application log, and a domain name origination graph corresponding to the domain name system log based on the system audit log corresponding to the electronic device, the application log corresponding to the target application in the electronic device, and the domain name system log corresponding to the electronic device; fusing the domain name origination graph into the application origination graph based on the domain name nodes in the application origination graph to obtain a sub-fused origination graph; fusing the audit origination graph into the sub-fused origination graph based on the event nodes in the audit origination graph to obtain a fused origination graph; and performing anomaly detection based on the fused origination graph. The multi-source fusion log compression method for anomaly detection in this application can alleviate the problems of semantic gap and dependency explosion.
Owner:INST OF ADVANCED TECH UNIV OF SCI & TECH OF CHINA

Small target double-stage detection and defect evaluation method for fence structure

The invention discloses a fence structure-oriented small target double-stage detection and defect assessment method, which comprises the following steps of: inputting preprocessed RGB image data of a fence into a trunk convolution module of a defect risk assessment model to obtain first feature maps with different scales, fusing the first feature maps with polarization physical feature maps of the fence respectively, and obtaining a defect risk assessment result through a plurality of FPN modules; performing feature fusion on the fused feature map by adopting a cross-level feature fusion strategy and a cascade connection strategy, and sequentially outputting a fused second feature map to a channel attention module, a space attention module, a texture attention module and a region candidate module to obtain a candidate box set; when a candidate frame meeting the image reconstruction requirement exists, candidate frame reconstruction is carried out through an image super-resolution reconstruction module, RGB image data and a polarization physical feature map, then a bounding box, a defect category and a pixel mask are determined through an RoI feature processing module and a defect detection and segmentation module, and defect risk assessment information is obtained in combination with a geometric model of a fence.
Owner:HESHENG ZHIHUI (XIAN) TECHNOLOGY CO LTD

Deep network structure based abrasive plate bonding detection method and system

The application discloses a deep network structure abrasive plate joint detection method and system, and a built-in multi-dimensional adaptive feature enhancement module of the detection method, which comprises a multi-scale feature extraction unit, a channel adaptive attention unit, a spatial adaptive attention unit and a residual connection unit. The multi-scale feature extraction unit is responsible for capturing feature information under different granularities. The channel adaptive attention unit strengthens the expression of key feature channels through weight distribution. The spatial adaptive attention unit is used for highlighting the spatial region related to the abrasive plate joint in the feature map. The residual connection unit fuses the enhanced features and the original input, which not only guarantees the effective propagation of the gradient, but also simplifies the learning process of the model. Through the synergistic effect of multi-scale perception and channel spatial double attention mechanism, the module can effectively enhance the weak and irregular feature performance of the abrasive plate joint in the complex background, thereby significantly improving the recognition ability and robustness of the detection system.
Owner:CHANGSHA RES INST OF MINING & METALLURGY CO LTD

Mixed identifier generation type recommendation method and system based on local collaborative context

The invention belongs to the technical field of artificial intelligence, and particularly relates to a mixed identifier generation type recommendation method and system based on local collaborative context, and the method comprises the steps: obtaining a user historical interaction sequence, and carrying out the multi-granularity clustering of users, and obtaining a group of each user; constructing a static identifier based on the article content, constructing a dynamic identifier in combination with the user group information and the local interaction sequence, and fusing to generate a mixed identifier of each article; constructing an instruction fine tuning task, embedding group information into an instruction prompt word, and learning by using a large language model to generate a mixed identifier of a next article from a historical sequence; and generating a recommendation result based on the optimized large language model according to the historical sequence of the target user and the group information thereof. According to the method, the local collaborative context is fused through multi-granularity group estimation, and the mixed article representation is constructed in combination with static and dynamic identifiers, so that the problems of single user interest modeling and semantic deficiency of article representation are effectively solved, and the accuracy and personalized level of generative recommendation are improved.
Owner:QINGDAO UNIV OF SCI & TECH

User circle selection method and related product

The invention discloses a user circle selection method and a related product. In the scheme, a user circle selection instruction in a natural language form is obtained, and the user circle selection instruction is converted into a query vector; based on the query vector, performing semantic retrieval in a user vector library to obtain a target user set; the user vectors in the user vector library are constructed based on the user behavior data and the brand knowledge base. Compared with the prior art that the user circle selection has the problems of label semantic deficiency and sparse behavior data, and the accuracy of the user circle selection result is low, the user circle selection method and device have obvious advantages.
Owner:XIAMEN NANXUN CO LTD

A lightweight service-aware SRv6 satellite routing method for smart grid

PendingCN122293156ABridging the "semantic gap"Eliminate the "semantic gap"
This invention relates to a lightweight service-aware SRv6 satellite routing method for smart grids. It includes: at the service semantic encapsulation layer, classifying grid service flows and generating structured metadata, encapsulating them into SRv6 data packets carrying Power Service Identifiers (PSIDs); at the intelligent policy control layer, parsing the PSIDs and calling a dedicated path calculation engine based on the service type, combining predicted topology to calculate the end-to-end path, encoding it into a PSID list for distribution; at the policy execution and forwarding layer, satellite nodes parse the PSID list for forwarding and execute corresponding policies based on the service semantic domain; simultaneously, satellite nodes monitor link status in real time, triggering local fast rerouting in case of faults, and feeding back the status to the control layer to form a closed-loop optimization. This invention achieves direct mapping between service requirements and forwarding policies through PSIDs, and combined with "one policy per category" path calculation and a local fast self-healing mechanism, solves the challenges of reliability and accurate QoS assurance for differentiated grid service transmission in highly dynamic satellite networks.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Contrastive learning code search method with reinforced multimodal semantics

ActiveCN117668159BSimple structureRich language information representation
The present application relates to a code search method based on contrastive learning of reinforced multimodal semantics, and belongs to the field of natural language processing and machine learning. Firstly, the code snippet is represented as token sequence, abstract syntax tree and program expression graph three modalities, and BERT model is used to generate the feature vectors of each modality and splice them into joint code feature vectors. Then, a contrastive loss function is constructed to reduce the distance between the query statement and the corresponding code snippet in the feature space. Finally, the cosine similarity is used to calculate the distance between the query statement feature vector and the joint code feature vector and sort them, and the code search result is output. The present application proposes a code search method based on contrastive learning of reinforced multimodal semantics to solve the problem that the existing method does not fully extract the code structure features and there is a semantic gap between the query statement and the code snippet, thereby improving the accuracy of code search.
Owner:BEIJING INST OF TECH

Body-equipped intelligent robot task decomposition system and method based on physical state deduction

The invention discloses a task decomposition system and method for an intelligent robot with a body based on physical state deduction, and relates to the technical field of artificial intelligence robots. The system comprises a task instruction receiving and initial decomposition module, a state sensing and formatting input module, a task feasibility deduction module, a feedback and optimization closed loop module and a decision and output module. According to the method, a technical framework of sensing, deducing and optimizing a closed loop is constructed, so that the robot can perform virtual rehearsal on a task sequence before executing actual physical actions, thereby identifying environment state conflicts and execution path risks in advance, and fundamentally improving the task execution mode from passive response depending on trial and error, and improving the task execution efficiency. The method is converted into active reliable planning based on physical state deduction, and the adaptability, the safety and the deployment efficiency of the robot for executing complex tasks in an open environment are remarkably improved.
Owner:GUANGZHOU SHUNQING ZHIHE TECHNOLOGY CO LTD

Training methods, devices, equipment, and storage media for self-supervised learning models

This application discloses a training method, apparatus, device, and storage medium for a self-supervised learning model, belonging to the field of computer and internet technology. The method includes: acquiring a sample set; for a target text sample in the sample set, concatenating the target text sample with other text samples in the sample set to generate a first negative sample corresponding to the target text sample; and using the first negative sample to perform self-supervised training on a text feature extraction model; wherein the text feature extraction model is used to obtain feature information of the input text based on the input text, in order to match retrieval text with semantically similar characteristics to the input text. This application improves the text feature extraction model's ability to distinguish text information with small semantic differences, thereby enhancing the retrieval capability of the text feature extraction model.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A cargo automatic classification method and system based on bilateral semantic mapping and hybrid retrieval

PendingCN122262337Aprecision recallBridging the Semantic GapDigital data information retrievalBiological modelsDigital dataLinguistic model
The application relates to the technical field of electric digital data processing, and discloses a cargo automatic classification method and system based on bilateral semantic mapping and mixed retrieval, which comprises the following steps: obtaining original cargo name data and a standard name database, carrying out text cleaning and structured splitting, reconstructing semantics of the two ends of data by using a large language model, projecting semantic features into a high-dimensional vector space by a vector coding model to calculate similarity and recall a candidate set, extracting fine-grained interaction features by using a cross encoder, combining character matching statistics to calculate a reordering score, verifying consistency of a target item based on a logical mutual exclusion matrix, and outputting classification data. The application effectively suppresses the semantic reconstruction illusion of a generative model, enhances the matching certainty of heterogeneous data through a hidden feature fusion and a logical verification mechanism, and ensures the industrial-level reliability of model-sensitive cargo classification.
Owner:CHANGSHA PURAN NETWORK TECH CO LTD

Dry eye detection method and system based on multi-modal prior and generalized contrast learning

PendingCN122266733AAddressing the lack of medical reasoning skillsImprove discrimination abilityMedical data miningImage analysisCosine similarityFeature extraction
This invention relates to the field of medical image analysis technology, specifically to a method and system for dry eye detection based on multimodal prior and generalized contrastive learning. The method includes: first, acquiring typical images, descriptive text, and relationship graphs of various dry eye diseases; extracting and stitching multimodal features to construct a medical prior knowledge base; second, extracting visual features from the anterior segment of the eye image and retrieving corresponding semantic features from the knowledge base according to disease labels, then projecting and aligning the two; third, constructing a medical similarity matrix based on multimodal features, mining difficult negative samples, and performing generalized contrastive learning on positive and ordinary negative samples to train a visual feature extraction network; finally, inputting the test image into the network for tear river segmentation, calculating the tear river height and pupil-tear river distance, and combining the cosine similarity of visual features and various semantic projection features to output the probability that the test image belongs to each dry eye disease level. This invention improves the accuracy of dry eye detection.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

A meteorological satellite observation data anomaly detection method based on a multi-modal semantic enhancement network

PendingCN122595134Aprecise positioningBridging the semantic gap
This invention discloses a method for anomaly detection in meteorological satellite observation data based on a multimodal semantic enhancement network. The method includes: acquiring multi-channel observation samples from meteorological satellites; constructing a multimodal semantic enhancement network, which includes a feature extraction network and a semantic mining network; training the multimodal semantic enhancement network; preprocessing the meteorological satellite multimodal observation data to be detected and inputting it into the multimodal semantic enhancement network to generate reconstruction errors for each modality corresponding to the feature extraction network and semantic reconstruction errors corresponding to the semantic mining network; determining the anomaly score of the meteorological satellite multimodal observation data to be detected based on the reconstruction errors of each modality and the semantic reconstruction errors; and outputting the anomaly channel, anomaly area heatmap, and alarm level if an anomaly is confirmed. Therefore, without relying on the labeling of anomaly samples or prior knowledge of anomaly types, this method can identify not only known anomaly types but also unknown anomaly types not seen during the training phase.
Owner:NAT SATELLITE METEOROLOGICAL CENT

Large language model evaluation method, device, equipment and storage medium

This application discloses a method, apparatus, device, and storage medium for evaluating large language models, relating to the field of artificial intelligence technology. The method includes: acquiring task configuration information and candidate evaluation model information; determining target evaluation model information based on stability evaluation index information and alignment evaluation index information of the candidate evaluation model information; injecting corresponding semantic evaluation elements into the model prompt word template of the target evaluation model information based on the task configuration information to determine semantic enhancement information of the target evaluation model; and generating corresponding evaluation results based on the semantic enhancement information of the target evaluation model. This application dynamically determines the optimal evaluation model through stability evaluation index and alignment evaluation index, and injects semantic evaluation elements of the business scenario into the corresponding model prompt word template to obtain evaluation results. It coordinates model selection and prompt enhancement, eliminates the semantic gap between general templates and business requirements, and improves the model evaluation's discriminative power, semantic consistency, and reproducibility.
Owner:CHINA MERCHANTS BANK

A hierarchical latent inference recommendation and model training method based on a large language model

PendingCN122711943AClear training goalsImprove generalization ability
The application belongs to the technical field of recommendation systems, and particularly relates to a hierarchical latent inference recommendation based on a large language model and a model training method. A category level and an item identification level hierarchical representation are constructed for items, a recommendation model is built based on a large language model, the recommendation model performs hierarchical latent inference according to a user historical interaction sequence, generates a category level latent state, and then generates a user collaborative filtering vector matching an item identification level representation; a phased progressive training strategy is adopted, and real category level representations and model self-generated category level latent state supervision vectors are used to generate vectors respectively. In the recommendation stage, the user historical interaction sequence is input into the trained model, the hierarchical latent inference is performed to output the user collaborative filtering vector, similarity calculation is performed between the user collaborative filtering vector and an item library identification level collaborative representation, and recommended items are determined.
Owner:UNIV OF SCI & TECH OF CHINA

An athlete action biomechanics evaluation method, system, device and medium based on deep learning

PendingCN122658554ABridging the Semantic Gapensure fidelityPattern recognitionAlgorithm
The application relates to an athlete action biomechanics evaluation method, system, device and medium based on deep learning, the method comprising: obtaining a motion signal sequence by time synchronization and cycle segmentation of a multi-modal technical action signal of an athlete, generating a joint representation tensor by kinematics and dynamics feature fusion, obtaining a bottom feature map by spatiotemporal coding of fused individual morphological parameters, performing cross-attention decoding based on a pre-constructed biomechanics semantic knowledge space to generate a semantic state matrix and a concept activation sequence with a concept node as a query vector, generating an evaluation result along a causal topological structure, and outputting an interpretable biomechanics evaluation report through natural language generation. The method can improve the interpretability, causal tracing ability and training guidance value of the athlete action biomechanics evaluation.
Owner:ZAOZHUANG VOCATIONAL COLLEGE OF SCI & TECH

Question method and device based on double-layer thinking chain and CTE operator library, computer equipment and medium

The invention relates to a question number method and device based on a double-layer thinking chain and a CTE operator library, computer equipment and a medium. The method comprises the steps of obtaining question number information and a related dynamic domain knowledge set, loading the question number information and the dynamic domain knowledge set into a prompt word template, generating a first prompt context and inputting the first prompt context into a large model, and obtaining query demand information; installing question number information, query demand information, an entity view file, a relation view file and a table selection example into a prompt word template, generating a second prompt context, inputting the second prompt context into the large model, and obtaining query plan information; and splicing the query plan information and a preset CTE operator library to a second layer of thinking chain cue word template to generate a third prompt context, inputting the third prompt context into the large model, obtaining an SQL output by the large model, and executing the SQL to obtain a question number result. According to the method, the accuracy of the SQL generated by the large model can be improved, the complex index calculation step can be traced back, and the repeatability of the generated result is high.
Owner:CHINA ASSET MANAGEMENT CO LTD

A medical image segmentation method, system, and device based on multi-attention and multi-scale fusion

This application discloses a medical image segmentation method, system, and device based on multi-attention and multi-scale fusion, belonging to the field of image segmentation technology. The method includes: constructing a MAMF-Net model, including an encoder and decoder with multiple skip connections; the encoder adopts a hybrid architecture of convolution and Transformer, and integrates an adaptive dilated convolution method; the decoder integrates dual-channel attention gating and multi-scale global channel feature enhancement methods to enhance the features transmitted by skip connections, and combines the features extracted by the encoder to fuse and reconstruct the segmentation result; training the MAMF-Net model using a historical medical image sample set to obtain a medical image segmentation model; acquiring any medical image to be identified and inputting it into the medical image segmentation model to determine the corresponding segmentation result. This application solves the problem of insufficient fusion of global semantics and local details in medical image segmentation.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Multi-scene-oriented voice-driven AI large model intention analysis method and action execution system

PendingCN121963734AAccurately differentiate needsAccurately distinguish operational needsSpeech recognitionInference methodsSemantic gapEngineering
The invention relates to the technical field of intention recognition, in particular to a multi-scene-oriented voice-driven AI large model intention analysis method and action execution system, which can accurately distinguish the inquiry demand and the operation demand of a patient by constructing an intention entity library containing inquiry and action deviation scores and combining threshold judgment and intention possibility scores; when the intention is unknown, a fine tuning language large model is introduced to generate a guide text, multi-round clarification is realized, and the intention recognition accuracy is ensured; after an inquiry intention is clarified, matching with a question and answer library to output answers, and after an action intention is clarified, mapping to an action instruction library and calling a back-end business system, so that closed-loop execution from patient languages to business actions is really realized. According to the method, a traditional intention-answer one-way mode is broken through, the blank of action intention recognition and execution is filled, a semantic gap in doctor-patient interaction is eliminated, a patient can directly complete operations such as registration, payment and sign-in through voice, and the medical treatment efficiency and the intelligent service level are remarkably improved.
Owner:BEIJING R&W ELECTRONICS TECH

A method for generating prompt words for large-scale models in power communication network operation and maintenance.

This application relates to the fields of computer data processing and artificial intelligence technology, specifically disclosing a method for generating large-scale model prompt words for the operation and maintenance of power communication networks. First, it standardizes, cleans, and aligns the original multi-source heterogeneous sensor data according to their time series, mapping it onto a network topology map to anchor the fault source. Then, it uses an anisotropic influence propagation model incorporating business protection logic to perform topology walks, extracting a context-aware subgraph containing fault propagation paths and primary / backup relationships. Next, the structural features of this subgraph are transformed into a serialized text description, combined with alarm facts and thought chain guidance instructions, to assemble a structured prompt word input large-scale model with rigorous logical constraints. This method solves the problem of attribution illusion in large-scale models caused by the lack of spatial constraints on prompt words, and the problem of missed detection of potential risk nodes due to neglecting business protection logic, achieving accurate localization and forward-looking prediction of complex faults in power communication networks.
Owner:STATE GRID HENAN INFORMATION & TELECOMM CO

Lightweight underwater semantic segmentation method based on fine-grained attention and adaptive feature fusion

The invention discloses a lightweight underwater semantic segmentation method based on fine-grained attention and adaptive feature fusion, and belongs to the field of computer vision and underwater image processing. Aiming at the problems of underwater image illumination attenuation, complex background and difficulty in small target segmentation in aquaculture, the invention provides a lightweight underwater semantic segmentation network and a matched training method, which mainly comprise three core improvements: 1, designing a fine-grained multistage feature attention module; superficial layer feature extraction is enhanced through cascade cavity convolution and space and channel attention mechanisms; secondly, a self-adaptive feature fusion module is introduced, and space details are reserved and redundant features are suppressed in combination with wavelet down-sampling; thirdly, a Gaussian Dice loss function is provided, and the problem of unstable training caused by small target boundary offset is solved through Gaussian weighting. According to the method, the calculation overhead is remarkably reduced, and meanwhile, the segmentation precision in the complex underwater environment is effectively improved.
Owner:DALIAN OCEAN UNIV

RNA (Ribonucleic Acid) far homologous detection method and system based on deep learning

PendingCN121963855Aimprove performanceEffectively bridge the semantic gapBiostatisticsBiological modelsData setVariome
The invention discloses an RNA (Ribonucleic Acid) far homology detection method and system based on deep learning, and belongs to the field of bioinformatics, an RNA sequence data set is divided into a plurality of sets and then a plurality of training data sets containing far homology pairs are constructed, so that an RNA far homology detection model is trained on the basis of each training data set, a plurality of model variants are obtained, and the RNA far homology detection model is obtained. Performance evaluation is carried out based on the independent test set to obtain a model with optimal performance; in the application stage, an RNA sequence data set is converted into vector representation based on the model, a vector database is constructed, after a to-be-detected sequence is received, the to-be-detected sequence is converted into vector representation based on the same mode, then the vector database is retrieved, similar RNA sequences including far homologous sequences are obtained, and an end-to-end deep learning framework is achieved. And the RNA sequence can be directly mapped to the structural similarity without depending on the structural information of RNA, so that the semantic gap between the sequence and the structure is effectively bridged.
Owner:HUAZHONG UNIV OF SCI & TECH

A digital resource retrieval system and method based on multi-modal and AI agent

The application provides a kind of digital resource retrieval system and method based on multi-modal and AI intelligent agent, it is related to electric digital data processing technical field, its method includes obtaining multi-modal query data and pre-processing, feature extraction and cross-modal semantic alignment, generates query vector set;Identify user category and accordingly the query vector set is semantically enhanced, and generates enhanced query vector;The enhanced query vector is matched with knowledge graph to determine the target node, and the extended query vector is generated along the semantic diffusion of knowledge graph;Based on the agent collaborative mechanism, the extended query vector is sequentially executed, and the parallel retrieval, fusion and sorting and confidence correction are carried out, and the modified sorting result is obtained;According to the scene adaptation factor, the modified sorting result is reordered, and the digital resource retrieval result is output in combination with the user category, so that user perception, semantic expansion and multi-agent collaborative retrieval for multi-modal education scene can be realized, and the accuracy and personalized adaptation degree of digital resource retrieval are improved.
Owner:LANGLANG CULTURE TECHNOLOGY CO LTD

Hybrid identifier generation-based recommendation method and system based on local collaborative context

ActiveCN121980091Bimprove interpretabilityGive full play to semantic understanding skillsPersonalizationLinguistic model
The application belongs to the technical field of artificial intelligence, and specifically relates to a hybrid identifier generation type recommendation method and system based on local collaborative context, acquires a user historical interaction sequence, carries out multi-granularity clustering on the user to obtain a group of each user; a static identifier is constructed based on item content, a dynamic identifier is constructed in combination with user group information and local interaction sequence, and a hybrid identifier of each item is fused and generated; an instruction fine-tuning task is constructed, group information is embedded into an instruction prompt, a large language model is used to learn to generate a hybrid identifier of a next item from a historical sequence; and a recommendation result is generated based on the optimized large language model according to a historical sequence of a target user and group information of the target user. The application fuses local collaborative context through multi-granularity group estimation, combines static and dynamic identifiers to construct a hybrid item representation, effectively solves the problems of single user interest modeling and missing semantics of item representation, and improves the accuracy and personalized level of the generation type recommendation.
Owner:QINGDAO UNIV OF SCI & TECH

Multi-modal data fusion processing method and system based on deep learning

The invention relates to the technical field of intelligent medical image auxiliary diagnosis, and discloses a multi-modal data fusion processing method and system based on deep learning, and the method comprises the steps: obtaining multi-modal medical data, carrying out the preprocessing of the multi-modal medical data, carrying out the feature extraction, and recording the quality meta-information; constructing positive and negative sample pairs, mapping the positive and negative sample pairs to a unified semantic space, and performing parameter optimization and soft alignment processing; generating a multi-scale time code, calculating a time sequence association weight matrix, and carrying out time sequence weighted fusion; quality evaluation is carried out, a preliminary quality score is calculated, and misjudgment is corrected through a clinical rule base to obtain a self-adaptive fusion weight; judging modal integrity and determining a fusion strategy, calculating a cross-modal attention interaction weight and carrying out weighted fusion; predicting disease categories, calculating the contribution degree of each modal feature and generating a multi-modal visualization result; according to the method, the problems of cross-modal semantic alignment, sequential relation modeling and data quality difference are solved.
Owner:ZHEJIANG CHINESE MEDICAL UNIVERSITY

Long video key frame retrieval method and device based on knowledge graph

The application relates to the technical field of multi-modal intelligent video understanding, and provides a long video key frame retrieval method and device based on a knowledge graph. A long video knowledge graph construction pipeline is constructed through a frame-level subtitle generation, frame-level knowledge graph construction, similarity video segmentation, segment and abstract generation process, and structured modeling of long video semantic content is realized. Higher retrieval precision can be obtained while ensuring efficiency through the setting of a two-level retrieval mechanism. Through vector matching and multi-hop neighbor expansion on a unified knowledge graph, a node set strongly related to positioning and problems is located, thereby narrowing the semantic gap between natural language problems and structured graphs and improving the accuracy of key frame selection. Through the setting of an iterative retrieval mechanism, the key frame obtained through retrieval can be used as the basis for answering the problem text.
Owner:NAT UNIV OF DEFENSE TECH

Semantic understanding-based multi-modal information visualization content generation method

PendingCN121997935Asolve semantic expressionSolve the problem of user intent disconnectionSemantic analysisNeural learning methodsEngineeringProcessing
The invention relates to the technical field of semantic processing, and discloses a multi-modal information visualization content generation method and system based on semantic understanding. The method comprises the steps of obtaining a natural language instruction, structured data and context information; respectively carrying out semantic analysis, data annotation and context coding; generating a unified representation through a multi-modal semantic fusion encoder; constructing a visual semantic decision tree, and providing an interactive adjustment interface for a user to dynamically adjust semantic node parameters; and updating the semantic representation in response to the adjustment operation, and driving a rendering engine to generate final visual content. The system comprises a multi-modal input acquisition module, a semantic analysis module, a fusion coding module, a decision tree construction module, an interaction adjustment module, a content generation module and the like. According to the system, interpretability and real-time controllability of a semantic level are realized, and the consistency of a visualization result and a user intention is remarkably improved.
Owner:HENAN UNIVERSITY