Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2709 results about "Semantic information" patented technology

Semantic information(Noun) (Research) the part of a message that is stored in the semantic memory system and can be tested with traditional verbal methods.

Brain tumor multi-modal large model construction method and device, equipment and storage medium

The invention discloses a brain tumor multi-mode large model construction method, device and equipment and a storage medium, and is applied to the technical field of brain tumor imagines.The method comprises the steps that pixel-concept level alignment is conducted on a multi-mode MRI image and a pathological text; constructing a multi-modal feature fusion network for fusing image features and text features by adopting an attention mechanism of pathology perception and combining medical semantic information; training the multi-modal feature fusion network to generate an analysis report and a segmentation result; according to the technical scheme of multi-task cooperation, cross-modal pathological semantic accurate alignment, pathological knowledge graph injection and lightweight and continuous optimization parallelization, full-process coverage of brain tumor accurate segmentation, analysis report generation and prognosis prediction is achieved, the problems that a traditional model lacks pathological semantic support and is insufficient in clinical adaptability are solved, and the clinical adaptability of the traditional model is improved. And the deployment feasibility and the dynamic optimization capability are also considered.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Context compression method based on multi-round dialogue intention graph construction

The invention provides a context compression method based on multi-round dialogue intention graph construction, which comprises the following steps: S1, dialogue data acquisition and preprocessing: carrying out natural language processing on each dialogue unit; s2, intention atlas construction is achieved through node design and edge design, and each node comprises original text content and structured semantic information; s3, carrying out context compression and graph structure cutting, and only retaining sub-graphs forming a core semantic link; s4, dynamic context management and topic jump processing: in a multi-topic dialogue, when a user jumps or switches to a new topic, a system records a sub-graph of a current active topic, and contextual nodes of an inactive topic are frozen; s5, context sequence generation and model input: linearizing node contents in the cut sub-graph according to a dependent link sequence to generate a compressed context sequence, and transmitting the context sequence and current user input to a large language model for reasoning; and S6, continuous updating and feedback optimization are carried out.
Owner:WUXI BAISHANG ZHONGWANG DATA TECHNOLOGY CO LTD

Hyperspectral image and laser radar data classification method based on dynamic fusion network

The invention relates to the technical field of artificial intelligence and remote sensing image processing, and particularly provides a hyperspectral image and laser radar data classification method based on a dynamic fusion network. The method comprises the following steps: preprocessing acquired multi-modal data, and constructing multi-scale input; a dual-scale local attention module is designed, and context information of different scales is fused in a self-adaptive weighted mode through gating soft pooling; a dynamic down-sampling feature enhancement module is designed, the down-sampling rate is dynamically adjusted according to the complexity of the feature map, and deep multi-scale interaction is carried out based on a Mama backbone; constructing a directional interactive attention module, extracting features in horizontal, vertical and diagonal directions through directional gating convolution, and capturing an anisotropic structure of a linear ground feature; through the design of a double-path classifier, fusing shallow space details and deep semantic information; and the model is trained, optimized and reasoned to obtain data classification, and the method improves the classification precision and the calculation efficiency.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Cross-modal document information extraction method based on space-semantic alignment

The invention relates to a cross-modal document information extraction method based on space-semantic alignment, and belongs to the field of artificial intelligence, computer vision and natural language processing. According to the method, the spatial feature and semantic information bidirectional alignment model is designed, by constructing the spatial feature and semantic feature bidirectional alignment model, the document layout information can dynamically adjust attention distribution of text semantic features, meanwhile, semantic information reversely optimizes the spatial features, collaborative modeling of spatial layout and semantic information is achieved, and the document layout efficiency is improved. Therefore, the accuracy and robustness of complex document information extraction are improved. According to the method, a hierarchical cross-modal information extraction model is designed, through the hierarchical cross-modal information extraction model, the overall structure of a document is recognized on the global level, local key content is focused on the regional level, fine modeling is conducted on fine-grained texts and visual elements on the entity level, and accurate recognition of a cross-modal entity and the semantic relation of the cross-modal entity is achieved; and the generalization ability and applicability of information extraction are enhanced.
Owner:BEIJING INST OF COMP TECH & APPL

Android system user interface interaction method and device based on large language model

The invention discloses an Android system user interface interaction method and device based on a large language model, and relates to the technical field of intelligent man-machine interaction, and the method comprises the steps: receiving a natural language instruction of a user, and synchronously capturing a real-time interface state of a current Android screen to generate a structured context containing semantic information of interactive elements; inputting the instruction and the context into a large language model, and generating a hierarchical task plan formed by a plurality of atomic operations with a logic dependency relationship; based on a dynamic adaptation mechanism, converting the atomic operation sequence into an operation instruction which can be executed by the Android system; receiving interface change information after instruction execution, and feeding back the interface change information as a new round of context to the large language model to determine a task execution state; and finally generating natural language feedback information to the user according to the execution state. By means of the mode, complex cross-application tasks can be understood and executed, and interaction intelligence of the Android system is remarkably improved through real-time interface perception and dynamic planning adjustment.
Owner:SHENZHEN Y-COM TECH CO LTD

Camouflage target detection method based on feature selection attention and frequency domain edge guidance

The invention discloses a camouflage target detection method based on feature selection attention and frequency domain edge guidance. According to the method, four-level features of a camouflage target image are extracted through a backbone network SMT and are respectively screened; the high-level features are input into a semantic information supplement module, and after semantic features are enhanced, the high-level features and the trunk features are sent into a spatial feature enhancement module together. And inputting the obtained fine-grained features into an edge feature sensing module, and finally fusing multi-scale features through a multi-scale jump connection technology to generate a mask pattern with higher discrimination. The method has the advantages that the network parameter quantity is reduced and key information is reserved through a feature selection mechanism; a spatial feature enhancement module is used for enhancing multi-scale feature representation and remote dependence modeling; the dilution of the semantic context is relieved by means of a semantic supplement module so as to improve the positioning precision; and an edge feature enhancement module is adopted to enhance edge semantic perception and improve boundary integrity. According to the method, the camouflage target detection performance is remarkably improved with relatively low calculation cost.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method and apparatus for processing image, electronic device, and storage medium

The disclosure provides a method and an apparatus for processing an image, an electronic device, and a storage medium, which relates to the field of artificial intelligence technologies, and particularly to a technical field such as computer vision, deep learning, and large-scale models. The solution includes: obtaining an input content adapted to an image processing task, in which the input content includes at least one of: a first text token sequence, a first image token sequence, or an image-text fusion sequence; obtaining a joint feature representation including multimodal semantic information by performing cross-modal semantic modeling on the input content, in which the multimodal semantic information indicates a semantic correlation relationship of the input content in different modalities; and generating an output content adapted to the image processing task based on the joint feature representation.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Three-dimensional graphic engine natural language interaction method and system based on MCP and AI Agent

The invention provides a three-dimensional graphic engine natural language interaction method and system based on an MCP and an AI Agent, and relates to the technical field of intelligent control, and the method comprises the steps: autonomously judging an MCP tool set needing to be called according to the analyzed intention and scene demands, and planning the precedence logic and cooperation mode of tool calling; converting the structured semantic information and the tool calling plan into a unified semantic data packet through an MCP protocol to form a cross-tool collaborative MCP instruction; based on an MCP instruction, the adapter module performs adaptation conversion according to API characteristics of a target three-dimensional engine, a corresponding interface is called to execute object operation in a three-dimensional scene, the AI Agent can automatically correct a subsequent instruction or tool calling strategy based on feedback, and closed-loop intelligent interaction is achieved. According to the invention, the interaction barrier between the natural language and the three-dimensional graphic engine is broken, so that the user can conveniently and intelligently control the three-dimensional object, the scene view angle and the dynamic effect directly through the natural language, and the interaction efficiency and flexibility are remarkably improved.
Owner:SHANGHAI URBAN CONSTR INFORMATION TECH CO LTD

Music stave sentiment classification method and system based on multi-level distillation

PendingCN121502446ASpeech analysisBiological modelsInformation processingApplying knowledge
The invention discloses a music stave sentiment classification method and system based on multi-level distillation, and belongs to the technical field of music information processing. The method comprises the steps of firstly collecting music stave data and converting the data into stave data vectors, then performing feature extraction by using a long short-term memory network, then constructing a teacher network and a student network for knowledge distillation, and realizing multi-level knowledge transmission through temperature scaling, KL divergence loss and mask feature distillation. And finally, training a lightweight classification model to complete sentiment classification. The knowledge distillation technology is creatively applied to staff sentiment classification, the classification accuracy is effectively improved through an online multi-level distillation mode, and the technical problems that a traditional method lacks semantic information and a self-supervised model is not suitable for sentiment tasks are solved. The method has the main advantages of high classification precision, light model weight, capability of effectively capturing music emotion features and the like.
Owner:NANCHANG HANGKONG UNIV COLLEGE OF SCI & TECH

Tunnel apparent disease detection method and system based on deep learning and knowledge distillation

The invention relates to the technical field of tunnel crack detection and artificial intelligence edge calculation, and provides a tunnel apparent disease detection method based on deep learning and knowledge distillation, which comprises the following steps: step 1, introducing spectral domain information enhancement to an original tunnel image, the edge texture features of the disease area in the image are enhanced through methods such as multi-scale wavelet transform and small-scale enhancement. Step 2, constructing a high-performance teacher model, introducing a flexible up-sampling structure to adapt to feature recovery requirements of different levels of semantic information, introducing an efficient visual coding module to enhance feature fusion capability of different scale channels, and designing a scale adaptive weighted loss function at the same time; by introducing a frequency spectrum enhancement mechanism, structural features of disease areas with low contrast, fuzzy edges and the like are remarkably enhanced in an image preprocessing stage, clearer information input is provided for a model, and the stable recognition capability of a system in environments of uneven illumination, complex background and the like is enhanced.
Owner:INST OF GEOLOGY CHINA EARTHQUAKE ADMINISTRATION

Multi-agent collaborative visual completion method and system based on semantic communication

The invention provides a multi-agent collaborative visual completion method and system based on semantic communication, and the method comprises the steps: a first agent collects an occluded overall original image through a visual sensor, generates an occluded region mask through a lightweight semantic segmentation network, extracts mask semantic information, and transmits the mask semantic information to a second agent; the second intelligent agent matches the global semantic feature of the second intelligent agent with the received mask semantic information, locates a sheltered area, then cuts the sheltered area to obtain a semantic feature vector of the area, and returns the semantic feature vector to the first intelligent agent; and after the first intelligent agent receives the image, performing high-fidelity reconstruction on the occlusion area by using a conditional diffusion model to obtain a complemented image with consistent semantics, and performing weighted fusion and local color matching and splicing on the complemented image and the original image to obtain a global perception result. According to the scheme, perceptual information sharing and occlusion area reconstruction with low bit rate, high robustness and high real-time performance can be realized under limited communication resources.
Owner:CRSC INST OF SMART CITY RES &DESIGN

Land unmanned equipment semantic preserving large model deployment method and system

The invention provides a land unmanned equipment semantic preserving large model deployment method and system, and relates to the technical field of navigation control. According to the deployment method, a preliminary control suggestion is directly output through a lightweight fusion model, optimization fusion is carried out through an MPC framework, underlying dynamics and environmental constraints, and seamless connection from semantics to control is achieved; in combination with a VILO odometer and instance segmentation, a global map containing dynamic obstacle semantic information is constructed and updated in real time, and an accurate context is provided for planning and decision making; model pruning, knowledge distillation and edge calculation scheduling are adopted, so that a complex multi-modal semantic model can run in real time on an embedded platform; the decision basis from instruction analysis to control execution is recorded and visualized in the whole process, and the transparency and credibility of the system are improved; on-line re-planning, multi-stage fault detection and switching strategies are integrated, and the robustness of the system in a dynamic environment and an abnormal condition is remarkably improved.
Owner:BEIHANG UNIV

Multi-scale feature and local detail enhancement fused low-illumination target detection method and system

The invention provides a low-illumination target detection method and system fusing multi-scale features and local detail enhancement. The system comprises a feature extraction network based on a multi-pooling pyramid and cross-stage double-mixed attention, a dynamic detail semantic fusion pyramid network and double groups of detection heads. The feature extraction network based on the multi-pooling pyramid and the cross-stage double-mixed attention comprises a convolutional layer, a C2PSA module, an MPSPPF module and a CSP-EDHAN module; the dynamic detail semantic fusion pyramid network is used for fusing shallow high-frequency details and deep semantic information through top-down and bottom-up multi-scale feature fusion and introducing a surface detail fusion module into multiple scales, so as to output a fused feature map; the double groups of detection heads adopt a decoupling detection branch design, and the position and the category of a target are directly predicted on a fused feature map. The method can remarkably improve the precision and robustness of target detection in a low-illumination environment, and is suitable for the fields of night monitoring, automatic driving, security and protection and the like.
Owner:FUZHOU UNIV

Safe real-time detection method in complex scene based on multi-scale feature fusion

The invention discloses a safety real-time detection method in a complex scene based on multi-scale feature fusion, and relates to the technical field of safety detection, and the method comprises the steps: S1, obtaining a to-be-detected complex scene image; s2, extracting multi-scale initial feature maps with different semantic information and spatial details; s3, inputting the initial feature maps of the plurality of scales into an adaptive feature fusion network; s4, inputting the enhanced feature pyramid into a lightweight decoupling detection head, and executing target classification and bounding box regression in parallel; and S5, based on the category and position information, generating and outputting a final security detection result. The method has the advantages that semantic and detail features of different levels are effectively integrated through a self-adaptive gating fusion mechanism and multi-scale context aggregation, and the detection precision and robustness of the model on a multi-scale target in a complex scene are remarkably improved.
Owner:GUANGZHOU RENHE SHICHUANG INFORMATION TECHNOLOGY CO LTD

Dynamic modeling method of geological structure three-dimensional model

The invention relates to the technical field of three-dimensional modeling, in particular to a dynamic modeling method for a geological structure three-dimensional model. The method comprises the steps that multi-source data are acquired and preprocessed, a voxel semantic fusion algorithm based on variational optimization is introduced, after semantic information is extracted from the preprocessed multi-source data, the preprocessed multi-source data are mapped to a target three-dimensional space grid, and an optimal semantic fusion vector is obtained; based on the optimal semantic fusion vector, generating a standard voxel data pool, and constructing a geological structure three-dimensional model; monitoring data change, calculating the position of a newly added data point, combining the standard voxel data pool to obtain a space updating area, and modeling the space updating area to realize model updating; and after the modeling of the space updating region is completed, optimizing the boundary continuity. The problems that multi-source geological data cannot be directly used for structure construction and semantic fusion of a three-dimensional model, dynamic response to newly-added data is lacked, and geometric discontinuity and structural logic discontinuity exist at the boundary are solved.
Owner:INNER MONGOLIA SHANJIN GEOLOGY & MINERAL EXPLORATION CO LTD

Visual language navigation method and system based on cross-task incremental semantic memory graph

The invention belongs to the field of artificial intelligence and robot navigation, and discloses a visual language navigation method and system based on a cross-task incremental semantic memory graph, and the method comprises the steps that an intelligent agent executes a zero-sample visual language navigation task in a continuous environment; performing cross-modal alignment on the natural language instruction and environment observation based on a multi-modal large language model, selecting candidate waypoints and updating task progress; a semantic memory graph is constructed and dynamically updated, wherein the semantic memory graph is used for structured storage and cross-task multiplexing of scene semantic information and a spatial topological relation sensed by an intelligent agent in historical tasks; and performing global path planning and local dynamic fine tuning based on the semantic memory graph. According to the method, the problem of task-by-task forgetting in a traditional method is solved, environment understanding and task reasoning capabilities are improved by constructing structured long-term memory, and navigation precision and robustness are optimized through a global-local collaborative strategy.
Owner:SHANDONG UNIV

Small sample remote sensing image classification method based on hierarchical spatial structure learning

The invention discloses a small sample remote sensing image classification method based on hierarchical spatial structure learning. The method comprises the following steps: firstly, extracting multi-scale features of a remote sensing image by using a ViT (Visual Transform) model, and capturing rich semantic information and spatial structure relationships in the image; secondly, constructing a graph structure based on spatial adjacency and attention weight to model a structured relationship between samples, and encoding graph node features through a graph convolutional network (GCN) so as to enhance the discrimination ability of the features in a structural semantic space; thirdly, a residual enhancement mechanism is introduced to fuse global semantic information, and the discrimination capability of graph embedding is improved; then, based on the structural similarity between the support set and the query set, performing classification decision, and realizing accurate classification under a small sample condition; and finally, carrying out joint optimization on the whole model by adopting a training strategy of a small sample meta learning task and a supervision loss function.
Owner:BEIJING INST OF TECH

Weak supervision scene understanding method, system and equipment for multi-modal information interaction

The invention discloses a weak supervision scene understanding method, system and equipment for multi-modal information interaction. The method comprises the following steps: extracting initial visual features and initial text features according to an input image and a finger expression; mapping the initial visual features and the initial text features to a shared semantic space, and performing mutual perception and alignment on the initial text features and the initial visual features to obtain text perception visual features and visual perception text features; enhancing the text perception visual features through an encoder; constructing a query text feature of the target object, and interacting with the query visual feature of the input image to generate an initial object query; inputting the initial object query and the enhanced text perception visual features into a decoder, and outputting to obtain a target frame on the input image; total loss is calculated, and iterative training is repeated. According to the method, rich semantic information in the finger expression is fully mined and utilized, and precise positioning of the target object in the image through semantic analysis described by the natural language is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Face image reconstruction method based on semantic identity feature decoupling and consistency retention of diffusion model

The invention discloses a face image reconstruction method based on semantic identity feature decoupling and consistency reservation of a diffusion model, and the method comprises the steps: 1, obtaining and preprocessing a face image set of identity labeling, and generating a face feature point distribution diagram, a semantic mask diagram and a description text; 2, multi-modal features are extracted and fused through a semantic identity extraction network; 3, carrying out noise adding and de-noising processing by utilizing a diffusion model, and combining a reconstructed network and semantic identity loss optimization; and 4, face image reconstruction is completed. According to the method, in the face image reconstruction process, the driving requirements of semantic information such as texts for image editing can be accurately captured, fine-grained semantic features and identity features are decoupled, the core identity features of the face can be effectively reserved, and loss of identity consistency caused by semantic editing is avoided; therefore, technical support is provided for application scenes with high requirements on face identity accuracy in the field of computer vision, and the reliability and practicability of face image reconstruction are improved.
Owner:ANHUI UNIV

Remote sensing image three-dimensional reconstruction method based on semantic information

The invention discloses a remote sensing image three-dimensional reconstruction method based on semantic information, and relates to the technical field of remote sensing image three-dimensional modeling, and the method comprises the steps: carrying out the processing of a remote sensing image covering a target region, obtaining point cloud data, and carrying out the semantic segmentation based on a semantic segmentation model, and determining a semantic category label; dividing voxel units based on a three-dimensional sparse point cloud distribution condition, and constructing to obtain an initial anchor point; the distribution density of the initial anchor points is adjusted in combination with the semantic category labels, and the adjusted initial anchor points are obtained; iteratively training the three-dimensional Gaussian splash model for multiple times to update the anchor points to obtain scene anchor points; and obtaining a three-dimensional reconstruction model of the target area based on scene anchor point rendering. According to the method, semantic information is introduced and a 3DGS three-dimensional modeling technology is fused, so that the densification quality of anchor points and the geometric boundary definition of the model are effectively improved, the modeling efficiency can be improved, and the structural rationality and semantic interpretation of a three-dimensional reconstruction result can be enhanced.
Owner:WUHAN UNIV

File processing method and device oriented to big language model retrieval enhancement generation

The embodiment of the invention provides a large language model retrieval enhancement generation-oriented file processing method and device, and the method comprises the steps: carrying out the information extraction and semantic information enhancement of a file according to different file types, and constructing a meta-information structure of the file in combination with an enterprise business scene; dividing the document content into a plurality of structured blocks based on the meta-information structure, and labeling the title and context information of each block; when the file is uploaded, according to the parent directory material quantity and content change of the directory where the file is located, performing dirty marking on the directory; when a data request is received, whether a dirty mark exists in a related directory or not is judged, if yes, the directory is requested to be locked, a summary is extracted from files extracted from bottom to top through a large language model, and the dirty mark is cleared after the summary is cached.
Owner:特赞(上海)信息科技有限公司

Digital transformation intelligent question-answering method, device and equipment for small and medium-sized enterprises and storage medium

The invention provides a small and medium-sized enterprise digital transformation intelligent question answering method and device, equipment and a storage medium, and the method comprises the steps: carrying out the semantic recognition and entity extraction of a question input by a user, and generating a question semantic vector; retrieving a corresponding target text fragment in a preset vector database according to the question semantic vector; searching a sub-graph structure associated with the question semantic vector in a preset knowledge graph according to the target entity of the question semantic vector; reordering the text segments and the sub-graph structures based on correlation scores to obtain an ordering result set; and inputting the sorting result set into a retrieval enhancement generation model, and generating question and answer output content. Through the implementation of the scheme of the application, association matching of question semantics and enterprise knowledge can be realized by utilizing the text semantic information and the knowledge graph structure information at the same time in the question and answer generation process, the consistency of question and answer contents in logic structure and semantics is ensured, and the accuracy of digital transformation question and answer of small and medium-sized enterprises is improved.
Owner:YUNDI SMART TECH CO LTD

Distributed computing modeling system design method for high-dimensional discrete data

The invention relates to the technical field of distributed computing, and discloses a distributed computing modeling system design method for high-dimensional discrete data, which comprises the following steps: acquiring a feature identifier access request in a high-dimensional discrete data stream, counting an access frequency, and marking a feature identifier of high-frequency access as a high-frequency access state; constructing a local parameter copy on a non-master node physical computing unit of the distributed cluster; routing gradient update data aiming at the feature identifier to a local parameter copy closest to the network and generating a local residual tensor; calculating the direction cosine similarity of the local residual tensor and the global gradient update vector; calculating a real-time access entropy based on the access source distribution and converting the real-time access entropy into a dynamic synchronous threshold value; according to the method, through a dynamic synchronization judgment mechanism based on semantic information gain, high-frequency but invalid redundant communication is effectively inhibited, and network congestion under long-tail distribution is reduced.
Owner:NINGBO DAHONGYING UNIV

Adaptive frequency domain adversarial training method and device for target detector

The invention belongs to the field of computer vision and artificial intelligence security, and discloses a self-adaptive frequency domain adversarial training method and device for a target detector, the target detector comprises a repair module and a pedestrian detector which are connected in series, and the input of the repair module is connected with the output of a patch detector; the self-adaptive frequency domain adversarial training method comprises the following steps: losses in joint training comprise standard target detection losses, repair consistency losses on a frequency domain based on a frequency domain image corresponding to a training image and a clean image, and repair dependence losses based on a detected average precision mean value; according to the invention, the end-to-end joint training is carried out through the restoration module and the subsequent pedestrian detector, and the optimization target of the restoration module is directly aligned with the improvement of the detection robustness, so that the confrontation disturbance is eliminated as far as possible, and meanwhile, the key semantic information of the detection task is reserved to the maximum extent. The separation of the performance of the repair module and the pedestrian detector is avoided, and the detection robustness is improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Dynamic digital twinning three-dimensional Gaussian splash rendering system and method

The invention discloses a dynamic digital twinning three-dimensional Gaussian splash rendering system and a dynamic digital twinning three-dimensional Gaussian splash rendering method. According to the system, a static scene is reconstructed by utilizing a three-dimensional Gaussian splashing technology through a semantic scene reconstruction module, semantic information is associated to Gaussian primitives in combination with spatial alignment and semantic injection, and a model with a semantic identifier is generated; multi-source dynamic data are fused and converted into a dynamic data field which can be accessed by a GPU through a multi-mode spatio-temporal data processing module; and through a rendering attribute dynamic modulation module, screening a target primitive according to the semantic identifier, and directly modulating internal rendering attributes such as color, transparency or shape of the target primitive based on a data field query value, thereby realizing internal rendering and visualization of dynamic data. According to the method, the problems that in the prior art, rendering reality and real-time performance are difficult to give consideration to and data and scene fusion is superficial are solved, and dynamic digital twinning rendering with high reality and strong immersion is realized.
Owner:XIAMEN UNIV ARCHITECTURAL DESIGN & RES INST CO LTD

Large model zero sample learning method for hierarchical semantic enhancement

The invention discloses a hierarchical semantic enhanced large model zero sample learning method, which is characterized by comprising the following steps: firstly, performing semantic enhancement on category names by using a large language model to generate rich text description, and constructing a dynamic and hierarchical semantic prototype by combining original semantic information, the prototype comprising global concepts, local attributes and relation representations; secondly, extracting global semantic features and local detail features of the image by adopting a vision-language large model and convolutional neural network double-branch structure, and fusing the global semantic features and the local detail features through an attention mechanism to obtain enhanced visual representation; finally, multi-level global alignment, local alignment and relation alignment are designed and jointly optimized, accurate mapping of enhanced visual features and hierarchical semantic prototypes is achieved on multiple granularities, and classification of invisible categories is finally completed. The objective of the invention is to solve the problem of limited model generalization ability caused by insufficient semantic representation and single vision-semantic granularity in the existing zero sample learning method, and to enhance semantic representation by introducing knowledge of a large language model and innovatively implement hierarchical alignment, so that the robustness of the zero sample learning method is improved. And the recognition precision and robustness of the model in traditional and generalized zero sample learning scenes are remarkably improved.
Owner:XIANGTAN UNIV

Recommendation method based on semantic enhancement and heterogeneous hypergraph network

The invention discloses a recommendation method based on semantic enhancement and a heterogeneous hypergraph network. The recommendation method comprises the following steps that semantic information in an explicit feedback text is coded and serves as an auxiliary signal of a recommendation task; classifying the articles into predefined categories by using LLM, constructing article-category association, and mining a potential co-occurrence relationship of the articles; constructing a heterogeneous hypergraph network; spreading and aggregating hypergraph information; performing semantic alignment and model training; and performing recommendation calculation based on the final representation of the user and the representation of the article, and outputting a recommendation result. According to the method, through technical paths of semantic coding, hypergraph modeling, information spreading and alignment supervision, comment semantics of LLM coding are aligned to the recommendation space through GAE, and the problem of degradation of LLM representation in the recommendation space is effectively solved.
Owner:HUAZHONG UNIV OF SCI & TECH

Historical image-text-based three-dimensional ancient city model construction method and device, electronic equipment and storage medium

The invention relates to a historical image-text-based three-dimensional ancient city model construction method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining historical image-text data of a target city, the historical image-text data comprises related historical literatures of the target city in a specified period, an electronic ancient map corresponding to an ancient map drawn in the specified period, an electronic old map corresponding to an old map surveyed and mapped in the latest period from the specified period, and a shot satellite image; based on the historical image-text data, generating a city plane restored map of the target city in a specified period; extracting semantic information of the target city in a specified period from at least one kind of data in the historical image-text data, wherein the semantic information comprises city information of the target city in the specified period and attribute information of constituent elements; and generating a three-dimensional ancient city model of the target city in a specified period based on the semantic information city plane restoration map. Therefore, a high-precision and high-integrity three-dimensional ancient city model of the target city in the specified period can be constructed.
Owner:TSINGHUA UNIVERSITY

Intelligent charging pile layout optimization method based on multi-scale space-time diagram neural network

The invention relates to the technical field of electric vehicle charging pile planning, in particular to an intelligent charging pile layout optimization method based on a multi-scale space-time diagram neural network. Comprising the following steps: S1, constructing a heterogeneous dynamic graph which represents a potential charging pile position node set, epsilon t represents a time-varying edge set, At represents a time-varying adjacent matrix, and Xt represents a node feature matrix; s2, calculating a self-adaptive adjacency matrix with a specific relationship, and calculating a multi-type dynamic relationship between capture nodes of the self-adaptive adjacency matrix based on the constructed heterogeneous dynamic graph model; s3, updating a structure bias matrix, and capturing dynamic change characteristics of the network; s4, calculating distance measurement between nodes, and fusing geographic information and semantic information; s5, through graph neural network learning, based on the constructed heterogeneous dynamic graph and the calculated adaptive adjacency matrix, executing a graph neural network learning process, and extracting node space-time representation; and S6, predicting a future charging demand based on node representation obtained by graph neural network learning, and generating based on the future charging demand.
Owner:GUIZHOU AUTO FEDERATION NETWORK TECH CO LTD

Multi-modal semantic-action alignment method and device for end-to-end automatic driving

The invention provides a multi-modal semantic-action alignment method and equipment for end-to-end automatic driving, and belongs to the technical field of automatic driving, and the method comprises the steps: carrying out semantic reasoning through a large language model based on various sensor data and language instruction information of a vehicle, and generating semantic information containing a driving intention; inputting the semantic information into a semantic-action alignment module, and converting the semantic information into corresponding driving action representation through a learned consistency mapping relation from a semantic space to an action space; and generating an executable control track of the vehicle according to the driving action representation. According to the method, the problem of insufficient semantic and action space alignment is solved, information distortion and precision limitation caused by post-processing depending on rules are avoided, the high-level driving intention can be generated based on complex multi-mode information (such as navigation instructions and traffic environments), it is ensured that the final execution action is highly consistent with the intention, and the driving intention is more accurate. And the decision-making rationality of the system in a complex scene is enhanced.
Owner:DONGFENG MOTOR GRP