Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3965 results about "Semantic feature" patented technology

Semantic features represent the basic conceptual components of meaning for any lexical item. An individual semantic feature constitutes one component of a word's intension, which is the inherent sense or concept evoked. Linguistic meaning of a word is proposed to arise from contrasts and significant differences with other words. Semantic features enable linguistics to explain how words that share certain features may be members of the same semantic domain. Correspondingly, the contrast in meanings of words is explained by diverging semantic features. For example, father and son share the common components of "human", "kinship", "male" and are thus part of a semantic domain of male family relations. They differ in terms of "generation" and "adulthood", which is what gives each its individual meaning.

Intelligent question-answering system optimization method and device based on knowledge graph

The invention relates to an intelligent question-answering system optimization method and device based on a knowledge graph, and the method comprises the steps: obtaining original knowledge data of a target knowledge domain, and constructing a knowledge graph structure model; extracting term information of entity nodes in the knowledge graph structure model, and constructing an entity term set; receiving a natural language question input by a user, executing a semantic understanding operation based on the standardized expression set to obtain a structured question semantic representation, and matching the question semantic representation with the case training set to obtain context semantic features; constructing a cue word template, and executing a query instruction generation operation to obtain a target query statement of the graph database; submitting the target query statement to a graph database to execute data retrieval operation, and obtaining query result data corresponding to the question semantic representation; and performing personalized rendering processing on the query result data based on the user portrait information to generate final question and answer return content. The method has the effect of improving the query accuracy.
Owner:PENGHUA FUND MANAGEMENT CO LTD

Vision generation method and device based on semantic association modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health, poster design and the like, and discloses a visual sense generation method and device based on semantic association modeling, equipment and a medium. Generating a demand text containing theme and style parameters; semantic features in the demand text are extracted, semantic association weights are constructed, and element layout coordinates are optimized in combination with spatial distribution constraints; and encoding the layout information into a control matrix, fusing the control matrix with the initial noise, adjusting a noise reduction process through an encoding and decoding network, and generating target visual content highly matched with the semantic meaning of the user instruction. According to the method, the layout optimization function is constructed, the diffusion model is guided to focus the semantic salient region in space, language model output and the visual generation process are closely combined, structured response and space mapping of user semantic requirements are achieved, and the expression consistency and personalized adaptation capacity of visual content generation are improved.
Owner:SHENZHEN PINGAN COMM TECH CO LTD

AIGC content generation method and system based on multi-modal fusion

The invention relates to the technical field of AIGC content generation, discloses an AIGC content generation method and system based on multi-modal fusion, and aims to solve the problems of decentralization, low efficiency and insufficient originality of a traditional content generation tool. Multi-modal data such as texts, images, videos and audios are integrated, user intentions are analyzed in combination with intelligent retrieval and a domain knowledge base, automatic generation from multi-modal input to high-quality creative content is achieved, a cross-modal collaborative generation technology is adopted, semantic features are dynamically aligned, and logically coherent content is generated. The content emotional value is enhanced through an emotional analysis and dynamic optimization strategy, the homogenization bottleneck is broken through, meanwhile, an automatic quality evaluation and format adaptation mechanism is integrated, deep application of scenes such as text travel, advertisement, e-commerce and interactive network television service is supported, marketing copywriting, short videos and cross-platform distribution schemes can be efficiently generated, and the market competitiveness is improved. And the content production efficiency and the creativity transmission are obviously improved.
Owner:HANGZHOU WANDIAN TECHNOLOGY CO LTD

Multi-modal content compliance auditing method and system

The invention provides a compliance auditing method and system for multi-modal content. The method comprises the following steps: performing feature extraction on unstructured to-be-audited multi-modal content to obtain a structured feature vector; performing image-text semantic association on the text semantic feature vector and the image visual feature vector to obtain a fusion feature vector involving image-text semantic contradiction; constructing a domain knowledge graph based on the compliance guidance data of the domain to which the to-be-audited multi-modal content belongs; inputting the fusion feature vector into a domain knowledge graph, and performing compliance rule retrieval by adopting a sub-graph matching algorithm to determine a violation type corresponding to the fusion feature vector and a violated compliance term; and generating an interactive compliance audit report. The system comprises functional modules for realizing the steps in a one-to-one correspondence manner. According to the technical scheme, the problem that cross-modal semantic analysis of an existing multi-modal content compliance auditing method is not accurate can be solved.
Owner:SHANGHAI CAIYUE XINGCHEN INTELLIGENT TECHNOLOGY CO LTD

Marketing video auditing method based on AI

The invention provides an AI-based marketing video auditing method, and relates to the technical field of AI marketing video auditing, and the method comprises the steps: obtaining a multi-modal data original structure set, and extracting image semantic features, voice expression features, text semantic features and scene label information, and obtaining an image semantic feature set, a visual rhythm feature set, a voice expression feature set, a voice and picture synchronous association vector structure, a text semantic feature set and a subtitle semantic and image main body linkage relation graph. By constructing an image semantic feature set, a voice expression feature set, a text semantic feature set and a visual rhythm feature set and fusing the image semantic feature set, the voice expression feature set, the text semantic feature set and the visual rhythm feature set into a multi-modal content fusion feature tensor, unified modeling of an AI marketing video at visual, auditory and semantic levels can be realized; and subsequent microscopic consistency detection, compliance knowledge graph and emotion semantic conflict identification are effectively performed, so that full-link risk perception and accurate auditing of video contents are realized.
Owner:SHANGHAI WANGMAI INFORMATION TECH GRP CO LTD

Enhanced generation method based on question matching retrieval

The invention provides an enhanced generation method based on question matching retrieval, and belongs to the field of matching generation, and the method comprises the following steps: S1, a semantic feature coding stage: carrying out real-time feature extraction and vector space mapping on a natural language query input by a user by adopting a deep neural network model, generating high-dimensional distributed representation with semantic representation capability; s2, a knowledge base intelligent retrieval stage: executing multi-dimensional semantic matching in the vectorized knowledge base based on an approximate nearest neighbor search algorithm, and screening out a candidate knowledge set highly related to query semantics through a similarity measurement function; s3, retrieval matching results are automatically associated to the structured knowledge base through the established semantic-knowledge mapping relation, the preprocessed standardized response content is directly obtained, and the response content adopts a multi-modal data organization form and comprises a structured data entity and retains a rich text expression form.
Owner:北京致链科技有限责任公司

Large language model construction method fused with spatial semantic understanding

The invention relates to a large language model construction method and system fused with spatial semantic understanding, and the method comprises the steps: obtaining a multi-source heterogeneous corpus, and extracting an entity, an attribute and a business rule; extracting a semantic feature vector set based on the multi-source heterogeneous corpus, and constructing an entity relationship network and an enhanced knowledge graph; generating an enhanced training sample, and training the general large language model to obtain a primary large language model; generating a verification sample set and performing verification; identifying a specific weakness pattern, and generating a corresponding confrontation sample and a knowledge enhancement sample; training the primary large language model to obtain an optimized large language model; in conclusion, the enhanced knowledge graph fusing the spatial semantic features and the business rules is constructed, and the gradient training samples are generated based on the graph to perform multi-stage model training and optimization, so that the method has the effects of improving the internalized understanding ability of the model for the spatial semantics and the business rules and enhancing the reliability of multi-step spatial reasoning.
Owner:URBAN PLANNING & DESIGN INST OF SHENZHEN UPDIS

Multi-round dialogue intention recognition method and system based on adaptive semantic understanding

The invention provides a multi-round dialogue intention recognition method and system based on self-adaptive semantic understanding, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a natural language dialogue text of a current round of a user, and taking the natural language dialogue text as original input data; based on original input data, multi-level semantic features are extracted through a dynamic semantic coding algorithm, and semantic vector representation of a current round of dialogue is generated; setting three fixed anchor points in a semantic vector space based on a current round semantic vector and a historical dialogue state vector to form a triangular analysis structure; performing gridding segmentation on the triangular analysis structure, and generating a feature adjustment value according to distribution characteristics of segmented grids; and dynamically correcting the extraction process of the context-related features by using the feature adjustment value to obtain the corrected context-related features. According to the method, end-to-end optimization is realized in multiple rounds of interaction scenes such as customer service and intelligent assistants through full-process design.
Owner:MEGAVIEW INTELLIGENCE TECH LTD

Multi-modal fusion rumor detection method and system based on dynamic graph convolutional neural network

The invention discloses a multi-modal fusion rumor detection method and system based on a dynamic graph convolutional neural network. According to the method, a dynamic feature graph of a language propagation path is constructed, and potential features in the language propagation process are extracted and analyzed by utilizing time sequence changes and key node relations between nodes in a propagation graph. A neural network is adopted to extract and enhance image data, text semantic features are extracted in combination with a text feature modeling network, text feature vectorization expression is achieved based on a BERT model, and rich semantic information is obtained. And a gating mechanism is introduced to dynamically adjust fusion weights of different modal features, and an information fusion strategy is optimized. A collaborative attention mechanism is further adopted for deep fusion, interactive learning of text, image and propagation path features is enhanced, and the relevance of cross-modal and time series data is improved. And finally, inputting the fused feature vectors into a classifier for accurate classification, thereby realizing accurate detection of the social media rumors. According to the method, the multi-modal features are effectively integrated, and the false information identification efficiency is remarkably improved.
Owner:CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY +1

Text sentiment analysis method and system based on dynamic semantic segmentation and feature perception

The invention provides a text sentiment analysis method based on dynamic semantic segmentation and feature perception, and belongs to the field of natural language processing. Inputting the semantic vector sequence into a knowledge retrieval and dynamic graph construction model for multi-path context enhancement through the knowledge retrieval and dynamic graph construction model to obtain semantic features, semantic knowledge and graph structure information; the semantic vector sequence is subjected to multi-path context enhancement, the complex relation between different parts in the text can be comprehensively considered, and more comprehensive and deep semantic features and knowledge can be mined. Unified representation of semantic features is combined with heterogeneous graph features, an antagonism training strategy is adopted to train a knowledge retrieval and dynamic graph construction model, an improved ATOSS + module is introduced to carry out hierarchical attention fusion, and multi-granularity semantic enhancement features are obtained; therefore, emotion clues and semantic association hidden in the text can be captured, and the accuracy and integrity of semantic understanding of the text are improved.
Owner:SHAANXI UNIV OF SCI & TECH

Video content semantic understanding and text description generation method based on deep learning

The invention discloses a video content semantic understanding and text description generation method based on deep learning, and relates to the technical field of multimedia information processing.The method comprises the steps that the semantic similarity of a text and a video frame is calculated through a CLIP model, related key frames are selected, and features are aggregated; respectively extracting audio, visual and semantic features; aligning different modal features by using self-attention, unifying dimensions of the LSTM, and then splicing and fusing; attention weights are calculated at a video level, a frame level and a channel level, and key information expression is enhanced; swin Transform encodes fusion features, and LSTM (Long Short Term Memory) decodes step by step to generate natural language description; and a text-video index database is constructed, and rapid retrieval is realized based on semantic similarity. According to the method, the mapping relation between the video features and the natural language is learned end to end through the deep learning model, dependence on a fixed template can be eliminated, and semantic description with various sentence patterns and coherent logic is generated.
Owner:CHINA UNIV OF MINING & TECH YINCHUAN COLLEGE

Multi-scale semantic guidance image compression method and system and storage medium

The invention discloses a multi-scale semantic guidance image compression method and system and a storage medium, and the method comprises the following steps: obtaining input image data, carrying out the preprocessing of an input image, and obtaining standardized image data; inputting the standardized image data into a pre-trained semantic segmentation network to generate a multi-scale semantic feature map and a semantic weight map corresponding to the multi-scale semantic feature map; a three-stage pyramid encoder is constructed, and the standardized image data is subjected to the following steps of: sampling under depth separable convolution to generate multi-scale features; the reversible neural network carries out nonlinear transformation on the multi-scale features; the multi-scale feature subjected to nonlinear transformation is decomposed into a low-frequency sub-band and a high-frequency sub-band through adaptive discrete wavelet transformation, dynamic selective state space modeling is executed on the high-frequency sub-band based on a semantic weight map, and a compressed code stream is generated; and inputting the compressed code stream into a decoder, decoding based on a lightweight Mama module, and reconstructing an image in combination with inverse wavelet transform and a semantic weight map.
Owner:XIANGJIANG LAB

Visual language navigation method for cross-modal alignment in dynamic shielding environment

The invention discloses a visual language navigation method for cross-modal alignment in a dynamic shielding environment, and the method comprises the steps: collecting multi-modal data through a visual sensor, an inertial measurement unit, a laser radar and the like, and carrying out the preprocessing and time synchronization; sensing the dynamic shielding object through a model composed of a convolutional neural network and a long-short-term memory network, and estimating the future change of the dynamic shielding object in combination with a space-time sequence prediction algorithm; a double-branch convolutional neural network and a Transform based on a dynamic attention mechanism are adopted to respectively extract visual and semantic features and fuse the visual and semantic features; on the basis of occlusion prediction, potential occlusion region features are extracted in advance from a time dimension, an occluded image is repaired by using a generative adversarial network and geometric constraints in a space dimension, and cross-modal feature alignment is optimized through an attention mechanism; planning a path by using a hybrid reinforcement learning algorithm based on a deep Q network-space and a fast exploration random tree, and dynamically adjusting according to real-time shielding; according to the method, the accuracy, adaptability and reliability of visual language navigation in a dynamic shielding environment are improved.
Owner:SHANGHAI JIAOTONG UNIV

Method and system for generating official document key abstract based on multi-modal feature extraction

The invention provides an official document key abstract generation method and system based on multi-modal feature extraction, and relates to the technical field of multi-modal artificial intelligence generation, and the method comprises the steps: firstly obtaining text modal data, image modal data and table modal data of a to-be-processed official document, then carrying out the hierarchical semantic analysis processing of the text modal data, and obtaining a to-be-processed official document key abstract; the method comprises the following steps: generating a text semantic feature set, performing visual element extraction processing on image modal data, generating an image feature set, performing structured analysis processing on table modal data, generating a table feature set, and performing cross-modal alignment processing on the text semantic feature set, the image feature set and the table feature set. According to the method, the dynamic association feature sets among the text semantics, the image elements and the table elements are determined according to the text semantics, the image elements and the table elements, the three feature sets are subjected to multi-modal fusion processing according to the feature sets, the target abstract content of the to-be-processed official document is generated, complementarity of multi-modal information in the official document is fully utilized, and the generated abstract is more complete, accurate and targeted.
Owner:STATE GRID SHANDONG ELECTRIC POWER COMPANY WEIFANG POWER SUPPLY

Multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion

The invention discloses a multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion. The method comprises the following steps: respectively extracting multi-scale features of two modal input images by adopting a double-branch encoder; sequentially executing frequency domain decoupling and fusion, mutual information constraint-based feature optimization and low-frequency guided cross-modal fusion processing on each scale feature to generate a fused semantic feature; and performing up-sampling and feature refining on the fused features through a decoder, and outputting a full-resolution segmentation prediction map. According to the multi-modal remote sensing image semantic segmentation method, modal sharing information and specific details are effectively separated through frequency domain decoupling, feature representation is optimized through mutual information constraint, adaptive feature fusion is achieved in combination with an attention mechanism, and the accuracy and robustness of multi-modal remote sensing image semantic segmentation are remarkably improved.
Owner:NORTHEAST FORESTRY UNIV

Cross-modal interaction image restoration method fusing text semantic guidance and visual structure prior

The invention discloses a cross-modal interactive image restoration method fusing text semantic guidance and visual structure priori, which comprises the following steps of: firstly, acquiring natural language description input by a user and an image to be restored, and generating a semantic segmentation map of the image through a semantic segmentation model; encoding the text and image semantics by using a pre-trained cross-modal encoding model to obtain text and semantic features; guiding a semantic alignment attention module through Prompt to realize deep fusion of multi-modal semantic features and image space features; structural enhancement and regulation of image features are realized by constructing a text guide weight graph, performing element-level modulation on the text guide weight graph and the optimized semantic segmentation graph, constructing a cross-modal structure semantic feature graph and generating a structural modulation factor; a four-stage image restoration network is adopted, and a high-quality restoration image conforming to semantic guidance and structure prior is generated step by step. According to the method, the semantic consistency, the structural integrity and the visual reality sense of an image restoration result are improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-modal fusion deep learning analysis method and system

The embodiment of the invention provides a multi-modal fusion deep learning analysis method and system. The method is applied to the technical field of multi-modal learning, and comprises the following steps: obtaining multi-modal original data, sequentially processing image, text, audio and video data, and extracting visual features of the image, semantic features of the text, frequency spectrum and time sequence features of the audio, image features and time sequence features of a video frame and time domain features of an audio sequence; and then, according to the complementary information of the multi-source features, fusion processing is carried out to form a unified multi-modal feature representation, the unified multi-modal feature representation is input to a preset deep learning analysis model, and finally a multi-modal analysis result of comprehensive expression is obtained. According to the scheme, information complementarity and robustness are enhanced through multi-modal feature fusion, the comprehensive analysis capability of the model on semantic understanding, behavior recognition and state judgment in a complex scene is remarkably improved, and a more accurate, efficient and stable decision basis is provided for a multi-modal intelligent sensing system.
Owner:JIANGSU FENGYUN TECH SERVICE CO LTD

Data processing method and apparatus, electronic device, computer readable storage medium and computer program product

The present application provides a data processing method and apparatus, an electronic device, a computer readable storage medium and a computer program product. The method comprises: acquiring historical interaction information and a predicted interaction text corresponding to the historical interaction information; extracting a first acoustic feature and a first semantic feature of the historical interaction information, and extracting a second semantic feature of the predicted interaction text; performing fusion mapping on the basis of the first acoustic feature, the first semantic feature and the second semantic feature to obtain a first paralanguage feature; denoising initial noise on the basis of the second semantic feature and the first paralanguage feature to obtain a second acoustic feature of the predicted interaction text; and on the basis of the second acoustic feature, generating a voice signal corresponding to the predicted interaction text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-agent cooperation strategy generation method and device, equipment and medium

The invention relates to a multi-agent cooperation strategy generation method, device and equipment and a medium, and the method comprises the steps: separating acoustic spectrum features and text semantic features of conference voice through environment perception processing, solving a cross-modal information conflict problem, and generating an accurate semantic understanding result; identifying the essence of the problem based on task analysis, associating the responsibility field, and constructing a classifiable problem point set; calling an agent capability library to dynamically match problem requirements, and generating a candidate agent list; quantifying a problem influence range and a decision time limit through weighted emergency scores, and generating a priority-sorted agent sequence; screening and confirming a core problem point and a primary agent; and finally, generating an executable cooperation scheme through multi-agent collaborative optimization. According to the method, the problems of incomplete feature extraction, task allocation delay and resource conflict in the prior art are solved, and the operability and decision-making efficiency of a cooperation strategy are remarkably improved.
Owner:SHAOGUAN XINGCHENG NETWORK TECH CO LTD

Dynamic sensitive information filtering system and method based on context semantic understanding

The invention discloses a dynamic sensitive information filtering system and method based on context semantic understanding, and relates to the technical field of information security and natural language processing. Comprising the steps of 1, creating a dynamic sensitive information filtering system, 2, carrying out cleaning, structuring and standardization processing on an input text through a text preprocessing module, 3, capturing deep semantic features of preprocessed text data through a semantic feature extraction module by utilizing a deep learning model, constructing a context-associated semantic representation space, and carrying out dynamic sensitive information filtering on the context-associated semantic representation space. 4, performing multi-level sensitive information detection based on the semantic features through a sensitive information identification module, and identifying the type, the position and the risk level of the sensitive content; 5, on-line iteration of knowledge base and model ability is carried out through a dynamic updating module to cope with dynamic changes of sensitive information types, and 6, safety disposal is carried out on detected sensitive information through a result output module, a filtering result is output, auditing tracing ability is provided, and the auditing tracing ability is provided. And 7, forming a system optimization closed loop through a feedback mechanism module according to user feedback and manual auditing, wherein the system optimization closed loop is used for continuously improving the detection accuracy and adaptability.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

Document processing method and system based on text content extraction

The invention relates to a document processing method and system based on text content extraction. The method comprises the steps that an original document containing text, image and format information is received, the encoding format of the document is automatically detected, character set conversion is executed, and hierarchical indexes including page numbers, paragraphs and tables are established for an unstructured document; the method comprises the following steps: synchronously processing text content and visual layout through a pre-trained visual-language model, extracting word-level and sentence-level semantic features by a text stream embedding layer, analyzing spatial distribution features of document elements by a visual encoder, and fusing text and visual features through a cross-modal attention mechanism; and loading the domain knowledge graph matched with the document type, and executing entity linking to associate the text mentions to the knowledge nodes. According to the document processing method and system based on text content extraction, through the synergistic effect of vision-text joint coding and knowledge enhancement, the accuracy of financial contract key clause recognition tasks is improved, the error rate is lower than that of industry benchmark products, and the semantic understanding precision is remarkably improved.
Owner:WIN THE BID HUIKANG TECH CO LTD

Illegal content auditing method and device based on multi-modal data, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses a violation content auditing method, device and equipment based on multi-modal data and a medium. Inputting the visual semantic features and the composite audio features into a multi-modal model, generating fusion features through model alignment and fusion, and analyzing the fusion features based on a knowledge base to judge whether illegal content fragments exist in the multi-modal data, and when the illegal content fragments exist, positioning the illegal content fragments in the multi-modal data and generating an auditing report. According to the method, the visual semantic features and the composite audio features are fused, cross-modal compliance analysis is realized in combination with the knowledge base, frame-level or time-axis-level positioning is performed on the illegal content segments, and the auditing report containing the evidence is generated, so that the problems of insufficient single-modal detection accuracy and poor positioning capability are solved, and the auditing accuracy is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Semantic recognition system and method based on heterogeneous graph attention network and dynamic normalization

The invention discloses a semantic recognition system and method based on a heterogeneous graph attention network and dynamic normalization, and belongs to the technical field of natural language processing and artificial intelligence. The system adopts a dual-channel architecture and comprises a general semantic channel and a domain semantic channel, semantic feature extraction is performed through a DIFF attention mechanism and a ToST statistical attention mechanism, and training stability is improved by adopting a DyT dynamic normalization module. Adaptive fusion of cross-channel semantic features is realized through a GeGLU gating mechanism, and high-precision semantic recognition is realized by combining an improved SimCSE + + comparison learning loss and a local minimization editing strategy of semantic perception. According to the method, the problems of inaccurate semantic expression, poor context adaptability and the like in the prior art are solved, the accuracy and applicability of cross-domain semantic recognition are remarkably improved, and the method can be widely applied to scenes of legal document processing, financial document analysis and the like.
Owner:GUANGZHOU ELECTRIC POWER ENG SUPERVISION CO LTD

Internet big data extraction method and device, equipment and storage medium

The invention relates to an internet big data extraction method and device, equipment and a storage medium, and the method comprises the following steps: carrying out distributed crawler collection on an internet data source, obtaining original network data, and converting the original network data into a structured data matrix; and performing multi-level semantic analysis on the matrix, constructing a semantic feature map, and performing topic segmentation and classification to form a topic domain knowledge tree. Association rules in the knowledge tree are further mined, and an implicit knowledge network is constructed. Performing semantic decomposition and expansion on query conditions based on the network to generate expanded query data, and performing similarity matching with the knowledge network to obtain a candidate data set; and finally, performing multi-factor sorting and extraction on the candidate data, and outputting target data, thereby solving the technical problem that in the second-hand car market, due to wide data sources and fuzzy semantics, an existing system has relatively large deviation when performing price prediction and maintenance cost analysis.
Owner:QINGDAO WULIANG TECHNOLOGY CO LTD

Abnormal scene detection method based on visual and semantic feature fusion

The invention discloses an abnormal scene detection method based on visual and semantic feature fusion, and relates to the technical field of safety monitoring and intelligent identification, and the method comprises the steps: carrying out the preprocessing of a collected original image, and obtaining a preprocessed image; forming a multi-modal input pair by the preprocessed image and a predefined structured prompt statement; inputting the multi-modal input pair into the visual language large model, and outputting semantic features including visual feature vectors and text vectors; inputting the preprocessed image into a target detection model, and outputting visual features; fusing the semantic features and the visual features through a cross-modal attention mechanism to obtain multi-scale fusion features; and inputting the multi-scale fusion features into detection heads of all scales, executing abnormal scene detection, and outputting an abnormal detection result. When the unconventional object is identified in the abnormal scene, the visual features and the semantic features are fused to perform abnormal scene detection, so that the strong perception capability of a complex scene is realized, and false alarm or missing alarm is effectively avoided.
Owner:CHONGQING UNIV OF ARTS & SCI

Text-driven CAD modeling method and system based on diffusion and visual language model

The invention relates to the technical field of computer aided design, in particular to a text-driven CAD modeling method and system based on a diffusion and visual language model.The method comprises the steps that natural language text description is obtained, and CAD semantic features of the natural language text description are extracted; carrying out geometric standardization on the CAD semantic features by adopting a fine-tuning diffusion model, and generating a CAD view image conforming to engineering specifications; carrying out fusion by adopting a fine-tuned visual language model to generate a parameterized CAD construction sequence; a three-mode alignment mechanism is adopted, and the semantic consistency of the CAD semantic features, the CAD view images and the CAD construction sequences is checked; performing verification and post-processing on the CAD construction sequence, and outputting an executable Python code or STEP file; the CAD modeling method disclosed by the invention performs explicit modeling based on flexible modal description, and has the characteristics of high geometric constraint and high usability.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Sentiment analysis method based on prototype guide mode fusion and prompt enhancement

The invention discloses a sentiment analysis method based on prototype guide mode fusion and prompt enhancement, and constructs a multi-mode sentiment analysis network which comprises a multi-mode coding module, a prototype guide mode fusion module, a dynamic mode weight adjustment mechanism and a context prompt generation module. The method comprises the following steps: firstly, extracting semantic features of each mode by using a multi-mode encoder, and constructing a prototype feature library based on a labeled sample to describe typical representations of different modes under each category; and then, dynamically evaluating modal contribution through prototype similarity to realize modal adaptive fusion. Furthermore, a context prompt is generated according to a similarity retrieval result of the input sample and the prototype library, and the pre-training language model is guided to complete sentiment classification. According to the method, the problems of modal inconsistency, information redundancy, weak small sample generalization and the like can be effectively relieved, and the accuracy and robustness of sentiment analysis are improved.
Owner:SOUTH CHINA UNIV OF TECH

AI Agent agent implementation method based on large language model and knowledge graph

The invention discloses an AI Agent implementation method based on a large language model and a knowledge graph, and relates to the technical field of artificial intelligence, and the method comprises the steps: taking a semantic feature embedding vector as a query basis, obtaining an entity node and a relation path in the knowledge graph, carrying out structure embedding coding, and generating a structure embedding vector; performing two-channel semantic structure alignment on the semantic feature embedding vector and the structure embedding vector, and performing alignment processing on entity nodes and relation paths in the knowledge graph to generate a path candidate set; performing structured coding on the path candidate set to generate prompt information, combining the prompt information with a natural language instruction, inputting the combined prompt information into a large language model, and generating a response draft; and performing path consistency verification on the response draft, if the response draft is not consistent, executing rollback and regenerating a new response text draft, and if the response draft is consistent, outputting a final response text. According to the method, the quality and the credibility of the response text are remarkably improved, and the interaction stability and the user experience of the intelligent agent are enhanced.
Owner:SHANGHAI INTERNATIONAL STUDIES UNIVERSITY

Recommendation method for enhancing semantics and interest perception by using large language model

The invention discloses a recommendation method for enhancing semantics and interest perception by using a large language model. The method comprises the following steps: firstly, performing semantic modeling on unstructured text information such as user comments, article description and the like by utilizing the powerful capability of a large language model in semantic comprehension and user preference modeling aspects, so as to improve the deep perception capability of a recommendation system on user interests and article semantic attributes; then, through a semantic feature alignment and discretization strategy, the problem that continuous semantic representation generated by a large language model is incompatible with features of a traditional recommendation system in the aspect of an expression structure is solved; finally, unified modeling of semantic information and traditional recommendation signals is achieved through a recommendation integration mechanism, and recommendation performance and model interpretability are improved.
Owner:SOUTHEAST UNIV

Defect identification and positioning method

The invention relates to the technical field of pipeline inspection, in particular to a defect identifying and positioning method. Comprising the following steps: generating a uniform node feature tensor through coordinate mapping and feature fusion by synchronously collecting a pipeline inner wall image, an ultrasonic echo and an electromagnetic eddy current signal; constructing a space-time heterogeneous feature graph, integrating three types of relationships of a space adjacent edge, a time evolution edge and a semantic similarity edge, and dynamically optimizing a graph structure by utilizing a trainable fusion factor; a heterogeneous edge decoupling convolution and dynamic attention mechanism is designed, space-time semantic features are extracted through channels, neighborhood information is aggregated, and high-resolution defect classification is achieved; based on a classification result and a residual tensor of an original feature, a defect space position is accurately predicted through a coordinate inversion network, and positioning robustness is improved by combining positioning confidence score and weighted aggregation; and finally fusing the equipment track and the pipeline three-dimensional model to realize defect geographic coordinate mapping and interactive visualization. According to the method, the defect identification precision and the positioning reliability in a complex pipeline environment are remarkably improved.
Owner:SHAANXI TAINUOTE TESTING TECH CO LTD