Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

906 results about "Semantic space" patented technology

Semantic spaces in the natural language domain aim to create representations of natural language that are capable of capturing meaning. The original motivation for semantic spaces stems from two core challenges of natural language: Vocabulary mismatch (the fact that the same meaning can be expressed in many ways) and ambiguity of natural language (the fact that the same term can have several ...

Knowledge graph construction method and system based on large language model technology

The invention relates to the technical field of knowledge graph construction, and discloses a knowledge graph construction method and system based on a large language model technology. The method comprises the following steps: receiving a multi-source heterogeneous data stream, and completing semantic space mapping and cross-modal feature fusion to generate a unified semantic representation vector set; constructing an initial knowledge graph skeleton; performing incremental optimization on the skeleton, and performing entity relationship disambiguation and conflict detection; and iteratively updating the knowledge representation, and outputting a target knowledge graph meeting semantic consistency. The system comprises a data receiving module, a semantic fusion module, a skeleton construction module, an optimization module and a knowledge updating module. According to the method, multi-source heterogeneous data is effectively processed, the accuracy, the dynamic updating capability and the semantic consistency of the knowledge graph are improved, and the method has wide application prospects in the fields of intelligent question answering, information retrieval and the like.
Owner:NAVAL AVIATION UNIV

Multi-modal AI data fusion processing method and device, equipment and medium

The invention relates to a multi-modal AI data fusion processing method, device and equipment and a medium, and the method comprises the steps: firstly extracting visual, auditory and text modal features through a pre-training encoder, executing dimension alignment, and generating a standard data feature set with unified dimensions; a cross-modal semantic graph is constructed based on a cosine similarity algorithm, and the problem of semantic mismatch of heterogeneous data is solved; residual enhancement is carried out on the map nodes, and noise interference is eliminated; fusing the optimized features and the semantic topology in combination with a graph convolutional network to generate aggregation graph representation; the fusion features are mapped to a low-dimensional semantic space through a variational auto-encoder, and cross-modal correlation essence is captured; the key dimension contribution degree is quantified, a visual report is generated, and semantic association rules among modals are disclosed, so that the dimension isomerism limitation of a traditional fusion technology is broken through, quantifiable cross-modal semantic mapping is established, the whole process traceability from feature fusion to decision interpretation is realized, and the method is suitable for popularization and application. And the multi-modal decision black box problem in the fields of medical diagnosis, automatic driving and the like is effectively solved.
Owner:罗林松

Rare disease knowledge graph construction method based on modal injection and multi-modal fusion

The invention relates to the technical field of medical artificial intelligence and knowledge graph construction, in particular to a rare disease knowledge graph construction method based on modal injection and multi-modal fusion. Comprising the following steps: S1, collecting multi-modal medical information including texts, images and genes; s2, standardization processing is carried out, and a three-layer metadata structure is constructed; s3, complementing missing modal data, and performing feature extraction and unified dimension conversion on the modal data to realize representation alignment in a shared semantic space; s4, performing multi-level semantic fusion to obtain a unified fusion semantic vector; and S5, constructing a double-layer structure system rare disease knowledge graph comprising an ontology layer and an instance layer. According to the method, multi-modal medical information of texts, images and genes is selected to construct the knowledge graph of the rare disease, the application range, coverage and accuracy of the knowledge graph are improved, correspondence adaptation of rare cases during clinical diagnosis and treatment of the rare disease can be achieved, and the method has high recognition capacity.
Owner:湖南工商大学

Fragmented data cross-modal label generation system and method based on deep transfer learning

The invention provides a fragment data cross-modal label generation system and method based on deep transfer learning. The method comprises the following steps: extracting first high-dimensional feature vectors in different modes; mapping the first high-dimensional feature vectors of different modals into the same semantic space through a cross-modal comparison loss function to realize multi-modal alignment and fusion to obtain second high-dimensional feature vectors; labeling semantic tags corresponding to the second high-dimensional feature vectors based on the fragmented data components by adopting a small sample transfer learning algorithm; a multi-channel Hash encoder is adopted, a self-adaptive encoding strategy is called according to different modal data combinations, and the second high-dimensional feature vector is encoded into a multi-channel binary Hash code; in combination with an incremental graph neural network, the binary hash codes and the corresponding semantic tags are dynamically expanded into the historical knowledge graph; matched fine-grained tags are established for semantic differentiation features of different entity combinations in the target knowledge graph, and a cross-modal tag tree is obtained by combining three-matrix hierarchical construction.
Owner:LONGMA ZHIXIN (ZHUHAI HENGQIN) TECH CO LTD

Judicial scene-oriented multi-modal data fusion method and system

The invention discloses a judicial scene-oriented multi-modal data fusion method and system. The method comprises the following steps: 1) collecting multi-modal data in a judicial application scene and converting the multi-modal data into data in a uniform format; 2) mapping the data to a unified semantic space to realize cross-modal alignment; then associating the multi-modal data to obtain the text description of the same case and the corresponding image evidence as the multi-modal features of the corresponding case; 3) constructing a knowledge graph of a judicial application scene based on the multi-modal features; driving multi-modal data fusion based on the knowledge graph and constructing a multi-modal evidence chain of each entity; 4) quantifying the integrity of the knowledge graph and the reliability of the multi-modal evidence chain according to a preset index, marking abnormal nodes in the knowledge graph according to a quantification result, and adjusting the abnormal nodes; 5) generating a structured knowledge graph according to the knowledge graph and creating a dynamic desensitization report; edges in the structured knowledge graph represent relationships between legal entities, and each relationship is bound with a multi-modal evidence chain.
Owner:CHINA NAT SOFTWARE & SERVICE

Multi-label electrocardiogram classification method based on self-supervised pre-training and multi-modal semantic alignment

The invention discloses a multi-label electrocardiogram classification method based on self-supervised pre-training and multi-modal semantic alignment, which belongs to the technical field of artificial intelligence, and comprises the following steps: realizing self-supervised pre-training of unlabeled data through a single-modal contrast enhancement network, generating global and local contrast views by adopting a multi-scale random cutting strategy, and classifying the global and local contrast views in a multi-scale random cutting mode; in combination with a teacher-student network architecture, the potential invariance features of the ECG signals are learned while negative sample dependence is avoided, the problem of annotation data scarcity is effectively relieved, and the feature robustness is improved. A multi-modal fusion mechanism based on label semantic guidance is provided, a time domain signal and a frequency domain time-frequency graph are mapped to a unified semantic space through fine-grained semantic alignment, local feature enhancement and cross-modal complementary information fusion are realized by using a cross attention mechanism, and the problem of semantic difference caused by modal heterogeneity in a traditional method is overcome. A multi-label comparison loss function based on a disease co-occurrence relation is proposed, a category discrimination boundary is dynamically optimized by modeling a label co-occurrence probability, the feature separability of a tail category is improved while the head category discrimination ability is enhanced, and the problem of sample category imbalance in a multi-label scene is remarkably relieved.
Owner:YANSHAN UNIV

Decision-making method and device guided by multi-modal semantic map, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a decision-making method and device guided by a multi-modal semantic map, equipment and a medium. Extracting a visual feature vector, a language feature vector and an action feature vector, splicing to generate a multi-modal initial feature, mapping the multi-modal initial feature to a shared semantic space, constructing a multi-modal semantic map, and inputting a map-guided attention mechanism to generate a cross-modal alignment feature; the cross-modal alignment features and task targets are input into a meta-learner to generate task adaptability features, the task adaptability features are input into a parallel reasoning network to execute subtasks in parallel, and a gating fusion network integrates output results to generate a global decision. According to the method, cross-modal semantic association and task adaptability are enhanced through the combination of shared semantic space mapping, map guiding attention and a meta learning device, and the accuracy and efficiency of multi-modal decision making are improved through the combination of parallel reasoning and gating fusion.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-modal automatic knowledge graph construction method based on large language model

According to the multi-modal automatic knowledge graph construction method based on the large language model, a multi-modal data stream is preprocessed, features are extracted, and the multi-modal data stream is mapped to a unified semantic space through a cross-modal alignment network after being processed through the large language model, a visual converter and a time sequence neural network. In the space, entities and categories are recognized based on a large language model, a triple is generated by combining a multi-modal feature judgment entity relationship, mapping fusion is performed through an ontology alignment algorithm driven by a graph neural network and a predefined domain ontology, finally knowledge is stored in a graph database, and dynamic updating is performed by means of incremental learning and online reasoning. Standardized APIs and visualization components are provided. According to the method, the construction efficiency and the automation degree of the knowledge graph are remarkably improved, the cross-modal information fusion and knowledge maintenance capability is enhanced, and the application requirements of intelligent retrieval, recommendation, decision support and the like are met.
Owner:BEIJING SPACEFLIGHT TUOPUGAO SCI & TECH CO LTD

Intelligent planning method and system for weak current system in smart park

The invention discloses an intelligent planning method and system for a weak current system in a smart park, and belongs to the technical field of weak current intelligent design. The method comprises the steps of performing feature extraction on the weak current multi-source data of the smart park to form a weak current feature set; a multi-dimensional semantic space is constructed, semantic association features are obtained, and node features, topological relations and constraint rules of the weak current system are determined; generating a weak current knowledge graph based on the information, and performing semantic alignment on the basic information of the park to obtain a final scene demand representation; performing graph reasoning and constraint calculation according to the representation to obtain a feasible region and constraint satisfaction condition, and generating a candidate construction scheme; and screening out an optimal construction scheme from the candidate schemes according to a preset comprehensive optimization strategy and sending the optimal construction scheme to a control center. According to the scheme, the weak current scheme is promoted from demand understanding to scheme optimization, and a coherent and verifiable automatic process is formed; therefore, the manual intervention is less, the design judgment is more accurate, and the finally output construction scheme has higher engineering reliability.
Owner:YITAIDA TECHNOLOGY CO LTD

VLM model intelligent decision-making-based driving method and device, and storage medium

PendingCN121291416AAlgorithmControl signal
The invention discloses a VLM model intelligent decision-making-based driving method and device and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: extracting key visual features from continuous multi-frame driving scene images based on a preset visual feature extraction algorithm; inputting a user instruction and a time sequence visual Token corresponding to the key visual features into a preset VLM model for multi-modal alignment, and generating a target planning Token; inputting the target planning Token and the time sequence vision Token into a preset trajectory generation model to obtain a predicted trajectory; and converting the predicted trajectory into a control signal, and controlling the mobile device to complete a moving action based on the control signal. The problem of modal difference between a semantic space and an action space is solved.
Owner:YOUDI ROBOT (WUXI) CO LTD

Cross-modal semantic alignment method based on multi-source heterogeneous data

The invention provides a cross-modal semantic alignment method based on multi-source heterogeneous data, and relates to the technical field of data processing.The method comprises the steps that a multi-source heterogeneous data set in a preset collection task scene is input, and modal feature extraction is executed; analyzing the multi-source heterogeneous feature data set, and outputting a plurality of information bearing capacity scores; obtaining a first heterogeneous feature data source, and generating a reference template for semantic alignment; and obtaining residual heterogeneous feature data sources, establishing an alignment mapping relationship between the reference template and the residual heterogeneous feature data sources, and outputting an aligned shared semantic space. By means of the method and device, the technical problem that in the prior art, due to the fact that data of different modalities are excessively different in structure and format, semantic alignment precision is uneven, semantic omission is likely to be caused, and semantic alignment accuracy and integrity are affected is solved, multi-source heterogeneous data are analyzed, and the semantic alignment accuracy and integrity are improved. And the data source with the highest information integrity is selected as the alignment template, so that the accuracy and integrity of semantic alignment are improved.
Owner:广州云趣信息科技有限公司

Knowledge graph agent construction method and system based on multi-modal fusion

The invention relates to a knowledge graph agent construction method and system based on multi-modal fusion, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring multi-modal data, and mapping different modal data to a unified semantic space through a cross-modal embedding technology to obtain multi-modal reference data; extracting features of the multi-modal reference data through an attention mechanism to obtain multi-modal joint features; constructing and updating a dynamic knowledge graph based on the multi-modal joint features to obtain a time sequence dynamic dependency knowledge graph; and according to the time sequence dynamic dependency knowledge graph, context-aware reasoning is carried out through a cross-modal reasoning engine to obtain a user input mode, and an interaction strategy is dynamically adjusted based on the user input mode to generate a multi-modal response. And the knowledge graph agent construction based on multi-modal fusion is realized.
Owner:SHANGHAI YUTA TECH CO LTD

Intelligent digital human training method and system based on multi-modal interaction

The invention discloses an intelligent digital human training method and system based on multi-modal interaction, and belongs to the technical field of semantic indexing.The method specifically comprises the steps that voice, vision and text data are analyzed and converted into high-dimensional feature vectors through a modal exclusive encoder, the high-dimensional feature vectors are projected to a unified semantic space through a cross-modal semantic mapping model, and the high-dimensional feature vectors are obtained; generating a semantic primitive containing a modal identifier, a core semantic tag and a feature weight; semantic primitives are used as nodes, directed edges and edge weight table association strength are established based on semantic similarity, typical scene node connection weights are strengthened, and a mesh map containing intra-modal hierarchy and inter-modal cross association is formed; constructing a double-layer index on the basis of the mesh map; semantic primitives are extracted from newly added data, the position of a new node in an association graph is determined through a graph matching algorithm, an association edge with an existing node is automatically established, and a lower-layer modal exclusive index is synchronously updated.
Owner:JIANGXI INST OF FASHION TECH

Radiology report generation method fusing disease perception comparison and cross-modal alignment

The invention relates to the field of radiology reports, and discloses a radiology report generation method fusing disease perception comparison and cross-modal alignment, and the method comprises the following steps: respectively extracting visual features of a medical image and text features of a radiology report through a visual encoder and a text encoder; utilizing a disease perception contrast learning module DACL to map the visual features and the text features to a unified semantic space; self-adaptive feature matching is carried out by using a dynamically updated memory matrix through a cross-modal memory alignment module CMA; a three-stage training strategy optimization model is adopted, and pre-training, reinforcement learning fine adjustment and knowledge distillation are included. Incremental knowledge is dynamically updated through text features, cross-modal feature alignment tasks are focused, coupling with feature representation is avoided, alignment efficiency is improved, a three-stage training strategy is adopted, if student model performance is better, weights are migrated, and clinical accuracy and generalization ability of the model are further enhanced.
Owner:CHONGQING NORMAL UNIVERSITY

Fine-grained zero-sample medical image classification method based on cross-modal feature alignment

The invention discloses a fine-grained zero-sample medical image classification method based on cross-modal feature alignment, and the method comprises the steps: 1, carrying out the partitioning of a full-section pathological image, and extracting the features of a local image block; 2, a cross-modal alignment module is used for designing a local window attention mechanism to enhance space interaction between image blocks; the semantic enhancement module is used for constructing a pathology prompt template based on a large language model to generate fine-grained category description, and expanding the distance between categories in a semantic space; and 4, performing weighted fusion on the image block features through coordinate sensing, and generating final slice-level classification prediction. According to the invention, through multi-scale space interaction of the cross-modal image block alignment module and semantic enhancement of the semantic refinement module based on the visual language model, the classification precision of the fine-grained medical image is significantly improved, and the limitation of the existing method on feature alignment and semantic differentiation is effectively solved; and an efficient solution is provided for zero-sample medical image classification.
Owner:UNIV OF SCI & TECH OF CHINA +1

Abnormal traffic detection and attack identification method and system based on deep learning

The invention belongs to the technical field of network security, and provides an abnormal traffic detection and attack recognition method and system based on deep learning, and the method comprises the steps: data preprocessing and feature extraction, cross-modal semantic alignment and knowledge graph construction, causal enhancement association reasoning, intelligent engine optimization, cloud edge collaborative resource scheduling, and result output. According to the method, statistical features and signature features are mapped to a unified semantic space through a cross-modal semantic alignment and knowledge graph construction module, a semantic barrier between heterogeneous features is broken through, time sequence causal discovery and transfer entropy calculation are introduced, a simple correlation and a reliable causal can be distinguished, and the method has a good application prospect. According to the method, the accuracy and credibility of attack chain reasoning are improved, the false alarm rate is reduced, online self-evolution of a detection model and dynamic optimal allocation of system resources are realized through intelligent engine optimization and cloud edge collaborative resource scheduling modules, and the overall adaptability, robustness and practicability of the system are enhanced.
Owner:BEIJING HENGAN JIAXIN SAFETY TECH CO LTD

Listed company operation risk early warning method based on multi-source auditing and text semantic fusion

The invention discloses a listed company operation risk early warning method based on multi-source auditing and text semantic fusion, and relates to the technical field of auditing, and the method comprises the steps: S1, crawling and converging multi-source heterogeneous data of listed company financial newspapers, auditing suggestions, supervision announcements, inquiry letters, news public opinions and market transactions; according to the method, unstructured texts are subjected to cleaning, blocking and semantic vectorization processing, each text segment is embedded into a high-dimensional semantic space, a vector index is established, a bottom-layer knowledge base of an RAG framework is formed, in the stage, it is ensured that the data structure is uniform, the source is traceable, standardized input is provided for subsequent semantic retrieval and modeling, and the reliability of the system is improved. S2, a query expression is constructed based on a target company, a time window and a risk topic, dense semantic retrieval and sparse BM25 retrieval methods are comprehensively used, a time decay and source credibility weighting mechanism is introduced, and the problems that a traditional method is single in data dimension and information is split are solved.
Owner:NANJING UNIV OF FINANCE & ECONOMICS

Multi-modal pre-training model construction method and system for monitoring video

The invention provides a multi-modal pre-training model construction method and system for monitoring videos, and the method comprises the steps: automatically constructing a high-quality multi-modal alignment data set through a single-modal description generation model and a large-scale language model, and remarkably reducing the marking cost; a special coding network and a shared projection layer are adopted to realize feature extraction and uniform semantic space alignment of video, audio and text modes; performing dynamic semantic fusion by using a modal collaborative attention mechanism; designing cross-modal contrast learning, mask prediction and time sequence consistency tasks to carry out multi-task pre-training; an external knowledge base is introduced, and the semantic reasoning ability is enhanced through a microretrieval mechanism; and optimizing model parameters by adopting a multi-task joint loss function and an end-to-end training strategy. According to the method, efficient and automatic construction, deep semantic alignment and fusion and intelligent reasoning of knowledge enhancement of monitoring video multi-modal data are realized, and the understanding and generalization ability of the model in a complex scene is effectively improved.
Owner:BEIJING JIAOTONG UNIV

Double-branch diffusion three-dimensional scene generation method based on multi-modal semantic graph

The invention belongs to the technical field of three-dimensional scene modeling, and discloses a dual-branch diffusion three-dimensional scene generation method based on a multi-modal semantic graph, which comprises the following steps of: firstly, receiving multi-modal data such as sketches, texts, automatic completion instructions and scene general knowledge, extracting features and fusing the features into a unified multi-modal semantic graph; utilizing a graph neural network and an attention mechanism to enhance semantic graph features, and optimizing physical constraints through a physical engine; complementing the missing visual modality and graph structure relationship; performing quality scoring on the scene based on semantics, spatial relationships and physical constraints; and finally, respectively generating a spatial layout and a geometric shape through a double-branch diffusion model, and ensuring the coordination of the layout and the shape. The method has the advantages of multi-modal information fusion, physical rationality guarantee, high structure complementation capability, high-quality score optimization and efficient generation process, and is suitable for three-dimensional scene modeling requirements in the fields of virtual reality, augmented reality, robots and the like.
Owner:CHINA JILIANG UNIV

File retrieval method, device and equipment based on multi-modal AI and storage medium

The invention discloses a file retrieval method, device and equipment based on multi-modal AI and a storage medium, relates to the technical field of data retrieval, and aims to solve the problem that traditional file retrieval is low in efficiency and precision. The method comprises the following steps: performing content analysis on a multi-modal file including a picture file, a video file and a document file of text and / or visual information, and extracting a text semantic feature, a picture visual feature and a video time sequence feature; mapping the text semantic feature, the picture visual feature and the video time sequence feature to a cross-modal semantic space used for representing a high-dimensional vector associated with different modal features, and generating a cross-modal association vector; according to the modal type of the retrieval information, a corresponding retrieval module in a mixed retrieval engine is called to conduct retrieval in a cross-modal semantic space, an initial retrieval result is obtained, and the mixed retrieval engine comprises a text retrieval module, a visual retrieval module and a cross-modal fusion module; and performing dynamic weight distribution sorting on the initial retrieval result, and outputting a target matching result.
Owner:SHANXI XINDINGCHEN TECH CO LTD

Animal scene-oriented adaptive multi-modal data fusion method

The invention relates to the technical field of data fusion, and discloses an animal scene-oriented adaptive multi-modal data fusion method, which comprises the following steps of: extracting spatio-temporal characteristics from multi-source heterogeneous data such as visual sense, auditory sense and physiological sensing, constructing an animal-environment-group ternary spatio-temporal relation graph, and constructing an animal-environment-group ternary spatio-temporal relation graph; a pilot frequency sampling problem is solved through an adaptive interpolation algorithm, cross-modal projection alignment is completed in a public semantic space, unified space-time representation is output, and confidence coefficient weight is dynamically calculated based on uncertainty measurement of each modal feature. According to the method, accurate alignment of multi-modal data is realized through the cross-modal space-time attention network, the multi-modal feature alignment error is reduced compared with that of a traditional LSTM method, the training data volume of a federated element migration reinforcement learning framework is reduced compared with that of a traditional migration learning method, and the cross-species generalization performance of the model is improved. A multi-level causal inference engine quantitatively reveals causal association between environmental factors and animal diseases, and in combination with a dynamic decision tree visualization technology, the decision recognition degree is improved.
Owner:INST OF SPECIAL ANIMAL & PLANT SCI OF CAAS +1

Moving target intelligent detection method and system based on weak supervision dynamic optimization

The invention provides a moving target intelligent detection method and system based on weak supervision dynamic optimization, and the method comprises the steps: respectively extracting video features and text features from an original video and a text, carrying out the fusion, and generating a frame-level semantic similarity score as a pseudo tag; utilizing learnable object query and fusion feature interaction to generate positive and negative proposal masks; guiding feature comparison learning of the positive proposal by using a pseudo tag, so that the positive proposal infinitely fits text features in a semantic space, and the negative proposal infinitely deviates from a related region of the text features; performing text reconstruction based on a mask condition Transform by using positive and negative proposal masks, and performing semantic consistency training on different proposals to obtain a video time domain positioning result; and dynamically optimizing a video time domain positioning result, generating a final positioning result, and completing intelligent detection of the moving target. According to the invention, by constructing a learnable negative proposal and a dynamic pseudo-label constraint mechanism, the time domain positioning precision under a weak supervision condition is significantly improved.
Owner:SHANGHAI SATELLITE ENG INST

Multi-modal content intelligent auditing and violation detection method and system

The invention discloses a multi-modal content intelligent auditing and violation detection method and system, and relates to the technical field of information processing. The method comprises the following steps: carrying out audio-picture separation on a video stream, and carrying out parallel processing on audio-to-text and visual key frame extraction; a space-time encoder is constructed to record the corresponding relation of the time stamps of all the modes; constructing a multi-modal resource target dictionary, and forming an inter-entity knowledge graph by using relation categories; calculating a confidence coefficient difference index between modals by comparing, learning and training the shared semantic space; and carrying out violation judgment, triggering a sensitive characteristic threshold value for any mode, and starting multi-mode evidence cross validation. According to the method, audio and picture separation is carried out on the video stream, parallel processing is carried out on audio-to-text and visual key frame extraction, a multi-modal resource target dictionary is constructed, violation judgment is carried out according to confidence coefficient difference indexes among modals, and false information auditing efficiency and detection efficiency are improved.
Owner:ZHENGZHOU JIERUAN INFORMATION TECH RES INST CO LTD

Unmanned aerial vehicle inspection system multi-modal data fusion and intelligent analysis platform and method for wind power plant

The invention discloses a multi-modal data fusion and intelligent analysis platform and method for an unmanned aerial vehicle inspection system for a wind power plant. The platform comprises a multi-modal data acquisition module, a feature extraction and standardization module, a multi-modal information fusion module, a joint learning and optimization module, a domain knowledge injection module and an intelligent decision and application module. The system processes multi-source heterogeneous data through an integrated learning and deep learning fusion strategy, projects features to a shared semantic space by using joint training and comparative learning to enhance the anomaly discrimination ability, and performs verification and semantic enhancement on a supervised retrieval result in combination with a knowledge base in the wind power field. And finally, outputting a high-reliability diagnosis report and a maintenance suggestion. According to the invention, accurate identification and positioning of the fan fault are realized, and the inspection efficiency and the system decision reliability are significantly improved.
Owner:CHINA RESOURCES NEW ENERGY (SUIXIAN TIANHEKOU) WIND ENERGY CO LTD

Robust audio and video speech recognition method and device based on multilayer perception fusion

The invention discloses a robust audio and video speech recognition method and device based on multilayer perception fusion, and belongs to the technical field of audio and video multi-mode semantic modeling and speech recognition. According to the method, audio and visual bimodal input is utilized, a teacher-student structure is introduced in a training stage, and a student model is guided to learn stable semantic representation under various noise conditions through a self-distillation mechanism. In order to enhance the alignment capability and anti-interference performance between audio and video features, a multi-layer suppression and enhancement interaction module is introduced into the joint encoder, layer-by-layer fusion and noise suppression between modes are realized, and a robust multi-mode fusion encoder (RMIE) is constructed. The RMIE models modal alignment and feature enhancement in a multi-level semantic space at the same time, and the semantic offset problem caused by modal difference and noise interference is effectively relieved. Furthermore, a decoder based on an attention mechanism is introduced on the basis of the RMIE, and an audio and video speech recognition model with end-to-end recognition capability is obtained through fine tuning.
Owner:SICHUAN UNIV

Code generation method, electronic equipment, readable storage medium and program product

The invention discloses a code generation method, electronic equipment, a readable storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the steps that code generation demand information and code generation cue words are input into a language model with a code generation function, and a code generation initial result is obtained; determining a semantic space model for defining a semantic space to which the code generation result belongs according to the code generation demand information, and converting the semantic space model into a mathematical model corresponding to the code generation scene to obtain a semantic verification model; inputting the code generation initial result into a semantic verification model, reasoning the code generation initial result, and determining a semantic verification result according to whether an end node can be reached or not; and determining a code generation result according to the semantic verification result and the code generation initial result. The problem that high-quality codes cannot be generated due to large model illusion in the prior art can be solved, and the code generation accuracy and reliability are effectively improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Entity alignment method based on large language model adaptive fusion

The invention discloses an entity alignment method based on large language model adaptive fusion. The method comprises the steps that multi-modal data is preprocessed, multi-modal data features are obtained, and the multi-modal data features comprise semantic features corresponding to text data, image features corresponding to image data and topological relation features corresponding to geographic space data; mapping the multi-modal data features to a unified semantic space to obtain an entity representation vector in the unified semantic space; for the entity representation vectors in the unified semantic space, dynamic weights are distributed according to confidence coefficients of different modal data, then weighted fusion entity representation is calculated based on the dynamic weights, alignment is carried out through comparative learning and attention mechanism matching, and aligned multi-modal fusion entity representation is obtained; and for the aligned multi-modal fusion entity representation, performing entity matching and geographic coordinate calibration to obtain an entity matching result. According to the invention, the accuracy and robustness of entity alignment are improved.
Owner:深圳市规划和自然资源数据管理中心(深圳市空间地理信息中心) +1

Multi-dimensional transmission data compression sensing method and system for urban emergency rescue

The invention relates to a multi-dimensional transmission data compression sensing method and system for urban emergency rescue, and the method comprises the steps: obtaining multi-modal data, carrying out the normalization processing and sparse representation, and obtaining the sparse coefficient of each modal; constructing a block diagonal random matrix for each modal sparse coefficient to perform independent compression sampling, generating a low-dimensional observation value before coding, and performing arithmetic coding and low-density parity check code error correction coding to generate an anti-interference code stream; and performing low-density parity check code decoding and arithmetic decoding on the anti-interference code stream to obtain a decoded low-dimensional observation value, performing iteration by using an orthogonal matching pursuit algorithm, recovering sparse coefficients of each mode, performing reconstruction to obtain text, image and audio features, and projecting the text, image and audio features to a shared semantic space to obtain a shared semantic space. And calculating a joint semantic weight through a cross-modal attention mechanism, and generating a multi-modal fusion result with consistent semantics. The method disclosed by the invention provides an efficient, robust and semantic collaborative end-to-end solution for urban emergency rescue.
Owner:GUANGDONG INTELLIGENT ROBOTICS INST

Multi-modal data identifier generation method and system based on semantic hash

The invention discloses a multi-modal information identifier generation method and system based on semantic hash. The method comprises the following steps: carrying out data preprocessing and multi-modal feature extraction on multi-modal data to obtain a unified representation containing rich semantic information; mapping the features to a shared semantic space by adopting an independent alignment projection network of each mode, and introducing cross-mode contrast learning and label supervision to realize semantic alignment among different modes; compressing the high-dimensional features after multi-modal data alignment to generate a Hash code with a fixed length; generating a unified semantic hash code containing semantic information of at least two modals; searching the number of times that the Hash code appears in a database through the unified semantic Hash code, generating a redundant code, and splicing the redundant code with the unified semantic Hash code to form an identifier of the sample; on the basis of the unified semantic hash codes, the hash codes related to the to-be-queried category in the test set are queried in the training set, and unified identification and efficient retrieval of different modes are achieved.
Owner:BEIHANG UNIV

Multi-modal large model incremental training data screening method

The invention provides a multi-modal large model incremental training data screening method, and relates to the technical field of data processing, and the method comprises the steps: executing modal structure analysis on newly added multi-modal data, extracting each modal vector, calculating a semantic matching degree, and removing samples lower than a preset first threshold value; calculating a multi-level semantic distance between a sample embedding vector and a historical clustering center in a unified semantic space, and dividing a core semantic region sample, a boundary semantic region sample and a discrete semantic region sample according to the change rate of the multi-level semantic distance; performing semantic fine-grained alignment on the boundary semantic region samples, when multimodal unstable distribution is detected, executing local context reconstruction to repair semantic deviation, and if the multimodal unstable distribution is still unstable, removing the semantic deviation; performing multiple rounds of small-batch reasoning, calculating a semantic stability coefficient based on a semantic prediction result, and when the semantic stability coefficient is lower than a preset second threshold value, determining that the sample is a potential drift sample and removing the potential drift sample; constructing an incremental training data set; according to the method, the autonomy and accuracy of incremental training data screening are improved.
Owner:ZHONGSHU (XIAMEN) INFORMATION TECH CO LTD +1