Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20 results about "Modal data" patented technology

An automated welding lack of fusion defect detection method

PendingCN122262800ABiological modelsFeature extractionModal data
The present application relates to the technical field of welding defect detection, and particularly relates to an automatic welding un-melted defect detection method; the present application provides an automatic welding un-melted defect detection method, based on multi-modal Transformer, through multi-modal data synchronous acquisition, space-time alignment, feature extraction and cross-modal correlation fusion, the detection precision and reliability are significantly improved; on the embedded platform, the single frame end-to-end inference time is about 10 ms, the power consumption is less than or equal to 40 W, and stable operation can be realized under the condition of 50 Hz sampling frequency, so that real-time monitoring of the welding process is realized.
Owner:AOTAI ELECTRIC

A historical textual research analysis method, device and equipment

PendingCN122364485AData matchingData set
This application relates to the field of digital humanities technology and discloses a method, apparatus, and device for historical verification analysis. The method includes: acquiring target data to be verified and determining modal data matching the target data, the modal data including first modal data and second modal data; querying target reference data matching the modal data to form a set of reference data matching the modal data; performing confidence screening on the reference data set to determine a set of candidate data matching the modal data, the candidate data set including a first candidate set and a second candidate set; performing cross-modal semantic consistency matching on the first and second candidate sets; constructing a text-image association graph based on the matching results; correcting the edge weights of the text-image association graph; and determining the historical verification information of the target data based on the corrected text-image association graph. The technical solution provided by this application can improve the reliability and accuracy of historical verification.
Owner:ZHEJIANG UNIV

Multi-modal data processing method, device and equipment based on uncertainty evaluation

PendingCN122451791AModal dataAlgorithm
The application provides a multi-modal data processing method, device and equipment based on uncertainty evaluation, and relates to the technical field of data processing. The method comprises: acquiring multi-modal data of a to-be-tested object. A plurality of perturbation samples are generated by performing a plurality of perturbation processes on the multi-modal data. The perturbation processes comprise an enhancement perturbation on the multi-modal data and / or a semantic equivalent perturbation on a preset input prompt word. Each perturbation sample is input into a multi-modal large model to generate a plurality of perturbation detection results. According to the plurality of perturbation detection results, a confidence and a target detection result corresponding to the confidence are generated and output; the target detection result is reference information obtained by evaluation. The method is used to achieve the effect of the accuracy of the confidence of uncertainty evaluation.
Owner:HUAGONG TECHNOLOGY CO LTD +1

Method, device, electronic equipment and product for generating multi-modal data

PendingCN122286077AData setModal data
This disclosure relates to methods, apparatuses, electronic devices, and computer program products for generating multimodal data. The method includes generating second multimodal data from a computing device based on a first data source (either unimodal or multimodal), random noise, and user input, wherein the number of modalities in the second data is greater than or equal to the number of modalities in the first data; wherein the computing device includes a first encoder, a generator, a decoder, and an editor. In this manner, an effective method for expanding datasets is provided, enriching the diversity of styles and content in synthetic data.
Owner:ROBERT BOSCH GMBH

A multi-modal data automatic labeling method, device, equipment, storage medium and program product for the power field

PendingCN122451480AModal dataSemantic matching
The application relates to a power field-oriented multi-modal data automatic labeling method, device, equipment, storage medium and program product. The method comprises the following steps: respectively extracting single-modal features of each data modality in power multi-modal data, and fusing to obtain multi-modal fusion features; calculating the semantic matching degrees of the multi-modal fusion features and each preset label in a pre-constructed domain knowledge base, and determining candidate labels; based on a pre-constructed domain rule base, checking the matching logic between the business attributes contained in the candidate labels and the scenes where the power multi-modal data is located, and obtaining a rule checking result; based on the contents reflected by different data modalities in the power multi-modal data, performing cross-modal consistency comparison on the candidate labels, and obtaining a consistency checking result; and determining a target labeling label of the power multi-modal data according to the rule checking result and the consistency checking result. The method can improve the accuracy of automatic labeling of power business multi-modal data.
Owner:SOUTHERN POWER GRID DIGITAL GRID RESEARCH INSTITUTE CO LTD

A multi-modal data processing method based on digital twinning

PendingCN122451800AModal dataAlgorithm
The application discloses a kind of based on digital twinning multi-modal data processing method, including the following steps: obtaining object basic data and standard multi-modal data;Construct digital twin, and determine target entity object, virtual twin object and twin state parameters in digital twin;According to the corresponding relationship between target entity object, virtual twin object, twin state parameters and standard multi-modal data, establish twin mapping relationship;Obtain multi-modal observation characteristics;According to twin mapping relationship, construct twin layer system association diagram;According to twin mapping relationship, obtain correction restriction mapping;According to the improved Sheaf neural network, obtain multi-modal state representation;Carry out consistency comparison and correction, obtain twin correction state representation;Update twin state parameters, and generate the twin processing result of digital twin.The application utilizes digital twinning and improved Sheaf neural network, realizes multi-modal data credible fusion, and is accurate in updating, strong in anti-interference.
Owner:BEIJING WANXIANG INSTANT ANALYSIS TECHNOLOGY CO LTD

Model pre-training optimization method and system for fusing multi-modal data

ActiveCN120579141BModal dataNoise
The application provides a model pre-training optimization method and system for fusing multi-modal data, and relates to the technical field of multi-modal technology. The method comprises the following steps: after obtaining voice modal data and other modal data, sampling at low resolution and high resolution two scales respectively, generating corresponding voice vector sets and calculating uncertainty metrics; globally aligning the voice and other modal data at the first resolution scale, and transferring the global alignment result to the second resolution scale to perform refinement correction, and finally returning to the first resolution scale, thereby forming a multiple mesh format back-and-forth iteration; by monitoring and observing noise and state noise in real time and integrating them into the uncertainty metric, the weighted coefficient can be dynamically reduced during training for high-noise sections, and a higher weight is given to low-noise or stable sections to strengthen effective features; the application can adaptively suppress noise interference and retain voice mutation details in a multi-modal scene, and has better robustness and generalization performance.
Owner:NANJING TORTOISE & HARE RACE SOFTWARE RES INST CO LTD +1

Video forensics method based on transformer and quantum characteristics

ActiveCN122027849BEliminate the effects ofachieve stable recognitionQuantum computersBiological modelsModal dataQuantum technology
The application provides a video forensics method based on a Transformer and quantum characteristics, and belongs to the technical field of multimedia content security and computer vision, and the method comprises the following steps: S1, collecting multi-modal original data; S2, pre-processing multi-modal data; S3, preliminarily detecting single-modal authenticity; S4, quantum-optimized multi-modal feature fusion; and S5, comprehensively determining multi-modal authenticity. Through multi-modal cooperation and quantum technology innovation, the application effectively solves the problems of insufficient robustness and inaccurate positioning of traditional forensics technology, and provides an efficient, accurate and feasible technical solution for video content authenticity verification.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Real-time processing method of multi-modal data for internet of things gateway

ActiveCN121691294BData setModal data
This invention relates to the field of IoT data processing technology, and more particularly to a real-time multimodal data processing method for IoT gateways. The method includes: acquiring multi-source operational data and encapsulating it into a standard operational dataset; acquiring high-bandwidth modal data from the standard operational dataset; when the maximum link latency time exceeds a preset latency anomaly threshold, acquiring the version effective fingerprint of the target model and the fingerprint of the local model; calculating information redundancy when obtaining a model version anomaly diagnosis result; acquiring a corrected latency time and comparing it with a preset normal baseline latency interval; and adjusting the current decision cycle when the corrected latency time is greater than or equal to the upper limit of the normal baseline latency interval. This invention achieves accurate processing of high-bandwidth modal data by real-time monitoring of latency, diagnosing model version consistency, eliminating redundant information, and performing link adjustments, thereby improving the real-time performance and reliability of IoT gateway data processing.
Owner:BEIJING KINGDOES RFID TECH

A multi-modal fusion method and system based on confusion perception and reliability gating

The application provides a multi-modal fusion method and system based on confusion perception and reliability gating, and relates to the technical field of modal fusion. The method comprises the following steps: extracting explainable features and deep semantic features of electroencephalogram and audio multi-modal data, fusing the two types of features of the same mode to obtain modal fusion features; obtaining initial posterior features of each mode based on the fusion features, and correcting the initial posterior features of each mode according to the confusion mode formed in the training stage and the category distance prior in the corresponding modal representation space; constructing a conflict-uncertainty joint gating mechanism based on the conflict information between modes, the uncertainty information of each mode and the modal global reliability prior, to obtain the dynamic fusion weight of each mode; and weighting and fusing the corrected posterior features of each mode based on the weight to obtain the final result, so as to improve the accuracy, stability and robustness of multi-modal classification in a complex scene.
Owner:SHANDONG UNIV

Intelligent exhibition hall interaction control method and system based on multi-modal data

The application discloses an intelligent exhibition hall interaction control method and system based on multi-modal data, and relates to the technical fields of human-computer interaction and signal processing. The method comprises the following steps: acquiring multi-source original data in an interaction space, extracting a body pointing space vector, an end motion speed sequence and a micro-motion high-frequency tremor sequence, and performing space-time synchronous alignment; performing power spectrum density analysis on 6Hz-12Hz signal components to generate a first compensation factor representing physiological load and latching; performing second-order difference operation on the end motion speed sequence to match an exponential decay model, and generating a first convergence probability; dynamically adjusting an intention judgment adjustment factor by using the first compensation factor, and realizing nonlinear coupling of the first convergence probability and the space vector to generate an intention index value; and triggering a control signal if the intention index value meets the standard. The application converts physiological tremor into a system gain variable, cooperates with a dynamic convergence constraint to repair a feature mismatch problem, and improves the certainty and stability of determination under complex working conditions.
Owner:JIANGSU ZHIYUAN MODERN SUPPLY CHAIN CO LTD

Model training method and system, data processing method and system, and electronic device

PCT designated stageWO2026123749A1Other databases queryingNeural learning methodsData setModal data
The present disclosure relates to the technical field of data processing techniques and large models. Disclosed are a model training method and system, a data processing method and system, and an electronic device. The model training method comprises: acquiring multi-modal training data from a target multi-modal data set; using a multi-modal pre-trained model to perform concatenation and mixing on the multi-modal training data, so as to obtain multi-modal representations; and using the multi-modal representations to adjust the multi-modal pre-trained model, so as to obtain a target multi-modal representation model, wherein the target multi-modal representation model is used for performing knowledge retrieval and response content generation on multi-modal query data input by a user, so as to obtain a target answer. The present disclosure solves the technical problems in the related art of heavy computational and storage burdens, limitations in information retrieval, and failing to be applied to composite retrieval scenarios caused by the use of a CLIP model for multi-modal representation.
Owner:ALIBABA (CHINA) CO LTD

Short video recommendation method and system based on heterogeneous graph neural network for fusing multi-modal data

ActiveCN119089004BModal dataTemporal context
The application discloses a short video recommendation method and system based on a heterogeneous graph neural network by fusing multi-modal data, and aims at the problem that current short video recommendation methods do not fully consider rich text, image, audio and other multi-modal information in the video, still adopt a simple multi-modal feature adding mode, and lack deep mining of the correlation between different modes. Meanwhile, the potential interest features in various interactive behaviors (such as browsing, liking, collecting and the like) of the user are neglected, resulting in the limitation of the recommendation result. The application constructs a heterogeneous content encoder containing an attention mechanism to fuse the features of different modes, uses a heterogeneous graph neural network to mine the potential features of the short video. At the same time, the time context information is introduced, the potential features in various behaviors of the user are mined by using a graph perception network, then the graph contrast is used to reduce the influence caused by the sparseness of the various interactive behavior supervision signals, and finally accurate short video recommendation is provided for the user.
Owner:INNER MONGOLIA UNIV OF TECH

Subway station multi-modal data noise processing method, computer device and program product

PendingCN122333304AModal dataEngineering
The application provides a subway station multi-modal data noise processing method, computer equipment and program product, which is applied to a subway station multi-source perception data processing scene, and is used for realizing noise identification and abnormality repair in a multi-modal data fusion process. First, each modal data from different devices is subjected to space-time alignment processing, multi-source data is uniformly mapped to a same time slice and a station space unit to obtain corresponding multi-modal data, on the basis of which, a dynamic fusion feature of a current time slice is constructed for each station space unit, noise embedded in the dynamic fusion feature is identified based on cross-modal consistency features in the dynamic fusion feature and modal data feature representations, an abnormal modal is located, finally, target modal data feature representations corresponding to the abnormal modal are repaired, and the repair result is written back to the dynamic fusion feature for updating, so as to suppress noise propagation and improve the accuracy of the fusion result.
Owner:BWTON TECH CO LTD

Hierarchical cim architecture system, scheduling method and device, equipment and storage medium

PendingCN122132350ADecoupling structural heterogeneous conflictsImprove efficiencyProgram initiation/switchingResource allocationModal dataBottleneck
The present disclosure relates to a layered CiM architecture system, a scheduling method, an apparatus, a device and a storage medium. The method comprises: a control layer and a 3D memory-in-compute array connected vertically, the 3D memory-in-compute array is integrated with a plurality of memory-compute layers for different modalities; the control layer is used for modal perception on acquired modal data, and the modal data is scheduled to a memory-compute layer in the 3D memory-in-compute array according to the result of the modal perception; the 3D memory-in-compute array is used for selecting a memory-compute layer corresponding to the modal data to infer the modal data according to the processing requirement of the modal data after receiving the modal data. The unreasonable bottleneck of bandwidth allocation of data loading is broken, the efficiency, energy efficiency and long-time running stability of edge multi-modal inference are significantly improved, and the resource constraints of edge scenarios are fully adapted.
Owner:HANGZHOU WEIHENG TECHNOLOGY CO LTD

A communication method, device and storage medium

A communication method, device and storage medium, relating to the technical field of communication. The method comprises: sending uplink assistance information (UAI), wherein the UAI comprises first indication information, the first indication information indicates a mapping relationship between a quality of service flow identifier (QFI) and a multi-modal service, and the first indication information comprises one or more bitmaps (bitmaps), each bitmap indicates the mapping relationship through the position of a bit with a first set value. The method can make the network device clear the data streams of multiple modes belonging to the same service, so as to facilitate the coordinated management and control of the data streams, for example, when performing cell switching, the data streams of multiple modes belonging to the same multi-modal service can be simultaneously switched to a target cell or released, thereby avoiding the loss of part of the multi-modal data streams after switching, causing the multi-modal service to be abnormal or even interrupted, and thus the user experience is ensured.
Owner:HONOR DEVICE CO LTD

Method, device and medium for fusing multi-modal data augmentation of industrial equipment mechanisms

PendingCN122451468AData setModal data
The application discloses a multi-modal data augmentation method fusing industrial equipment mechanism, equipment and medium, and relates to the technical field of data processing. The method comprises the following steps: collecting multi-modal data of multiple industrial data sources, and performing timestamp unification alignment processing on the multi-modal data to generate an aligned multi-modal data set; according to a preset multi-dimensional labeling system, performing label processing on the multi-modal data set to generate a labeled multi-modal data set, and constructing a mechanism knowledge base based on the mechanism knowledge of the industrial equipment; an initial condition diffusion model is constructed, and the knowledge in the mechanism knowledge base is taken as condition information and embedded into the training process of the initial condition diffusion model to generate a target loss function; based on the labeled multi-modal data set and the target loss function, the initial condition diffusion model is iteratively trained, and random noise data and target conditions are input into the trained mechanism condition model to output synthetic multi-modal data samples conforming to the industrial equipment mechanism.
Owner:SHANDONG JINGBO POLYOLEFIN NEW MATERIALS CO LTD

Apparatus, operation method of apparatus, and computer-readable non-volatile recording medium on which program for performing operation method of apparatus is recorded

PCT designated stageWO2026146731A1Computer hardwareModal data
An apparatus according to one embodiment of the present disclosure may comprise: an output interface; a memory for storing pieces of multi-modal processed data of a preset time unit, generated on the basis of pieces of multi-modal data; and one or more processors for acquiring a voice command, retrieving, from the memory, multi-modal processed data related to the acquired voice command, generating a prompt on the basis of the retrieved multi-modal processed data and the acquired voice command, acquiring a response result from the prompt through an artificial intelligence model, and outputting the acquired response result through the output interface.
Owner:LG ELECTRONICS INC