Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

58 results about "Multimodal learning" patented technology

The information in real world usually comes as different modalities. For example, images are usually associated with tags and text explanations; texts contain images to more clearly express the main idea of the article. Different modalities are characterized by very different statistical properties. For instance, images are usually represented as pixel intensities or outputs of feature extractors, while texts are represented as discrete word count vectors. Due to the distinct statistical properties of different information resources, it is very important to discover the relationship between different modalities. Multimodal learning is a good model to represent the joint representations of different modalities. The multimodal learning model is also capable to fill missing modality given the observed ones. The multimodal learning model combines two deep Boltzmann machines each corresponds to one modality. An additional hidden layer is placed on top of the two Boltzmann Machines to give the joint representation.

Agricultural load prediction method and system based on multivariate time sequence decoupling multi-modal learning

The invention discloses an agricultural load prediction method and system based on multivariate time sequence decoupling multi-modal learning, and belongs to the technical field of agricultural load prediction. Comprising the steps of collecting historical agricultural load and meteorological data; decomposing historical agricultural load and meteorological data by using multivariate variational mode decomposition to obtain cycle, trend and residual mode components; and respectively constructing a time convolutional neural network, a bidirectional gating cycle unit and a support vector regression model for the decomposed period, trend and residual modal component data set, and fully mining feature information of each mode after decomposition, thereby realizing accurate prediction of agricultural load. According to the method, the potential nonlinear space-time coupling relationship between the agricultural load and the meteorological factor is captured, the prediction effect in a seasonal periodic fluctuation scene of the agricultural load and a long-term trend and agricultural load abnormal scene is improved, the agricultural load prediction precision is improved, and a support is provided for reliable and stable operation of a power grid.
Owner:WUXI POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD

Multimodal interleaved image-text generative model based on dynamic feature synchronizer

The present disclosure relates to the technical field of multimodal learning, and particularly relates to a multimodal interleaved image-text generative model based on a dynamic feature synchronizer. An image encoder extracts multi-resolution multi-scale feature maps from input images of interleaved image-text data; and a dynamic feature synchronizer in a multimodal large language model acquires fine-grained information, so as to determine output feature data corresponding to the interleaved image-text data and then generate a target image and / or target text associated with the interleaved image-text data.
Owner:TSINGHUA UNIVERSITY

Load prediction method and system for dividing multiple tasks based on error, medium and equipment

PendingCN121809758Aavoid forgettingSimplify the multi-modal learning processForecastingNeural learning methodsData setLoad forecasting
The invention discloses a load prediction method and system based on error division multiple tasks, a medium and equipment, and the method comprises the steps: combining an error division task with a continuous learning mechanism, firstly calculating a sample relative error in a training process, classifying low-error data into a data set corresponding to a learned load feature mode, and carrying out the calculation of a data set corresponding to the learned load feature mode; the high-error data is delimited as a new task to be learned; measuring model parameter importance through a Fisher information matrix of EWC, introducing an L2 regularization item to constrain core parameter updating, and avoiding forgetting a learned old mode; and circularly executing task division and model training, so that a single model gradually masters various load characteristic modes in continuous learning. A division basis does not need to be set manually, a multi-mode learning process is simplified, and prediction accuracy and robustness in a complex load scene are improved.
Owner:GUOHUA ENERGY INVESTMENT +1

Method for predicting the electrical characteristics of semiconductor devices

This invention provides an electrical characteristic prediction method for predicting the electrical characteristics of semiconductor devices from a process list. [Solution] A method for predicting the electrical characteristics of a semiconductor device using a feature calculation unit and a characteristic prediction unit, wherein the feature calculation unit has a first learning model 210 and a second learning model 220, and the first learning model has the steps of learning a process list for producing a semiconductor device and generating a first feature. The second learning model has the steps of learning the electrical characteristics of a semiconductor device produced by the process list and generating a second feature. The characteristic prediction unit has a third learning model 230, and the third learning model has the steps of performing multimodal learning using the first feature and the second feature and outputting the values ​​of variables used in the calculation formula for semiconductor device characteristics. Furthermore, the first to third learning models each have different neural networks.
Owner:SEMICON ENERGY LAB CO LTD

Multi-modal learning method and device based on gradient polyhedron volume modulation and medium

The invention relates to the field of multi-modal learning in machine learning and computer vision, and discloses an unbalance multi-modal learning method and device based on gradient polyhedron volume modulation and a medium, and the method comprises the steps: optimizing the model parameters of a multi-modal model through iterative training; in each iteration, calculating the gradient of the total loss function relative to the model parameters through back propagation, and obtaining a single-mode gradient vector and a multi-mode gradient component vector of each mode; based on the single-mode gradient vector and the multi-mode gradient component vector, constructing a gradient matrix for each mode, and calculating an unbalance ratio representing an optimization unbalance degree between the modes; generating a gradient modulation weight based on the imbalance ratio, and modulating the gradient vector of each mode by using the modulation weight to obtain a modulated gradient; model parameters are updated by applying the modulated gradient so as to relieve optimization imbalance between modes and promote collaborative optimization of a multi-mode target and a single-mode target. According to the invention, an explainable and accurate quantification tool is provided for multi-modal imbalance.
Owner:STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST

An RGB-D saliency object detection device and method based on a depth-adaptive inverse thinning network architecture.

This invention discloses an RGB-D saliency object detection device and method based on a depth-adaptive inverse thinning network architecture. It employs a dual-stream architecture (RGB stream and depth stream) to extract RGB-related features and depth features respectively, and uses a cross-fusion method to achieve multimodal feature fusion. It determines whether the fused features are high-level features; if so, it obtains intermediate saliency detection results; otherwise, it classifies them as low-level features and waits for the generated intermediate saliency results to be processed by an inverse thinning module for inverse thinning of low-level features. Finally, it generates saliency prediction results. This invention addresses the problem of how to achieve cross-level multimodal learning and fully explore the correlation and complementarity between RGB images and depth maps.
Owner:HARBIN INST OF TECH

A multi-modal learning method, device and equipment under different perception sources and a medium

The application discloses a multimodal learning method and device under different perception sources, equipment and medium, relates to the technical field of machine learning, and utilizes a learnable temperature parameter to adjust the semantic energy score of different modes, and utilizes the semantic energy weight score to weight the interaction features. In this process, the semantic quality of each mode is dynamically evaluated through the energy score, noise modes are suppressed, key information is highlighted, a temperature adjustment and dynamic gating mechanism are introduced to eliminate the quality difference and semantic asymmetry between modes, and adaptive feature fusion is realized. Then, a gradient adjustment factor is obtained according to the confidence ratio, and a direction consistency adjustment factor is generated according to the cosine similarity. In this process, the gradient score perception represented by the gradient adjustment factor obtained based on the confidence ratio and the gradient direction alignment represented by the direction consistency adjustment factor generated according to the cosine similarity are used to adjust the mode gradient from two dimensions, so that the overall performance of the multimodal learning is finally improved.
Owner:INNER MONGOLIA UNIV OF TECH

Method and device for improving word learning efficiency of primary and secondary school students

The application relates to a method and device for improving the word learning efficiency of primary and secondary school students, and relates to the technical field of recommendation learning, which comprises the following steps: S1, defining a feature vector of a word and a word list of learning materials; S2, establishing the feature vector of all words in an English learning textbook, initializing a word recommendation value dictionary and performing learning material warehousing; S3, dynamically updating the feature vector of the word and the recommendation value dictionary according to the performance of students in different learning types in the learning process; and S4, screening out candidate learning materials, calculating the recommendation value of the candidate learning materials through the word dictionary of the candidate learning materials and the word recommendation value dictionary, and pushing the learning materials with the highest recommendation value to the students; the application solves the problems that traditional word learning lacks personalized dynamic recommendation and is difficult to adapt to multimodal learning scenes and individual memory decay rules, realizes long-term learning effect visualized closed-loop optimization through the combination of dynamic feature engineering and multi-scene recommendation, and effectively improves the word learning efficiency.
Owner:读书郎教育科技有限公司

A multi-modal balanced learning method based on label reshaping

This invention discloses a multimodal balanced learning method based on label reshaping. Starting from label-side design, this invention proposes for the first time to balance multimodal learning by reshaping the cross-modal label space. A reshaping matrix is ​​generated using the prediction output of complementary modalities to inject cross-modal inter-class information. Simultaneously, the reshaping intensity is adaptively calculated based on the confidence differences between modalities to control the degree of label space transformation, thereby balancing the mapping difficulty from the feature space to the label space for different modalities. Finally, a target parameter update strategy is employed, applying different optimization objectives to different parts of the model to ensure training stability. This method dynamically reshapes the label space corresponding to each sample and each modality, constructing a balanced soft supervision signal rich in inter-class relationships. This balances the learning difficulty of different modalities at the source and injects cross-modal inter-class information to enhance the model's discriminative ability.
Owner:SOUTHEAST UNIV

Thyroid nodule image classification method based on ultrasound and infrared multi-modal images

The application discloses a thyroid nodule image classification method based on ultrasound and infrared thermal image multimodal images. The application is based on ultrasound and infrared thermal images, combined with a multimodal learning method, and provides a self-adaptive multimodal hybrid model, which is composed of an intra-modal mixed encoder, a self-adaptive cross-modal encoder and a classification head. The intra-modal mixed encoder can model global features while extracting local features; the self-adaptive cross-modal encoder can formulate personalized modal weights according to different cases, and simultaneously perform information interaction and fusion of inter-modal features; and the classification head classifies the obtained fusion features. The method automatically classifies the thyroid nodule images of a subject based on ultrasound and infrared thermal images of the subject by using an AmmH model, and provides objective and accurate classification results for doctors, so that auxiliary diagnosis is realized.
Owner:WUHAN UNIV

Document image classification method and system based on multi-modal contrast learning framework, and readable storage medium

The invention discloses a document image classification method and system based on a multi-modal contrast learning framework and a readable storage medium, and relates to the field of computer vision and natural language processing. Comprising the following steps: constructing a training data set, extracting positive and negative sample pairs, sending the positive and negative sample pairs into a multi-modal large model, setting cue words for model training, and optimizing the model by supervising and comparing loss; respectively selecting examples from a plurality of categories as templates, respectively sending the templates into a model, and storing a plurality of output feature embedding vectors and corresponding categories into a template library; and sending a document image of which the category is to be confirmed into the model, outputting feature embedding vectors, reading the template, and taking the category corresponding to the feature embedding vector with the maximum similarity as a classification result. According to the method, algorithms such as multi-modal learning and comparative learning are combined, multi-source information such as images and texts is fused, discriminative features in the multi-source information are extracted, accurate classification of document images is achieved, and the method has the advantages of being universal, efficient and high in precision.
Owner:BEIJING YIDAO BOSHI TECH

Classification and identification method based on interference signals

The invention discloses a classification and identification method based on interference signals, which relates to the technical field of signal processing, and comprises the following steps: receiving interference signal data, and further obtaining a complexity evaluation value by using information entropy, fractal dimension and spectral characteristics; key parameters of the ant colony algorithm are dynamically configured based on the complexity evaluation result of the interference signal data features and the application scene information, and the key parameters comprise pheromone volatilization coefficients, the number of ant colonies and the step length; when the complexity evaluation value is lower than a data feature complexity threshold value, classification features are marked by directly utilizing dynamically configured ant colony algorithm key parameters, and when the complexity evaluation value is higher than or equal to the data feature complexity threshold value, the classification features are marked by utilizing multi-modal learning; classifying the interference signals based on the marked classification features; and the accuracy of the classification result is evaluated, and then key parameters of the ant colony algorithm are adjusted and optimized. Through evaluation of classification results and optimization of key parameters, the accuracy and adaptability of interference signal classification identification are improved.
Owner:CHINA ORDNANCE IND GRP NO 207 RES INST

A copper electrolytic short-circuit detection method and system based on modal enhancement and fusion mechanism

This invention discloses a method and system for detecting short circuits in copper electrolytic capacitors based on modal enhancement and fusion mechanisms, relating to the field of multimodal learning technology. The method includes: acquiring the RGB and HSV modal images of the copper electrolytic capacitor electrode plate to be tested; using the YOLO model framework as the target detection framework, setting two parallel branches in the backbone network and replacing the C3 module in the backbone network with a dynamic C3 module to construct an improved YOLO model framework; introducing an adaptive cross-modal fusion module and an enhanced perception fusion module in the modal fusion stage of the improved YOLO model framework to construct a multimodal detection model; and inputting the RGB and HSV modal images into the multimodal detection model to perform short circuit detection on the copper electrolytic capacitor electrode plate to be tested. This invention alleviates the technical problems of false detections and missed detections in existing technologies.
Owner:UNIV OF SCI & TECH BEIJING

Human pose estimation method based on multimodal attention network

This invention relates to the fields of wireless sensing and deep learning, specifically to a human pose estimation method based on a multimodal attention network. In video modality learning, a multi-resolution network (FCN) is designed to extract pose features. By applying a local self-attention network to the generated heatmap, the weights and attention levels of different regions are adaptively adjusted to capture more refined keypoint information. In CSI modality learning, a spatiotemporal attention network and a multimodal-guided linear spatial position variation layer are designed to facilitate self-learning of its spatiotemporal features. A teacher-student architecture is adopted in the overall multimodal learning network to allow the Wi-Fi signal-based pose estimation model to learn more from the deep learning capabilities of the video modality. By transferring the correct human pose estimation information from the video learning network to the Wi-Fi signal learning network, accurate human pose information can be estimated using only the CSI modality as input.
Owner:TIANJIN UNIV

Learning situation report generation method and system based on AI language model

The invention belongs to the technical field of learning condition management, and particularly relates to a learning condition report generation method and system based on an AI language model, and the system specifically comprises a multi-modal learning condition data aggregation module, a learning condition data noise reduction and regularization module, a knowledge point mastery degree decision module, a knowledge point enhancement necessity analysis module, a learning condition report AI generation and output module and a teacher decision end. Individual skilled and weak knowledge points of students are accurately distinguished through a knowledge point mastery degree decision-making module, a knowledge point reinforcement necessity analysis module provides a basis for a teacher to accurately position knowledge points needing key reinforcement and reasonably distribute teaching resources, and a learning enthusiasm quantification output module accurately quantifies learning attitudes of the students. The course teaching management evaluation module comprehensively evaluates the course teaching management condition, and the learning condition report AI generation and output module generates a two-dimensional learning condition report of the whole class and the individual students, thereby providing powerful technical support for personalized teaching and precise teaching management.
Owner:SICHUAN LETU ZHIXING TECH CO LTD

Image-text cross-modal retrieval method based on instance enhancement and causal intervention

The invention discloses an image-text cross-modal retrieval method based on instance enhancement and causal intervention, and belongs to the technical field of multi-modal learning. Aiming at the problem that a co-occurrence deviation cannot be effectively solved when an existing image-text cross-modal retrieval method is used for processing semantic alignment of image-text data, a series of preprocessing is carried out on paired image-text data, and a data set is divided; through an instance enhancement mechanism, in combination with context enhancement in instances and association interaction between the instances, feature expression is enhanced, and a semantic gap between modals is reduced; a causal component and an interference component are distinguished by adopting an embedded decoupling technology, and an anti-fact sample is constructed to optimize and improve the discrimination robustness of a model; and constructing an objective function optimization model containing semantic association stability loss, instance enhancement loss, embedding decoupling loss and anti-fact loss. According to the method, interference of background information in image-text data is effectively weakened, deep semantic association is accurately captured, and higher generalization and robustness are shown in scenes such as intelligent retrieval and content recommendation.
Owner:SHANXI UNIV

A dementia classification method based on knowledge-guided image diffusion reasoning

This invention relates to the field of multimodal data fusion and neurodegenerative disease classification, specifically a dementia classification method based on knowledge-guided image diffusion reasoning. First, initial features are extracted from magnetic resonance imaging (MRI) data using a 3D neural network, and a conditional diffusion model is used to generate category-enhanced image features. Simultaneously, a dementia knowledge graph is constructed from clinical metadata to integrate domain prior knowledge. Then, the enhanced image features and the knowledge graph embedding representation are input into a multimodal learning framework. This framework aligns and fuses information from the two modalities through bimodal consistency learning and modality-specific learning achieved through contrastive learning. Finally, the fused features are fed into a classifier for dementia classification. This invention provides a novel knowledge-guided deep learning framework for dementia classification, effectively overcoming the challenges of limited data and poor interpretability faced by traditional models, and significantly improving the accuracy and robustness of classification.
Owner:ZHEJIANG SCI-TECH UNIV

Missing modal multi-modal learning method based on koopman conditional progressive diffusion

A missing modality multimodal learning method based on Koopman conditional progressive diffusion is proposed, relating to computer vision, natural language processing, and multimodal information fusion. The method includes: a Koopman structured multimodal encoder that separates each modality representation into a shared subspace and a specific subspace through Koopman feature decoupling, and applies orthogonal constraints to ensure effective decomposition, suppressing modal bias at its source; a conditional fractional diffusion network that constructs multi-scale noise perturbations in the feature space based on variance amplification forward diffusion, combined with a Koopman structured conditioner to maintain stable conditional signals, and a time-varying fractional network performing inverse diffusion reconstruction, while introducing a Koopman guidance mechanism to maintain cross-modal linear consistency, and using residual channel self-attention to further refine the recovery results; and a Koopman-aware progressive training strategy that achieves phased optimization through sample difficulty measurement and dynamic weight scheduling, enabling the model to first establish a robust structure on simple samples and then improve generalization ability on complex samples.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A multi-modal data alignment and enhancement method based on semantic consistency

This invention belongs to the field of multimodal data processing technology and discloses a multimodal data alignment and enhancement method based on semantic consistency, aiming to solve the problems of semantic gap and insufficient generalization of multimodal data models. The method includes: acquiring original multimodal data on the same semantic topic, performing noise and anomaly processing and format normalization on the multimodal dataset; extracting semantic features of each modality from the multimodal dataset and normalizing and adapting them to construct a shared semantic space; achieving semantic consistency alignment between modal feature vectors of each modality through a cross-modal alignment mechanism; accordingly, performing intra-modal and cross-modal enhancement on each modality and its shared feature vectors, and integrating and deduplicating to obtain an aligned and enhanced multimodal dataset. This invention achieves accurate semantic alignment of multimodal data, improves data diversity, effectively enhances the generalization ability of multimodal learning models, and is applicable to artificial intelligence scenarios that rely on multimodal data collaboration, demonstrating strong practicality.
Owner:KUAIJI XINYUN (QINGDAO) TECHNOLOGY CO LTD

Assisted reproduction IVF outcome prediction model based on LSTM and LSTM-CNN fusion architecture and construction method

PendingCN121439251AMedical data miningEnsemble learningSperm morphologyFeature vector
The invention provides an assisted reproduction IVF outcome prediction model based on an LSTM and LSTM-CNN fusion architecture and a construction method, and is applied to the field of data processing. The method comprises the following steps: acquiring assisted reproduction clinical layering core data (including IVF period original data and divided into conventional fertilization IVF and ICSI fertilization IVF according to a fertilization mode), key parameters (including sperm motility and the like in conventional IVF and additional sperm morphological images in ICSI) and target patient information; preprocessing the data, carrying out feature engineering, and generating multi-modal feature vectors adaptive to the two fertilization modes; then, an optimization prediction model is constructed, and an initial model and performance evaluation information are obtained; iteratively optimizing the model, introducing a new module to improve the multi-modal learning ability, and generating an SOTA level model through multi-index evaluation; and finally, processing target patient information by using the model, and outputting a personalized prediction reference scheme and clinical application value analysis information.
Owner:PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY) +2

A noise-aware unbiased dynamic multimodal emotion recognition method

PendingCN122313590APattern recognitionLow noise
This invention discloses a noise-aware, unbiased dynamic multimodal emotion recognition method, belonging to the field of emotion recognition technology. The invention employs an unbiased dynamic multimodal learning framework. First, it maps the emotion representation of each modality to a Gaussian distribution, using the mean and variance of the distribution to decouple information content from noise. The noise intensity is then estimated based on the modality distribution variance, achieving accurate estimation of noise intensity under both extremely low and high noise conditions. Next, a modality dropout algorithm is used to quantify the inherent dependence of the output on each modality, and modality weights are calibrated accordingly to ensure the reasonable participation of each modality in emotion recognition and mitigate the double inhibition effect. Finally, a progressive learning strategy is adopted, dividing the training process into an emotion representation stabilization learning stage and a noise-aware modeling stage. While ensuring sufficient learning of the emotional semantic representation, modality quality and dependency modeling are gradually introduced, thereby improving training stability and convergence reliability.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A hate video detection method based on a multi-modal pattern library and a hierarchical complementary gating mechanism

This invention provides a hate video detection method based on a multimodal pattern library and a hierarchical complementary gating mechanism, belonging to the field of multimodal learning technology. The method includes: acquiring textual, visual, and audio features of a video through a multimodal feature extraction module and mapping them to a unified space; constructing a multimodal pattern library using a large language model, refining massive training samples into a compact and interpretable high-level semantic prototype set; injecting semantic priors through a pattern library retrieval and feature enhancement module; employing a hierarchical complementary gating mechanism to dynamically allocate modal weights with text as the semantic anchor, achieving selective fusion of audio and visual features; and finally training the model using a composite loss function to complete hate video detection. This invention solves the problems of modal competition, visual noise sensitivity, semantic ambiguity in instance-level retrieval, and high computational overhead caused by symmetric fusion in existing methods, significantly improving detection accuracy and efficiency, and enhancing model robustness and generalization ability.
Owner:XINJIANG UNIVERSITY

Intelligent course management system and method

This invention relates to the field of course management technology, and more particularly to a smart course management system and method, comprising: a data acquisition module for real-time acquisition of multimodal learning behavior data during the learning process; a teaching resource management module for storing and managing multi-source teaching resources and constructing a unified index map of teaching resources; a cognitive state triggering module for generating cognitive state modeling instructions; a cognitive state modeling module for constructing a dynamic knowledge graph and sending recommendation update instructions; an adaptive learning recommendation module for generating candidate knowledge point sequences based on the dynamic knowledge graph, filtering and matching resource types to generate recommendation results; and a negative feedback adaptive module for calculating the deviation rate between actual and expected learning gains, determining the source of deviation, and generating adjustment instructions. This invention improves the accuracy and adaptability of learning recommendations, and enhances course management efficiency and learning effectiveness.
Owner:HEBEI QISI EDUCATION TECH CO LTD

Information extraction traceability method and system based on native multi-modal large model, and readable storage medium

The invention discloses an information extraction traceability method and system based on a native multi-modal large model and a readable storage medium, and belongs to the technical field of artificial intelligence and multi-modal learning. Comprising the steps of determining a file representation form for a to-be-processed original image, performing coordinate labeling by adopting a mantissa outward rounding mode, and constructing a space labeling data set; inputting an image by using the spatial annotation data set and adopting a dynamic resolution input strategy, and training a multi-modal large language model; inputting a sample image to the trained multi-modal large language model, outputting text content and coordinate information of the sample image, and performing coordinate post-processing on the coordinate information to realize edge compensation; and outputting the text content and the post-processed coordinate information. According to the method, the problems of coordinate dislocation, information loss and the like in a traditional OCR + language model two-stage architecture can be solved, and meanwhile, the defects of precision deviation and non-uniform format of an existing MLLM in space coordinate generation are overcome.
Owner:BEIJING YIDAO BOSHI TECH

Artificial intelligence-based teaching interactive content intelligent generation method

This invention relates to the field of artificial intelligence education and teaching technology, specifically to an intelligent method for generating interactive teaching content based on artificial intelligence. The method includes: receiving a personalized learning profile of a target learner containing a sequence of knowledge mastery, learning behavior patterns, and cognitive ability tags; generating a dynamic learner state vector through multimodal learning feature extraction and fusion; and constructing a dynamic knowledge subgraph by selecting related knowledge nodes from a pre-defined knowledge graph based on this vector. The dynamic learner state vector and the dynamic knowledge subgraph are input into a content generation engine, which simultaneously generates interconnected explanatory texts, interactive questions, and contextualized cases. The internal logical consistency of these three types of content is checked, and the difficulty is adjusted. After integration, a complete interactive teaching content unit is formed. This method generates related teaching content based on the learner's dynamic state and dynamic knowledge structure, optimizing the content logic and difficulty suitability.
Owner:SHANGHAI YICAO YIMU EDUCATION TECHNOLOGY CO LTD

Skill micro certificate comprehensive evaluation system fusing multi-modal learning results

PendingCN121436786AInstrumentsMindsetEngineering
The invention discloses a skill micro certificate comprehensive evaluation system fused with multi-modal learning achievements, and relates to the technical field of comprehensive evaluation systems, the skill micro certificate comprehensive evaluation system comprises a text module, an audio and video module and an emotion module, the text module judges answers written by trainees through a homework unit, and a log unit in the text module judges answers written by trainees from answering time, the longest number of input characters and the maximum number of input characters. The audio and video module judges behaviors of the trainees in an examination and brings illegal operations into scores, so that bad behaviors of the trainees in an examination room can be directly displayed by the scores, cheating behaviors of the trainees can be monitored at the same time, the method is more friendly to private enterprises with few faculty and staff, and the experience of the students is improved. The system can better save the labor cost, records the psychological state of the student during answering through the arranged emotion module, can finally evaluate the psychological state of the student during answering, and judges the psychological state of the student in the actual operation of subsequent work, thereby more comprehensively judging the ability of the student in multiple aspects.
Owner:HEBEI VOCATIONAL & TECH UNIV OF SCI & TECH

Multimodal learning behavior analysis method, system, and storage medium

This invention discloses a method, system, and storage medium for multimodal learning behavior analysis, applied in the field of artificial intelligence technology. It enables high-performance and highly interpretable multimodal learning behavior analysis, providing reliable analytical explanations for educational decision-making processes. The method includes: acquiring multimodal behavior data of the learning process of the object to be analyzed and preprocessing it to obtain multimodal sequence data; performing collaborative embedding representation on the multimodal sequence data to obtain initial feature representations; constructing a learning behavior data association graph based on the initial feature representations; decoupling the learning behavior data association graph using a graph decoupling neural network to obtain a multimodal decoupling graph; constructing an attribute routing mechanism based on the relationships between nodes in the multimodal decoupling graph; updating the node embedding representations and graph structure through the attribute routing mechanism to obtain target feature representations and target graph structures; performing learning behavior analysis based on the target feature representations and visualizing the target graph structure to obtain visualized learning behavior analysis results.
Owner:ZHEJIANG NORMAL UNIV

Distributed energy cluster state perception method and system fusing multi-modal learning

PendingCN122332877ANew energySimulation
This invention provides a distributed energy cluster state perception method and system that integrates multimodal learning, belonging to the field of state perception technology. It generates a dynamic operating condition baseline range in real time through multimodal fusion features. The judgment boundary can be synchronously and adaptively updated according to equipment operating conditions, environmental conditions, load fluctuations, and changes in renewable energy output. This effectively matches the highly time-varying, highly fluctuating, and nonlinear dynamic operating characteristics of distributed energy clusters, eliminating misjudgments and omissions caused by fixed thresholds, and significantly improving the matching accuracy across all operating conditions. It eliminates the perception lag in scenarios of sudden load changes, achieving millisecond-level real-time state perception. It uses a transient operating condition perception operator to capture transient operating condition features, eliminating the need for the traditional sliding window data accumulation process and completely eliminating perception delay. It can accurately capture transient operating condition changes such as sudden load changes, sudden changes in renewable energy output, and equipment start-up and shutdown, fully meeting the engineering requirements for real-time scheduling and rapid fault early warning of distributed energy clusters.
Owner:NANJING CORNERSTONE DATA TECH CO LTD

Portable multimodal learning analytics smart glasses

A portable multimodal learning analytics smart glasses, which can monitor, analyze and feedback the multimodal data, including expressions, voice, physiology, eye and head movements, and the data analysis result in real time during the learning process of the learner. The chip of the smart glasses integrates real-time data monitoring function, multimodal data analysis function and data visualization function. Through the data monitoring function, the variation conditions of expressions, voice, physiology, eyeball movement and head movements of the learning user during the learning process can be obtained in real time; through the multimodal data analysis function, the real-time acquired data is stored in the preset data structure for multimodal learning analysis; through the data visualization function, the processing result of data analysis is displayed to the learning user in the form of a visual graphic.
Owner:ZHEJIANG UNIV