Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

120 results about "Multimodal learning" patented technology

The information in real world usually comes as different modalities. For example, images are usually associated with tags and text explanations; texts contain images to more clearly express the main idea of the article. Different modalities are characterized by very different statistical properties. For instance, images are usually represented as pixel intensities or outputs of feature extractors, while texts are represented as discrete word count vectors. Due to the distinct statistical properties of different information resources, it is very important to discover the relationship between different modalities. Multimodal learning is a good model to represent the joint representations of different modalities. The multimodal learning model is also capable to fill missing modality given the observed ones. The multimodal learning model combines two deep Boltzmann machines each corresponds to one modality. An additional hidden layer is placed on top of the two Boltzmann Machines to give the joint representation.

Decision optimization method fusing enhanced multi-modal learning and knowledge graph

The invention discloses a decision optimization method fusing enhanced multi-modal learning and a knowledge graph, and relates to the technical field of artificial intelligence and knowledge graphs. According to the decision optimization method for fusing enhanced multi-modal learning and the knowledge graph, dynamic fusion of multi-modal features is realized through a dynamic weight adjustment and semantic alignment constraint mode, semantic precision and interpretability are improved, the semantic deviation problem caused by traditional static feature fusion is effectively solved, and the method has the advantages of being high in robustness and high in reliability. And by recording an intelligent reasoning path selected by a hierarchical reinforcement learning agent, interactive graph structure display can be carried out, the advantage of transparency is achieved, meanwhile, the interpretability of the path is further improved in cooperation with a multi-target reward function, intelligent updating of the knowledge graph is carried out in cooperation with comprehensive confidence, intelligent growth of the knowledge graph is achieved, and the intellectual property of the knowledge graph is improved. And a fine-grained interpretable report is generated through an adversarial training mechanism and anti-factual reasoning, so that the false alarm rate of an output result is further reduced, and the decision transparency is improved.
Owner:BEIJING SHANGCHENG ZHIYIN ROBOT TECHNOLOGY CO LTD

Online resource adaptive recommendation method for multi-modal learning behavior analysis

The invention discloses an online resource adaptive recommendation method based on multi-modal learning behavior analysis, and relates to the technical field of resource recommendation. The method comprises the following steps: firstly, dynamically collecting multi-modal data by using a heterogeneous sensor array, and carrying out noise reduction, probability distribution matching normalization and time-space alignment preprocessing; features are extracted through a hierarchical network, modeling learning behaviors such as a variational auto-encoder are combined, and the learning state is evaluated from multiple dimensions; recommendation decisions are generated based on reinforcement learning, recommendation is optimized in combination with personalized presentation and multi-source feedback analysis, meanwhile, the system has the functions of dynamic strategy adjustment, intelligent resource creation, cross-scene migration recommendation and the like, and accurate self-adaptive recommendation is achieved. According to the method, multi-modal data are comprehensively collected and deeply processed, learning behaviors and evaluation states are accurately analyzed, personalized resource recommendation is provided through intelligent recommendation and dynamic optimization strategies, recommendation accuracy and learning effects can be improved, user experience can be enhanced, and the utilization rate and competitiveness of platform resources can be improved.
Owner:SHANDONG LENSI EDUCATION TECH (GRP) CO LTD

Unified multi-modal alignment method and system based on hybrid expert and low-rank adaptation

The invention discloses a unified multi-modal alignment method and system based on mixed experts and low-rank adaptation, and belongs to the field of artificial intelligence and multi-modal learning. According to the method, firstly, a CLIP model serves as a basic model, and a unified modal encoder is constructed in combination with a modal perception hybrid expert strategy and a low-rank adaptive strategy; the unified modal encoder comprises an anchor modal marker corresponding to each anchor modal, an extended modal marker corresponding to each extended modal and a plurality of stacked multi-modal hybrid expert modules; then, on the basis of universal modal alignment measurement and knowledge distillation and cross-modal optimization strategies, multi-modal data samples are sampled in batches, and fine tuning training is conducted on the unified modal encoder. Through a single encoder and single training, alignment of any number of anchor modes and extension modes is realized, a special model does not need to be trained for each mode independently, the transferability of cross-domain and downstream tasks can be improved, and the method is suitable for tasks such as multi-mode understanding, generation and reasoning.
Owner:ZHEJIANG UNIV

Accurate learning data mining method based on cognitive calculation driving

PendingCN120523850AData processing applicationsRelational databasesCognitive intervention strategiesBehavioral data
The invention provides a learning data accurate mining method based on cognitive calculation driving. The learning data accurate mining method comprises the following steps of S1, performing multi-modal learning behavior data acquisition and heterogeneous integration; s2, a dynamic feature weight optimization step based on calculus; s3, performing cognitive state differential equation modeling; s4, a cognitive diagnosis hybrid model based on statistics; s5, incremental construction of the dynamic knowledge graph is carried out; s6, constructing a federated learning framework for privacy protection; s7, a cognitive intervention strategy is generated; s8, constructing a multi-granularity effect evaluation system; s9, a step of constructing an interpretability enhancement module; s10, a step of carrying out adaptive iterative optimization; the learning data accurate mining method based on cognitive calculation driving has the following advantages that the data utilization rate breaks through the limitation of a traditional method through federated learning and heterogeneous graph fusion; the attention prediction error is reduced through differential equation modeling, and the method is superior to all existing ARIMA / LSTM baseline models.
Owner:XINHUA WINSHARE PUBLISHING & MEDIA CO LTD

Multi-modal learning resource intelligent recommendation method and system based on AI large model

The invention relates to the technical field of computers, and provides a multi-modal learning resource intelligent recommendation method and system based on an AI large model, and the method comprises the steps: carrying out the intention analysis of interaction behavior data and user information based on the AI large model, and obtaining a user feature vector; obtaining multi-modal resource data based on the user feature vector, and integrating the multi-modal resource data to obtain a resource representation vector; mining implicit association between resources and knowledge points and a mapping relation between user requirements and target skills based on a knowledge graph in combination with user feature vectors and resource representation vectors to obtain a knowledge network; generating candidate recommended learning resources based on the context awareness information and the knowledge network; and performing sorting optimization on the candidate recommendation learning resources based on the user feature vector and the resource representation vector in combination with the current stage index of the multi-modal resource data, and generating an initial resource recommendation list. According to the embodiment of the invention, the content features of the learning resources are comprehensively captured, and the resource recommendation accuracy is improved.
Owner:SHENZHEN JINLONGFENG TECH CO LTD

Open environment-oriented missing modal gamma collaborative retrieval diffusion method

The invention discloses a missing mode gamma collaborative retrieval diffusion method for an open environment, and belongs to the field of multi-mode learning and missing mode processing. According to the method, a human brain multi-source context completion mechanism is simulated, and robust multi-modal learning is realized through three innovative modules: a context retrieval enhancement module: a multi-modal memory library is constructed, related instances are retrieved through similarity calculation under a gating mechanism, and context representation of a missing mode is enhanced; the prompt drive diffusion generation module is used for constructing a semantic prompt based on a retrieval result, fusing a de-noising diffusion probability model through an attention mechanism, and realizing context-aware knowledge migration and missing modal generation; and the inverse gamma noise optimization module is used for establishing a mixed normal-inverse gamma distribution model, dynamically sensing noise, realizing uncertainty estimation in multi-modal fusion and ensuring robustness and reliability of a regression result. According to the method, the dependence of the model on the available modal quality is effectively reduced, and the cross-modal knowledge migration effect and the multi-modal learning task performance are improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Ecological environment prediction system and method

The invention relates to the technical field of ecological environment prediction, in particular to an ecological environment prediction system and method, and the system and method employ artificial intelligence technologies such as multi-source data fusion, dynamic space-time modeling, cross-modal collaborative optimization, space-time deep learning, multi-modal learning, adaptive optimization, and prediction branches employing a Transform architecture neural network model. And the comprehensiveness, precision and practicability of ecological environment prediction are remarkably improved.
Owner:JIANGSU HANSHENG MEASUREMENT & CONTROL TECHNOLOGY CO LTD

Emotion conflict relieving method and system based on expert module and attention redistribution

The invention relates to the technical field of multi-modal learning, and discloses an emotion conflict relieving method and system based on an expert module and attention redistribution. The method comprises the steps that in the training process, multi-modal tokens are input into a large language model, and in each middle layer, modal expression is enhanced layer by layer through multi-modal specific experts: the tokens of different modals are distributed to corresponding experts through intra-modal routing; calculating token average representation of different modes through inter-mode routing dynamic fusion expert output; in the reasoning process, inputting the multi-modal emotion sample into the large language model, and performing reasoning attention redistribution in each middle layer of the large language model; and mapping the output of the last layer of the large language model through a linear classification layer, and predicting to obtain the emotion category of the multi-modal emotion sample. According to the method, a complex emotion conflict scene can be processed, and the modal prejudice of a multi-modal large language model is reduced, so that a more robust emotion reasoning and recognition effect is achieved.
Owner:UNIV OF SCI & TECH OF CHINA

Multi-modal molecular representation learning method for predicting permeability of cyclic peptide

The invention relates to the field of computer-aided drug design (CADD) and molecular informatics, in particular to a multi-modal representation learning method based on cyclopeptide molecules, which is used for predicting cell membrane permeability of cyclopeptide. The method mainly comprises the following steps: (1) data collection: integrating cyclic peptide permeability data from a ChEMBL database, a CycPeptMPDB database and a CyclicPepeda database and patent literatures; (2) multi-modal learning: for different modal data, a deep learning model is adopted to extract feature representations of the data; the method comprises the following steps of: encoding an SMILES sequence by using ChemBERTa (ChemBERTa); using Vision Transform to extract molecular image features, and learning a molecular image structure and 3D coordinate information based on GNN; (3) multi-modal feature fusion: adopting a self-adaptive extensible fusion mechanism, integrating SMILES feature information into image, graph and 3D coordinate features through a cross-modal feature fusion mechanism, and splicing all modal features to obtain multi-modal molecular representation; and (4) permeability prediction: sending the multi-modal molecular representation into a full connection layer for regression prediction so as to evaluate the permeability of the cyclopeptide.
Owner:HUNAN UNIV

Context modal completion multi-modal learning method based on semantic matching

The invention discloses a context modal completion multi-modal learning method based on semantic matching, and relates to the technical field of multi-modal learning, and the method comprises the following steps: for a multi-modal sample with modal damage, matching a semantic-related complete sample from a database based on reserved intact modal data; according to modal data of the semantic related complete sample, combining semantic information of the remaining intact modals to generate complementation representation of the damaged modals; inputting the complemented modal data and the remaining intact modal data into a multi-modal fusion model to generate a prediction result; and the performance stability under the modal damage condition is improved by optimizing model parameters through the combination of the task-related loss and the complementation loss. According to the method, through an innovative framework of semantic matching and context modal completion, the modal damage problem in multi-modal learning is effectively solved, and the performance and stability of the multi-modal fusion model under the modal damage condition are remarkably improved.
Owner:郑州埃文科技有限公司

Multi-modal English learning interaction system and vocabulary memory training method

The invention discloses a multi-modal English learning interaction system and a vocabulary memory training method, and relates to the technical field of English learning, the system comprises the following components: a data acquisition module, a data analysis module, a strategy adjustment module, a resource push module and an interaction learning module; multi-modal learning behavior data, including text input, voice reading, handwritten notes, video learning behaviors, interactive operation and the like, of learners are collected through the data acquisition module, the learners are subjected to group division by applying a group intelligent algorithm, and the behavior pattern and performance of each group in vocabulary learning are analyzed for each group, so that the learning efficiency of the learners is improved. Based on the analysis, the system can automatically adjust teaching strategies and push customized multi-modal learning resources and training methods, so that personalized requirements of different learners are met, and the learning effect and experience are remarkably improved.
Owner:XINXIANG VOCATIONAL & TECHN COLLEGE

Agricultural load prediction method and system based on multivariate time sequence decoupling multi-modal learning

The invention discloses an agricultural load prediction method and system based on multivariate time sequence decoupling multi-modal learning, and belongs to the technical field of agricultural load prediction. Comprising the steps of collecting historical agricultural load and meteorological data; decomposing historical agricultural load and meteorological data by using multivariate variational mode decomposition to obtain cycle, trend and residual mode components; and respectively constructing a time convolutional neural network, a bidirectional gating cycle unit and a support vector regression model for the decomposed period, trend and residual modal component data set, and fully mining feature information of each mode after decomposition, thereby realizing accurate prediction of agricultural load. According to the method, the potential nonlinear space-time coupling relationship between the agricultural load and the meteorological factor is captured, the prediction effect in a seasonal periodic fluctuation scene of the agricultural load and a long-term trend and agricultural load abnormal scene is improved, the agricultural load prediction precision is improved, and a support is provided for reliable and stable operation of a power grid.
Owner:WUXI POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD

Lithium battery K value real-time prediction method based on neural symbol reasoning and multi-modal learning

The invention provides a lithium battery K value real-time prediction method based on neural symbol reasoning and multi-modal learning, and the method specifically comprises the steps: S1, carrying out the coding of an industrial signal through a binary coding mode, converting an input voltage sequence into a binary sequence, and extracting a frequency domain feature through a wavelet transformation technology; s2, using a neural symbol inference engine to realize verifiable feature selection, and executing feature selection under logic constraints; s3, by means of a multi-head potential attention mechanism fusion process knowledge graph, a potential space projection matrix is constructed, and the number of attention heads is adjusted through a dynamic head number adjusting mechanism; and S4, completing online knowledge migration by utilizing a dynamic distillation expert system, and realizing knowledge transfer and model optimization by combining a hybrid expert architecture and an online distillation technology and applying an expert dynamic activation function and knowledge distillation loss. According to the method, advanced technologies such as neural symbol reasoning and multi-modal learning are fused, the prediction precision is high, the response delay is small, the energy consumption ratio is low, and the interpretability score is high.
Owner:GUANGDONG YIZHILIAN TECHNOLOGY CO LTD

Multimodal interleaved image-text generative model based on dynamic feature synchronizer

The present disclosure relates to the technical field of multimodal learning, and particularly relates to a multimodal interleaved image-text generative model based on a dynamic feature synchronizer. An image encoder extracts multi-resolution multi-scale feature maps from input images of interleaved image-text data; and a dynamic feature synchronizer in a multimodal large language model acquires fine-grained information, so as to determine output feature data corresponding to the interleaved image-text data and then generate a target image and / or target text associated with the interleaved image-text data.
Owner:TSINGHUA UNIVERSITY

Study performance precise teaching management method based on time sequence behavior modeling

The invention provides a learning performance precise teaching management method based on time sequence behavior modeling. The learning performance precise teaching management method comprises the following steps: S1, carrying out multi-modal learning behavior data acquisition; s2, aligning and segmenting the time sequence data; s3, performing knowledge retention rate differential modeling; s4, carrying out time sequence feature extraction; s5, constructing a mixed time sequence model; s6, performing dynamic learning ability evaluation; s7, carrying out adaptive resource recommendation; s8, carrying out cross-correction knowledge diffusion optimization; the method has the following advantages: the method is real-time and accurate; the data is acquired to acquire the intervened closed-loop delay lt; compared with the traditional method, the time is increased by 40 times. And cost optimization: the workload of teachers is reduced by 58% and the resource purchase cost is reduced by 42% through an automation strategy. The scale effect is that federal learning supports ten-thousand-person-level concurrence, and the model updating period is shortened to the hour level from the quarter level. Education fairness: cross-school knowledge diffusion enables the superior rate of weak schools to be improved by 29%.
Owner:XINHUA WINSHARE PUBLISHING & MEDIA CO LTD

A garlic seed sprouting method based on multimodal learning network

The present application discloses a garlic seed sprouting correction method based on a multimodal learning network, which relates to the field of garlic sowing technology, including: capturing dynamic video frames of garlic seeds before sprouting through an image acquisition device; characterizing and identifying the garlic seeds in the dynamic video frames and judging the orientation of the garlic seed bulbs through a multimodal learning fusion network that integrates knowledge graphs and computer vision; controlling a sprout correction device to correct the sprouts of the garlic seeds based on the judged orientation of the garlic seed bulbs, ensuring that the garlic seeds can stand upright in the soil with the bulbs upward when mechanically inserted. The captured dynamic video frames of the garlic seeds are judged through a multimodal learning fusion network to judge the orientation of the garlic seed bulbs, and finally, the orientation of the garlic seeds is adjusted using the sprout correction device based on the judgment result, ensuring that the garlic seeds can stand upright in the soil with the bulbs upward.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Load prediction method and system for dividing multiple tasks based on error, medium and equipment

PendingCN121809758Aavoid forgettingSimplify the multi-modal learning processForecastingNeural learning methodsData setLoad forecasting
The invention discloses a load prediction method and system based on error division multiple tasks, a medium and equipment, and the method comprises the steps: combining an error division task with a continuous learning mechanism, firstly calculating a sample relative error in a training process, classifying low-error data into a data set corresponding to a learned load feature mode, and carrying out the calculation of a data set corresponding to the learned load feature mode; the high-error data is delimited as a new task to be learned; measuring model parameter importance through a Fisher information matrix of EWC, introducing an L2 regularization item to constrain core parameter updating, and avoiding forgetting a learned old mode; and circularly executing task division and model training, so that a single model gradually masters various load characteristic modes in continuous learning. A division basis does not need to be set manually, a multi-mode learning process is simplified, and prediction accuracy and robustness in a complex load scene are improved.
Owner:GUOHUA ENERGY INVESTMENT +1

A Multimodal Learning Automatic Complaint Content Analysis and Classification Method and System

The present invention discloses a multi-modal learning automated complaint content analysis and classification method and system, which relates to the technical field of semantic analysis, and includes: collecting multi-modal complaint data and preprocessing the collected multi-modal complaint data; extracting the features of the preprocessed multi-modal complaint data and fusing the features of different modal data; using the fused features to construct a complaint content classification model, performing model training, evaluating and optimizing the trained model; and deploying the optimized model for real-time classification. The multi-modal learning automated complaint content analysis and classification method provided by the present invention classifies complaint content quickly and accurately, and proposes targeted processing strategies according to the main categories, emotional tendencies and specific topics of complaints. Automatically trigger an emergency response mechanism to ensure that problems are processed immediately and ensure the effective utilization of resources. Learn from historical data, continuously optimize decision rules, and improve the intelligence level of processing strategies and customer satisfaction.
Owner:JIANGSU HUCHUAN TECH CO LTD

Adaptive differential privacy budget adjustment method for multi-modal learning

The invention discloses a multi-modal learning-oriented adaptive differential privacy budget adjustment method, which dynamically allocates a differential privacy budget in a training process by evaluating the feature distribution difference, task contribution degree and noise sensitivity of multi-modal data, and specifically comprises the following steps of: extracting a multi-modal feature vector and calculating modal correlation; fusing features and predicting downstream task results; determining a contribution ratio and a sensitivity difference ratio based on the prediction accuracy and the noise sensitivity; the privacy budget of each round is finely adjusted in combination with the correlation, so that the high-contribution mode distributes the large budget to guarantee the performance, and the low-contribution mode distributes the small budget to strengthen privacy protection; gaussian noise is injected according to a budget in gradient back propagation.
Owner:GUANGZHOU UNIVERSITY

Method and device for improving word learning efficiency of middle and primary school students

A method and device for improving word learning efficiency of middle and primary school students relates to the technical field of recommendation learning, and comprises the following steps: S1, defining feature vectors of words and a word list of learning materials; s2, establishing feature vectors of all words in the English learning textbook, initializing a word recommendation value dictionary and performing learning material storage; s3, dynamically updating the word feature vectors and the recommendation value dictionary according to the expressions of the students in different learning types in the learning process; s4, screening out alternative learning materials, calculating recommendation values of the alternative learning materials through a word dictionary and a word recommendation value dictionary of the alternative learning materials, and pushing the learning material with the highest recommendation value to the student; the problem that traditional word learning lacks personalized dynamic recommendation and is difficult to adapt to multi-modal learning scenes and individual memory attenuation laws is solved, dynamic feature engineering is combined with multi-scene recommendation, visual closed-loop optimization of long-term learning effects is realized, and word learning efficiency is effectively improved.
Owner:读书郎教育科技有限公司

A self-evolutionary evaluation system for teaching effectiveness of educational agents

ActiveCN120317721BForecastingBiological modelsFeature vectorEvolutionary learning
The present invention discloses a self-evolutionary evaluation system for the teaching effectiveness of an educational agent, which relates to the field of educational artificial intelligence. The system includes a feature fusion module, a model construction module, a trigger judgment module, a self-evolution module, and a mapping graph update module. The present invention utilizes a multimodal learning algorithm for unified coding and feature extraction, thereby constructing a multidimensional feature vector that comprehensively reflects the cognitive and emotional states of students, thereby improving the accuracy and scientific nature of teaching effectiveness evaluation. By introducing an evolutionary learning mechanism, through mutation, crossover, and screening operations on inefficient or fluctuating strategies, combined with a simulation environment for strategy pre-evaluation, the teaching strategy is continuously optimized, thereby improving the strategy adaptability and continuous teaching optimization capabilities of the educational agent.
Owner:MIANYANG TEACHERS COLLEGE

Method for predicting the electrical characteristics of semiconductor devices

This invention provides an electrical characteristic prediction method for predicting the electrical characteristics of semiconductor devices from a process list. [Solution] A method for predicting the electrical characteristics of a semiconductor device using a feature calculation unit and a characteristic prediction unit, wherein the feature calculation unit has a first learning model 210 and a second learning model 220, and the first learning model has the steps of learning a process list for producing a semiconductor device and generating a first feature. The second learning model has the steps of learning the electrical characteristics of a semiconductor device produced by the process list and generating a second feature. The characteristic prediction unit has a third learning model 230, and the third learning model has the steps of performing multimodal learning using the first feature and the second feature and outputting the values ​​of variables used in the calculation formula for semiconductor device characteristics. Furthermore, the first to third learning models each have different neural networks.
Owner:SEMICON ENERGY LAB CO LTD

Dual-modal learning slope risk detection method integrating laser ranging and monitoring images

The present invention discloses a dual-modal learning slope risk detection method that integrates laser ranging and surveillance images, and relates to the technical field of slope safety monitoring. The present invention provides a dual-modal learning slope risk detection method that integrates laser ranging and surveillance images. The method first combines surveillance camera image data with three-dimensional position data collected by a single, single-point laser rangefinder to fuse image features with laser ranging features. Laser ranging supplements the third-dimensional features, significantly improving the three-dimensional perception of slopes, providing greater information and stronger recognition capabilities. The method then constructs a dual-modal network and performs multimodal learning to detect the type and area of ​​slope risks, particularly local risk changes. This method provides accurate, real-time early warning services for slope detection systems, reducing the probability of false alarms.
Owner:FUJIAN HUICHUAN DIGITAL TECH

Multi-modal learning method and device based on gradient polyhedron volume modulation and medium

The invention relates to the field of multi-modal learning in machine learning and computer vision, and discloses an unbalance multi-modal learning method and device based on gradient polyhedron volume modulation and a medium, and the method comprises the steps: optimizing the model parameters of a multi-modal model through iterative training; in each iteration, calculating the gradient of the total loss function relative to the model parameters through back propagation, and obtaining a single-mode gradient vector and a multi-mode gradient component vector of each mode; based on the single-mode gradient vector and the multi-mode gradient component vector, constructing a gradient matrix for each mode, and calculating an unbalance ratio representing an optimization unbalance degree between the modes; generating a gradient modulation weight based on the imbalance ratio, and modulating the gradient vector of each mode by using the modulation weight to obtain a modulated gradient; model parameters are updated by applying the modulated gradient so as to relieve optimization imbalance between modes and promote collaborative optimization of a multi-mode target and a single-mode target. According to the invention, an explainable and accurate quantification tool is provided for multi-modal imbalance.
Owner:STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST

SAM2 visual basis model assisted vehicle-mounted laser point cloud semantic segmentation method and system

The invention relates to the technical field of point cloud processing, in particular to a vehicle-mounted laser point cloud semantic segmentation method and system assisted by an SAM2 visual basis model, and the method comprises the steps: constructing a semantic segmentation multi-mode learning model for learning complementary information from different modes through employing an SAM2 model and a point cloud network; a training sample pair is utilized to perform joint training on the multi-modal learning model, and the point cloud network after joint training is used as a point cloud semantic segmentation target model, extracting a target segmentation mask by using an SAM2 model, and constraining the point cloud network to predict the semantic consistency of the point cloud by using a semantic score of the target segmentation mask based on a semantic classifier, so as to integrate image color and texture knowledge into point cloud network training through knowledge distillation in modal interaction; and inputting the to-be-processed vehicle-mounted laser point cloud into the point cloud semantic segmentation target model so as to perform semantic prediction output on the to-be-processed vehicle-mounted laser point cloud. According to the method, the semantic segmentation performance of the point cloud can be improved, and a multi-modal interaction gain effect can be obtained.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Multimodal learning resource intelligent recommendation method and system based on AI big model

The present invention relates to the field of computer technology and provides a method and system for intelligently recommending multimodal learning resources based on an AI large model. The method comprises: performing intent analysis on interactive behavior data and user information based on the AI ​​large model to obtain a user feature vector; acquiring multimodal resource data based on the user feature vector, and obtaining a resource representation vector based on the integration of the multimodal resource data; mining the implicit associations between resources and knowledge points and the mapping relationship between user needs and target skills based on a knowledge graph combined with the user feature vector and the resource representation vector to obtain a knowledge network; generating candidate recommended learning resources based on contextual perception information and the knowledge network; and ranking and optimizing the candidate recommended learning resources based on the user feature vector and the resource representation vector combined with the current stage indicators of the multimodal resource data to generate an initial resource recommendation list. The embodiments of the present invention achieve comprehensive capture of the content features of learning resources and improve the accuracy of resource recommendations.
Owner:SHENZHEN JINLONGFENG TECH CO LTD

An RGB-D saliency object detection device and method based on a depth-adaptive inverse thinning network architecture.

This invention discloses an RGB-D saliency object detection device and method based on a depth-adaptive inverse thinning network architecture. It employs a dual-stream architecture (RGB stream and depth stream) to extract RGB-related features and depth features respectively, and uses a cross-fusion method to achieve multimodal feature fusion. It determines whether the fused features are high-level features; if so, it obtains intermediate saliency detection results; otherwise, it classifies them as low-level features and waits for the generated intermediate saliency results to be processed by an inverse thinning module for inverse thinning of low-level features. Finally, it generates saliency prediction results. This invention addresses the problem of how to achieve cross-level multimodal learning and fully explore the correlation and complementarity between RGB images and depth maps.
Owner:HARBIN INST OF TECH

Group insurance policy processing method and device based on multi-modal learning, equipment and medium

The invention relates to the technical field of data processing, and discloses a group insurance policy processing method and device based on multi-modal learning, equipment and a medium. According to the scheme, a first semantic embedding feature of a target field is extracted from a group insurance policy text and image through a multi-modal large model; the second semantic embedding feature is extracted from the system filling text and the target field in the group insurance policy image, and the text and image information is fully fused, so that the key information in the group insurance policy can be captured more comprehensively and accurately. The feature similarity of the first semantic embedding feature and the second semantic embedding feature is calculated through a cross-modal semantic alignment algorithm, so that the semantic difference is recognized more accurately. The difference detection result is generated according to the semantic similarity, the problem field in the group insurance policy is positioned, the auditing efficiency and accuracy are improved, and potential risks are avoided. According to the scheme, claim settlement disputes caused by information errors of group insurance policies in the financial field and the medical field can be reduced, and the insurance service efficiency and quality are improved.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

A multi-modal learning method, device and equipment under different perception sources and a medium

The application discloses a multimodal learning method and device under different perception sources, equipment and medium, relates to the technical field of machine learning, and utilizes a learnable temperature parameter to adjust the semantic energy score of different modes, and utilizes the semantic energy weight score to weight the interaction features. In this process, the semantic quality of each mode is dynamically evaluated through the energy score, noise modes are suppressed, key information is highlighted, a temperature adjustment and dynamic gating mechanism are introduced to eliminate the quality difference and semantic asymmetry between modes, and adaptive feature fusion is realized. Then, a gradient adjustment factor is obtained according to the confidence ratio, and a direction consistency adjustment factor is generated according to the cosine similarity. In this process, the gradient score perception represented by the gradient adjustment factor obtained based on the confidence ratio and the gradient direction alignment represented by the direction consistency adjustment factor generated according to the cosine similarity are used to adjust the mode gradient from two dimensions, so that the overall performance of the multimodal learning is finally improved.
Owner:INNER MONGOLIA UNIV OF TECH

Method and device for improving word learning efficiency of primary and secondary school students

The application relates to a method and device for improving the word learning efficiency of primary and secondary school students, and relates to the technical field of recommendation learning, which comprises the following steps: S1, defining a feature vector of a word and a word list of learning materials; S2, establishing the feature vector of all words in an English learning textbook, initializing a word recommendation value dictionary and performing learning material warehousing; S3, dynamically updating the feature vector of the word and the recommendation value dictionary according to the performance of students in different learning types in the learning process; and S4, screening out candidate learning materials, calculating the recommendation value of the candidate learning materials through the word dictionary of the candidate learning materials and the word recommendation value dictionary, and pushing the learning materials with the highest recommendation value to the students; the application solves the problems that traditional word learning lacks personalized dynamic recommendation and is difficult to adapt to multimodal learning scenes and individual memory decay rules, realizes long-term learning effect visualized closed-loop optimization through the combination of dynamic feature engineering and multi-scene recommendation, and effectively improves the word learning efficiency.
Owner:读书郎教育科技有限公司