Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

59 results about "Active learning (machine learning)" patented technology

Active learning is a special case of machine learning in which a learning algorithm is able to interactively query the user (or some other information source) to obtain the desired outputs at new data points. In statistics literature it is sometimes also called optimal experimental design.

Efficient High-Entropy Alloys Design Method Including Demonstration and Software

Embodiments relate to system and methods involving use of a technique for managing a database for producing a material composition having a thermodynamic phase. The technique can include: receiving a binary phase diagram for each material to be used as a component of a high-entropy alloy (HEA); using one or more active learning machine learning techniques for generating a feature, the feature including: a primary feature that is representative of a probability that an HEA will exhibit a solid solution phase and / or an intermetallic phase, and a physics-based feature that is representative of a factor related to formation of a desired intermetallic HEA phase; encoding the primary feature and the physics-based feature; generating an output representation of a HEA alloy composition and phase of a predicted materials composition; and selecting a HEA composition and phase that will meet a material design criterion.
Owner:UNIV OF VIRGINIA PATENT FOUND

Multi-component catalyst active site prediction system and method fused with quantum embedding

The invention discloses a multi-component catalyst active site prediction system and method fused with quantum embedding, and relates to the technical field of catalysis and material informatics, and the system comprises a structure and site enumeration module which generates candidate sites; the adaptive quantum embedding calculation module obtains key reaction microcosmic parameters; the unified site fingerprint and feature engineering module constructs and fuses standard site fingerprints; the physical consistency machine learning module predicts adsorption energy and other parameters and uncertainty thereof; the active learning and sample selection module selects a high-value sample optimization model; the microdynamics evaluation module calculates index values such as activity; and the multi-objective optimization and sorting module generates an optimization sorting list. According to the method, the unification of calculation precision and efficiency is realized, the problem of non-unification of locus characterization is solved, the model interpretability and extrapolation reliability are improved, and the comprehensive evaluation and optimization sorting of multi-target performance are completed.
Owner:BEIJING ZHONGKE ARCLIGHT QUANTUM SOFTWARE TECH CO LTD

Micro-emulsion interfacial tension efficient prediction method and system based on active learning and molecular dynamics

The invention discloses a microemulsion interfacial tension efficient prediction method and system based on active learning and molecular dynamics. According to the method, 217 molecular descriptors corresponding to each molecular structure are calculated by adopting an RDKit software package, and the descriptors are used for representing molecular structure characteristics and serve as input variables of a machine learning model, so that key structure information including molecular branching degree, polarity and the like is transmitted. For an oil-water-surfactant ternary interface system, the oil-water interfacial tension in the presence of a surfactant is simulated and calculated through molecular dynamics, and an IFT value is set as a model prediction target. An active learning mechanism is introduced, and iterative sample labeling in the molecular dynamics simulation process is guided; and integrating the obtained IFT data with the molecular descriptor features, constructing a machine learning data set, and training a random forest model. According to the method, the problem of screening a high-performance surfactant layer by a middle-phase microemulsion system can be solved, and the ultra-low oil-water interfacial tension can be rapidly and efficiently screened.
Owner:SICHUAN UNIV

Dataset labeling using large language model and active learning

PCT designated stage expiredWO2025151197A1Biological modelsData setLinguistic model
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and / or the like for labeling data by (i) generating, using a natural language machine learning model, a labeled dataset from unlabeled data, (ii) training one or more instances of a classification machine learning model based on the labeled dataset, (iii) generating, using the one or more instances of the classification machine learning model, a plurality of validation classifications, and (iv) generating a refined labeled dataset that is based on the labeled dataset and a plurality of uncertainty scores associated with the plurality of validation classifications.
Owner:OPTUM INC

Montmorillonite thermal conductivity prediction method based on NEP machine learning potential

The invention relates to the technical field of machine learning potential molecular simulation, and provides a montmorillonite thermal conductivity prediction method based on NEP machine learning potential, and the method comprises the following steps: constructing a montmorillonite crystal structure model, carrying out first principle molecular dynamics calculation on the montmorillonite crystal structure model, introducing lattice distortion and atomic disturbance, and generating an initial structure set; performing static single-point DFT calculation on the initial structure set, and constructing a training data set; based on an NEP machine learning framework, utilizing the training data set to train and generate a potential function; optimizing the generation process of the potential function by adopting an active learning strategy; and performing thermal conductivity prediction verification on the potential function, and obtaining a thermal conductivity prediction result of the montmorillonite based on the potential function passing the verification. The defect that a traditional classical potential function depends on an empirical or semi-empirical model and is limited in precision and applicability is overcome, the problem that the dynamic behavior and the heat conduction mechanism of a complex material are difficult to accurately capture is solved, the calculation cost is reduced, and meanwhile the precision is kept.
Owner:ZHENGZHOU UNIV

Managing a model trained using a machine learning process

A computer implemented method of managing a first model that was trained using a first machine learning process and is deployed and used to label medical data. The method comprises determining (202) a performance measure for the first model, and if the performance measure is below a threshold performance level, triggering (204) an upgrade process wherein the upgrade process comprises performing further training on the first model to produce an updated first model, wherein the further training is performed using an active learning process wherein training data for the further training is selected from a pool of unlabeled data samples, according to the active learning process, and sent to a labeler to obtain ground truth labels for use in the further training.
Owner:KONINKLIJKE PHILIPS NV

AI-Based System and Method for Generating Enhanced Radiology Reports

The present invention relates to an AI-based system and method for generating enhanced radiology reports. The system comprises a database for storing multimodal patient data, a natural language processing (NLP) module for extracting clinical information, and a machine learning module for correlating the clinical information with radiology images to identify diagnostic insights. An AI-based report generation module analyzes the images and clinical information to generate a preliminary report, which is refined based on radiologist input. The generated report is then integrated into the patient's electronic health record. The system employs techniques such as multimodal deep learning, active learning, explainable AI, and federated learning to enhance diagnostic accuracy, capture expert feedback, provide transparency, and enable multi-institutional collaboration. The invention aims to improve the accuracy, efficiency, and value of radiology reporting in patient care.
Owner:DAVIS ALEXANDER

Dataset labeling using large language model and active learning

PendingUS20250232009A1Data setLinguistic model
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and / or the like for labeling data by (i) generating, using a natural language machine learning model, a labeled dataset from unlabeled data, (ii) training one or more instances of a classification machine learning model based on the labeled dataset, (iii) generating, using the one or more instances of the classification machine learning model, a plurality of validation classifications, and (iv) generating a refined labeled dataset that is based on the labeled dataset and a plurality of uncertainty scores associated with the plurality of validation classifications.
Owner:OPTUM INC

Active learning model validation

Methods, apparatus, and computer-implemented methods are provided for training machine learning (ML) techniques to generate a property model for predicting whether a compound has a particular property. An iterative process / feedback loop can be performed to generate a property model, the process including generating, based on the property model, a list of prediction results for a plurality of compounds and their association with the particular property, validating the property model based on compounds from the list of prediction results having an association with the particular property, and updating the property model based on the property model validation. The process / loop can be repeated using the updated property model until it is determined that the property model has been effectively trained. The property model validation can include selecting a list of compound candidates, performing simulation analysis and / or laboratory analysis on the list of compound candidates for association with the particular property, and updating the property model using the simulation and / or laboratory results.
Owner:BENEVOLENTAI TECH LTD

Alloy LIBS quantitative analysis method based on bayesian convolutional neural network and active learning

The application provides an alloy LIBS quantitative analysis method based on a Bayesian convolutional neural network and active learning. The method comprises the following steps: S1, spectrum data acquisition and preprocessing; S2, model construction and training; S3, active learning iteration optimization. Compared with a traditional machine learning model, the alloy LIBS quantitative analysis method provided by the application has high accuracy and good interpretability. Under the same number of labeled samples, the application shows better regression performance compared with a random sampling and a representative sampling strategy. Even if the number of labeled samples is reduced to 65% of the data set, the prediction performance of most components can reach 90% of full supervised learning, high accuracy is maintained, and good engineering application prospects are achieved.
Owner:CENT SOUTH UNIV

A machine learning potential energy model construction method based on hierarchical active learning

This invention relates to a method for constructing a machine learning potential energy model based on hierarchical active learning, comprising the following steps: constructing a database of background solvent molecules and lithium salt molecules; initializing a machine learning potential energy model committee; outer active learning automatically exploring the chemical species composition space of the electrolyte, generating electrolyte formulations and initial structures; determining whether a structure needs to be included in the labeling range based on uncertainty calculations; automatically constructing a liquid phase environment based on classical force field simulation methods; performing high-precision simulations based on first-principles calculations and labeling the energy and force values ​​of the structures; inner active learning driving the machine learning potential energy model to explore the configuration space and label the energy and force values ​​of structures with high uncertainty; obtaining a first-principles structure database of the real liquid phase environment through several iterations; and training to obtain the final high-performance machine learning potential energy model. This invention systematically and comprehensively samples the chemical species composition space of the electrolyte and the molecular simulation configuration space based on hierarchical active learning, thereby ensuring the predictive performance and simulation accuracy of the machine learning potential model on unknown and complex electrolyte systems.
Owner:CHONGQING UNIV

Efficient deep batch active learning method and system based on gradient regularization

The invention discloses an efficient deep batch active learning method and system based on gradient regularization, and relates to the technical field of machine learning and artificial intelligence, and the method comprises the steps: obtaining an unmarked sample, inputting the unmarked sample into a pre-established deep neural network model, and carrying out the output to obtain an output probability vector and a feature vector of a penultimate layer; according to the output probability vector and the feature vector of the penultimate layer, allocating a pseudo tag to the unmarked sample, calculating a gradient block of each category of the output layer based on the pseudo tag, performing regularization processing on the gradient blocks, retaining the gradient blocks corresponding to the category of the preset output probability, performing zero setting on the other gradient blocks, and generating a sparse gradient matrix; the method comprises the following steps: performing classification-level and sample-level hierarchical distance calculation on the basis of a sparse gradient matrix to obtain a sample-level distance, performing diversity sampling according to the sample-level distance, selecting representative samples by using a k-means + + clustering algorithm for labeling, and realizing acceleration and performance optimization of sample selection in deep batch active learning.
Owner:NANJING UNIV OF POSTS & TELECOMM

Training mRNA property prediction machine learning models using active learning

PCT designated stageWO2025264552A1BiostatisticsInstrumentsData miningBiology
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a machine learning model. In one aspect, a method comprises: determining a respective importance score for each mRNA codon in a set of possible mRNA codons; at each of a plurality of training iterations in a sequence of training iterations, selecting a current batch of mRNA sequences for training the machine learning model at the training iteration using the importance scores for the mRNA codons, and training the machine learning model on at least the current batch of mRNA sequences; and outputting the trained machine learning model.
Owner:SANOFI SA(FR)

Infertility patient pregnancy prediction method based on active learning and integrated model

The invention provides an infertility patient pregnancy prediction method based on active learning and an integrated model, and the method comprises the following steps: carrying out the preprocessing of collected data, including classification variable coding, standardization processing and multi-dimensional abnormal value detection; generating a few categories of synthetic samples through an SMOTE technology, and optimizing data distribution in combination with a Tomek Links method; training a pregnancy prediction model based on machine learning models such as a random forest and XGBoost; the prediction results of the multiple basic models are fused through the Blending technology; based on an uncertainty sampling strategy, selecting samples for labeling and carrying out mixed training; and the pregnancy success prediction probability and credibility are output by inputting patient data. According to the method, the prediction accuracy and stability of the model can be improved, the robustness of the model to noise data and boundary samples is improved, the generalization ability of the model on new data is improved to a certain extent, doctors and patients can obtain more accurate treatment decision support, and the success rate and efficiency of IVF-ET treatment are improved.
Owner:NANJING MEDICAL UNIV

Automating cryo-electron microscopy data collection

A method of automated control of a microscope in cryogenic electron microscopy (cryo-EM), wherein the microscope is configured to collect high-magnification micrographs of particles suspended in vitreous ice. Such particles are found in grid squares, and a square contains holes from which high-magnification micrographs are imaged. The method is carried out during an active data collection session, leveraging a pipeline that comprises a set of models. The pipeline evaluates a set of collection locations to determine whether to continue collection at a current grid / square or instead at a new grid / square. The evaluation is based on a set of one or more quality scores derived from one or more pretrained models and machine learning-based active learning. Based on the determination, control information is provided to automatically control the microscope to move to a next target for data collection.
Owner:NEW YORK STRUCTURAL BIOLOGY CENT

An intelligent collaborative optimization design method for a flexible OLED module protection structure

The application discloses an intelligent collaborative optimization design method for a flexible OLED module protection structure, and belongs to the technical field of flexible electronic protection and intelligent structure optimization. In view of the problems of the prior art, such as dependence on manual experience iteration, difficulty in taking into account the impact resistance and bending resistance, long design cycle, high trial and error cost, lack of physical constraints in traditional data-driven models, and insufficient prediction accuracy, the application first constructs a material-interface-module full-level mechanical response database, and collects test and simulation data under multiple working conditions. Then, the calibration and accuracy verification of material constitutive parameters are carried out, and a high-fidelity finite element simulation model is established to expand the data samples. On this basis, an agent model and a multi-objective optimization model are constructed by combining a machine learning algorithm, a high-precision mapping relationship between the structure parameters and the comprehensive protection performance is established, and finally, according to the target performance, reverse deduction and collaborative optimization are carried out, and the optimal structure parameters are output, and an active learning mechanism is introduced to realize the iterative updating of the model. The application realizes the paradigm shift of the flexible OLED module from "forward experience trial and error" to "target-driven reverse intelligent design", can simultaneously improve the impact resistance and bending resistance of the module, realizes the lightweight of the structure, greatly improves the design efficiency, and is suitable for the optimization design of the module protection structure of the folding screen, the crimping screen and the wearable flexible display terminal.
Owner:BEIJING ZHONGKE LIXIN TECH CO LTD

A supervised data generation method based on video semantic structural analysis

The present application belongs to the technical field of computer vision and machine learning, and discloses a supervised data generation method based on video semantic structured analysis, which comprises the following steps: initialization, construction of a labeled set and an unlabeled set, training of a video script generation model, preparation of a large visual-language model, determination of the number of active learning selection stages and the budget selected in each stage; sampling the video in the unlabeled set to obtain frames, generating image descriptions through the large visual-language model, and simultaneously generating scripts through the video script generation model; converting the image descriptions and scripts in text form into scene graphs by referring to SPICE; and calculating the learnability score of the unlabeled video by internally and externally measuring the samples, wherein the method follows a pool-based active learning process, effectively utilizes learnability, diversity and uncertainty to supervise data labeling of the most valuable videos.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN) +1

Active learning techniques for developing map prediction machine learning models to generate maps of biology

The present disclosure relates to systems, non-transitory computer-readable media, and methods for efficiently acquiring biological data via active machine learning. In particular, in some embodiments, the disclosed systems generate, utilizing a map prediction machine learning model, similarity prediction confidence scores for a plurality of perturbation pairs. In addition, in some embodiments, the disclosed systems determine, from the similarity prediction confidence scores utilizing an acquisition function, a perturbation pair for developing a ground truth similarity. Moreover, in some embodiments, the disclosed systems generate a tuned map prediction model by comparing a similarity prediction for the perturbation pair generated by the map prediction machine learning model with the ground truth similarity. Furthermore, in some embodiments, the disclosed systems utilize, in response to determining that a measure of confidence for the tuned map prediction model satisfies a stopping criterion, the tuned map prediction model to generate a blended map of biology.
Owner:RECURSION PHARMACEUTICALS INC

Method and apparatus for position parameter estimation

The application provides a position parameter estimation method and device, comprising: obtaining a semi-supervised data set and a received waveform at a target time; inputting the semi-supervised data set and the received waveform into a pre-constructed semi-supervised position parameter estimation model to obtain a distribution inference result; wherein the semi-supervised position parameter estimation model is trained based on a deep neural network using a sample data set using a variational inference theory, and at least comprises an encoder and a regressor. The semi-supervised position parameter estimation model provided by the application has the ability of active learning in inference, can effectively utilize the semi-supervised data set for parameter inference, can significantly reduce the communication, calculation and data set acquisition costs generated by the deployment of the machine learning algorithm in the positioning system, and realizes higher precision and higher generalization of the position parameter estimation.
Owner:TSINGHUA UNIVERSITY +1

Method, device and equipment for estimating reaction etching selection ratio of laminated material

The invention discloses a method, a device and equipment for estimating a reaction etching selection ratio of a laminated material, relates to the technical field of etching in microelectronic manufacturing, and is used for solving the problem that a micro-scale etching selection ratio is difficult to efficiently obtain in the prior art. Comprising the steps of obtaining a basic data set of a laminated material; integrating and screening the basic data set to obtain an initial sample; based on the initial sample, optimizing the basic data set by adopting an iterative optimization mechanism of a machine learning dynamic situation function of active learning iterative sampling to obtain a final potential function; and performing molecular dynamics simulation based on the final potential function, and estimating the reaction etching selection ratio of the laminated material in combination with an adaptive time step algorithm. According to the invention, a high-precision machine learning potential function can be formed and a dynamic model of an etching process can be completed; and quantitatively calculating the reaction rate difference of each layer of material, and deducing the etching selection ratio.
Owner:SEMICON TECH INNOVATION CENT(BEIJING) CORP +1

An open set image recognition method and system based on active learning

The present application relates to the technical field of machine learning open set image recognition, and provides an open set image recognition method and system based on active learning, which can fully utilize pseudo-labeled information of unlabeled data, reduce the influence of open set samples on the classification model, expand the labeled data set, and improve the accuracy of open set image recognition.
Owner:SHANXI UNIV

Data-driven reverse development method for low-melting-point fused salt heat storage material

The invention relates to a data-driven reverse development method for a low-melting-point fused salt heat storage material, which is based on an advanced machine learning algorithm, realizes directional development of low-melting-point fused salt through an arranged existing fused salt data set, predicts the components of the low-melting-point fused salt through a built active learning framework, and verifies through experiments. And the development period of the novel molten salt material is greatly shortened. An active learning strategy for directional development of the low-melting-point molten salt material is designed, the time from adjustment of hyper-parameters of a machine learning algorithm to final experimental characterization is short and can be completed in one day, while a traditional experimental development strategy is too high in randomness and basically takes several months as a development period, and the development period is shortened. According to the method, the development efficiency of the high-performance molten salt is greatly improved, the number of experiments is greatly reduced, and the development cost of the molten salt is reduced.
Owner:ZHEJIANG UNIV +2

Method and system for automatically detecting errors in industrial plant configuration table

To automatically detect errors in an industrial plant configuration table (CT), a table encoder (TE) is trained to estimate an estimation probability that cells in the table contain errors. According to an embodiment, a table encoder initially trains with a set of configuration tables. An active learning query policy (ALQS) selects candidate cells (CCs) from a set of configuration tables using a hybrid policy that combines uncertainty sampling with penalties for selecting multiple candidate cells from the same configuration table. A user interface receives a tagged cell (LC), wherein the tagged cell includes a candidate cell and a tag indicating whether the candidate cell is erroneous. A training component (TC) performs gradient updates (GU) on the table encoder using the marked cells. Thus, after each user interaction, the table encoder is retrained to become a table token classification model. Active learning is used, which is novel for token classification in table data, allowing efficient reduction of the amount of tagged data required. The active learning query strategy balances uncertainty and diversity and ensures that candidate cells are selected from configuration tables of different ranges. The embodiment provides a novel active learning workflow for table data and a deep machine learning model. Still further, this embodiment saves the time of an expert in annotating cells in a configuration table.
Owner:SIEMENS AG

Cross-batch power data intelligent labeling method

The invention discloses a cross-batch electric power data intelligent labeling method, which comprises the following steps of: constructing a hierarchical label scene tree, scientifically dividing data category levels according to an electric power safety supervision business process, formulating a unified label standard, integrating industry specifications and actual demands, and defining label meanings, ranges and values; automatic standard tools such as basic rules and machine learning are combined with interactive labeling methods such as active learning and crowdsourcing labeling to improve labeling efficiency; a Word2Vec word vector model and a semantic similarity calculation, dynamic updating and consistency detection method are adopted to realize semantic efficient mapping, and the method has the capabilities of quickly adjusting a label system and adaptively optimizing an annotation and mapping algorithm for a new data type and format, improves the cross-batch data processing capability of electric power safety supervision, and improves the data processing efficiency of the electric power safety supervision. The defects in the aspects of label consistency, labeling efficiency, semantic mapping accuracy and new data adaptability in the prior art are overcome.
Owner:FUJIAN YIRONG INFORMATION TECH +1

Two phase flows for reactions and separations

Disclosed herein is a method for designing a liquid-liquid biphasic micro-fluidic flow channel reactor for continuous extraction or reactive extraction, where chemistry happens in one phase and the product is removed to the other. The method comprises developing random forest and symbolic genetic regression machine learning (ML) models to predict flow patterns and the mass transfer rate, respectively, using a combination of experimental and computational fluid dynamics (CFD) data and literature-mined data while accounting for the effects of solvent properties and channel diameter. This enables rapid prediction for efficient scale-up of microchannels to millichannels. To minimize the number of CFD simulations and maximize model accuracy, the method comprises using active learning techniques.
Owner:UNIVERSITY OF DELAWARE

High-entropy alloy structure performance collaborative prediction method adopting active learning strategy

The invention provides a high-entropy alloy structure performance collaborative prediction method adopting an active learning strategy, and belongs to the technical field of machine learning application. The method solves the problems that the prior art depends on a fixed data set, cooperative and accurate prediction of the structure performance of the high-entropy alloy cannot be achieved under the conditions of sample scarcity and uneven distribution, and the generalization ability is weak. The technical scheme comprises the following steps: constructing and preprocessing a data set; constructing a feature set based on the properties of materials, and determining a key feature subset through multi-stage screening; training an initial collaborative prediction model; an active learning strategy is adopted, and high-value samples are screened through uncertainty calculation and clustering analysis for experiments; and an experimental result is fed back to the model for iterative optimization. According to the method, the prediction precision and generalization ability of a complex component space are remarkably improved, the model interpretability is enhanced, continuous autonomous optimization under the small sample condition is achieved, the method is suitable for various high-entropy alloy systems, and the research and development efficiency of materials can be greatly improved.
Owner:CHINA UNIV OF MINING & TECH

A water quality detection method and system based on machine learning combined with spectrophotometry

The application provides a water quality detection method based on machine learning combined with spectrophotometry, wherein, in the embodiment of the application, the absorbance of a water sample is measured at multiple preset wavelengths by a spectrophotometer to form a multidimensional absorbance data set, the data set is mapped to a high-dimensional space through a kernel function technology to obtain a data set with enhanced nonlinear characteristics, known water samples and the pollutant concentrations corresponding to the known water samples are extracted from historical water quality data to generate a training sample set, a machine learning model is applied to predict the pollutant concentration of the data set with enhanced nonlinear characteristics, an active learning strategy is used to dynamically adjust the training sample set to obtain a pollutant concentration prediction value, the pollutant concentration prediction value is used in combination with water quality safety standards to evaluate the safety level of the water sample, and a water quality report is output, the technical solution provided by the application improves the efficiency and accuracy of water quality detection, and provides strong technical support for environmental protection and public health protection, and the entire process realizes integrated operation from data acquisition to result interpretation, ensuring the scientificity and reliability of water quality management decisions.
Owner:CSSC HAISHEN MEDICAL TECH CO LTD

Automated phenotypic analysis of microscopic images using a cascade ml architecture with iterative active learning

The present disclosure relates to systems and methods for automated phenotypic analysis of microscopic images using a cascade machine learning architecture with iterative active learning. The systems and methods include a cascading flow of data through a plurality of machine learning models from images to specific phenotypic detection. The systems and methods use a feedback loop for continually improving the predictions of the plurality of machine learning models.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC