Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

160 results about "Unbalanced data" patented technology

Unbalanced data. In this context, unbalanced data refers to classification problems where we have unequal instances for different classes. Having unbalanced data is actually very common in general, but it is especially prevalent when working with disease data where we usually have more healthy control samples than disease cases.

Method for training llms based recommender systems using knowledge distillation, recommendation method for handling content recency with llms, solving imbalanced data with synthetic data in impersonation and deploying state of the art generative ai models for recommendation systems

A system and method for facilitating training of large language model based recommender systems are provided. The system may utilize one or more LLMs to create probability distributions for binary classification tasks associated with specific user-item pairs. The probabilities may be utilized to rank one or more tasks directly. The training of the one or more LLMs may involve the use of Knowledge Distillation methods and may be based on incorporating a dual-label system such as, for example, hard labels and soft labels. The one or more LLMs training data may consist of user-item pairs and their corresponding features. The labels used in the training process may include binary classification labels and their respective probabilities. The system may further implement the trained one or more LLMs to determine rankings or recommendations associated with user engagement of one or more content items.
Owner:META PLATFORMS INC

Electric power system lightning disaster risk detection method and device oriented to unbalanced data

The invention relates to an unbalanced data-oriented electric power system lightning disaster risk detection method, which comprises the following steps of: S1, gridding a target detection land parcel to obtain a plurality of grid units; S2, based on latitude and longitude coordinates of each grid unit, obtaining line parameters and environment characteristics of each grid unit in combination with a line database and a geographic information base, wherein the line parameters comprise tower height, span, grounding resistance, loop number, lightning arrester number and transformer capacity, and the environment characteristics comprise soil conductivity, building height and land utilization type; and S3, inputting the line parameters and the environment characteristics of each grid unit into a trained first detection model to obtain the thunder and lightning probability and the maximum peak current of each grid unit. Compared with the prior art, the multi-source data such as the tower height, the span, the grounding resistance and the soil conductivity are comprehensively considered, and the risk assessment result has higher regional adaptability and accuracy.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Motor fault diagnosis method for unbalanced data in noise environment based on dual-scale network

The invention provides a motor fault diagnosis method for unbalanced data in a noise environment based on a dual-scale network, and the method comprises the steps: collecting acceleration signals of motor vibration in a normal state and a fault state, and constructing a data set; building a fault diagnosis model based on a dual-scale network, wherein the fault diagnosis model comprises a wide convolution kernel 1D convolution module, an attention-based dual-scale module, a convolution-based feature fusion module, an attention 2D convolution layer and a classification module which are connected in sequence; based on the data set, using a category weighted loss function to train a fault diagnosis model based on a dual-scale network to obtain a trained fault diagnosis model; and inputting a motor vibration signal acquired in real time into the trained fault diagnosis model for fault diagnosis to obtain a fault type. According to the method, the influence of noise interference and uneven data on the neural network model is weakened by mining deep fault features and endowing different weights to the loss function.
Owner:ZHENGZHOU UNIV +1

Generative adversarial network unbalanced data processing method based on dynamic density guidance

The invention relates to a dynamic density guided generative adversarial network unbalanced data processing method (DAG-WGAN). The DAG-WGAN realizes unbalanced data processing through data preprocessing, dynamic density estimation and weight distribution, potential structure learning based on a variational auto-encoder (VAE), and density guide generation and dynamic feedback optimization based on WGAN-GP. The DAG-WGAN adaptively evaluates the sample density by using kernel density estimation (KDE) and a Gaussian kernel function, and allocates a weight for a generation process, thereby emphatically enhancing the low density and discriminating the sample generation of a difficult region. The VAE learns a potential manifold structure of a minority class of samples, realizes density-guided generation of a potential space under a WGAN-GP framework, and ensures diversity and manifold consistency of generated samples. In addition, a dynamic feedback mechanism is introduced, the weight and the gradient penalty coefficient are adaptively adjusted and generated, and the training stability and the sample generation robustness are improved.
Owner:HARBIN UNIV OF SCI & TECH

Cooperative medical prediction system oriented to heterogeneous data center

The invention discloses a cooperative medical prediction system for a heterogeneous data center, and relates to the field of intelligent medical treatment, and the system comprises a distributed client set which is used for carrying out the localized training based on a local private medical image and dose data; the centralized coordination node is used for managing and coordinating a federation training process and executing aggregation of cross-client parameters; the decoupling model architecture comprises a globally shared feature encoder and feature adapters unique to a plurality of clients; the alternative training scheduling module is configured to periodically switch between a global aggregation mode and a localization adaptation mode, and the adaptive aggregation weighting module is used for dynamically calculating and distributing the weight of each client in global aggregation based on the statistical difference between data distribution and overall distribution of each client. According to the scheme, the overall generalization performance and prediction stability of the model on heterogeneous multi-center data can be improved, and the performance difference between centers caused by unbalanced data distribution is relieved.
Owner:ZHEJIANG CANCER HOSPITAL

Multi-generator adversarial network intrusion detection method based on imaging variational enhancement

The invention discloses a multi-generator adversarial network intrusion detection method based on imaging variational enhancement in the technical field of network security, which comprises the following steps of: 1, preprocessing data and encoding images, converting network flow data into a two-dimensional image format, and reserving spatial-temporal characteristics and protocol characteristics of the data for subsequent model training; step 2, constructing a multi-generator adversarial network, adopting a plurality of generators to work in parallel, each generator being responsible for generating attack samples of required categories, and optimizing model parameters of the generators and discriminators through an adversarial training process to enable the distribution of the generated attack samples to be close to the distribution of real attack samples; according to the method, the sample and the classification model are generated through collaborative optimization, the robustness and generalization ability of a network intrusion detection system on an unbalanced data set are remarkably improved, and an innovative solution is provided for network security detection.
Owner:YANGZHOU UNIV

Traffic accident detection method based on VAE model

The invention discloses a traffic accident detection method based on a VAE model, and the method comprises the following steps: obtaining original data, and carrying out the primary processing of the obtained data; dividing a data set; converting a data format into a tensor format; performing cross validation on the data; pre-training the VAE model, adjusting and training a classifier, calculating a sample reconstruction error, KL divergence, a potential spatial distance and an output probability of the classifier, and calculating a mixed score through a mixed scoring formula; according to the method, the accuracy of a traffic accident detection model is improved, the accuracy of accident detection in an unbalanced data scene is effectively improved by fusing the characterization learning ability of the VAE and the discrimination ability of the classifier, and a mixed scoring mechanism is further introduced, so that the accuracy of the traffic accident detection in the unbalanced data scene is improved. Information of multiple dimensions such as VAE reconstruction error, KL divergence, potential spatial distance and classifier output probability is integrated, and the robustness of the model is enhanced.
Owner:SICHUAN POLICE COLLEGE +1

Two-stage air conditioning system fault diagnosis method and system based on improved deep residual network

The invention discloses a two-stage air conditioning system fault diagnosis method and system based on an improved deep residual network, and relates to the technical field of air conditioners. The method comprises the following steps: collecting historical operation data of equipment and parts of the heating ventilation air-conditioning system, and obtaining normal samples and fault samples; after the samples are preprocessed, labels are set for the samples according to categories, and normal-category samples and multi-category fault samples are obtained; constructing a deep residual network model fusing long-tail learning and an attention mechanism, inputting normal samples and various fault samples for training, outputting a predicted value of a data category label and determining a predicted fault type, thereby obtaining a trained heating ventilation air-conditioning system fault diagnosis model; and acquiring actual operation data, inputting the actual operation data into the trained heating ventilation air-conditioning system fault diagnosis model to carry out two-stage fault diagnosis, and outputting a fault type. According to the invention, rapid detection and accurate positioning of the fault of the heating ventilation air-conditioning system in a high imbalance data scene are realized.
Owner:UNIV OF SCI & TECH BEIJING

Unbalanced disaster risk prediction method based on WGAN-CNN

The invention discloses an unbalanced disaster risk prediction method based on a WGAN-CNN, relates to the technical field of disaster risk prediction, and aims to improve the prediction capability by optimizing feature learning of WGAN and CNN models in order to solve the problem of predicting rare disasters by using unbalanced and heterogeneous data. Data sources comprise satellite images, radars, social media and the like, and data consistency is ensured by unifying multi-modal data timestamps and aligning time sequences through dynamic time warping; then, using an improved WGAN to generate a rare disaster sample so as to enhance unbalanced data, and realizing data security sharing through encryption gradient and differential privacy technologies; in addition, based on geological similarity cross-neg region mapping weights, risk levels are evaluated and pushed in real time in combination with a CNN discriminator. According to the method, the data quality is improved through the WGAN, accurate prediction is realized by using the CNN, a real-time risk assessment tool applicable across regions is finally formed, and the rare disaster prediction capability is effectively improved.
Owner:INFORMATION RES INST OF EMERGENCY MANAGEMENT DEPT

LED automobile lamp failure prediction method and system based on data analysis

PendingCN122346653APredictive methodsSimulation
The present application relates to the technical field of fault detection, more particularly, the present application relates to a LED automobile lamp fault prediction method and system based on data analysis, comprising: acquiring multi-dimensional time series data of LED automobile lamps in running state; constructing direction drift index of each moment, for quantifying the direction consistency of signal drift in the corresponding window at this moment; based on the correlation between multi-dimensional time series data in each detection window, the multi-dimensional linkage coefficient of each moment is constructed. The present application combines the physical degradation mechanism to construct the direction drift index and the multi-dimensional linkage coefficient, accurately quantifies the consistency and synchronism of multi-dimensional parameter degradation, and then adaptively divides the degradation stage and evaluates the learning difficulty of each stage; In the training of the fault prediction model, the stage difficulty is used as the weight to dynamically correct the gradient. This effectively overcomes the problem of extremely unbalanced data ratio, realizes the sensitive capture and accurate early warning of the weak degradation characteristics of the lamp.
Owner:GUANGZHOU ETHER AUTOMOTIVE LIGHTING LTD

Database sharding method, apparatus, and computer program product

PendingCN122654213AShardCustomer information
The application provides a database sharding method, device and computer program product, the method comprising: obtaining a virtual account information table, the virtual account information table comprising customer information and virtual account attribute information; generating a plurality of intermediate table records according to the data of the virtual account information table, the intermediate table records comprising a type label of a job task of a batch job flow and a sharding value of the job task; obtaining the data of the virtual account information table corresponding to the sharding value of the job task of the intermediate table record, obtaining a sub-table of the virtual account information table corresponding to the intermediate table record, and determining the number of job shards corresponding to each sub-table according to the type label of the job task of the intermediate table record; and performing processing on each sub-table according to the corresponding job task by using threads of the number of job shards corresponding to each sub-table according to the flow of the batch job flow, and obtaining a processing result, thereby solving the problem of low distributed service processing efficiency caused by unbalanced data sharding of a database in the prior art.
Owner:中国邮政储蓄银行股份有限公司

Rusboost island detection method applied to direct current microgrid

This invention discloses a RUSBoost islanding detection method applied to DC microgrids. Specifically, it involves: collecting historical data of electrical characteristics under both grid-connected and islanded states to form an unbalanced dataset; using the historical data as a training set and preprocessing it to form a training sample set; constructing a weak islanding classifier based on classification and regression decision trees; evaluating the model's correctness using a confusion matrix; and establishing a RUSBoost-based islanding detection model; applying the constructed RUSBoost-based islanding detection model to the microgrid system to classify grid-connected and islanded states based on real-time voltage and current data. This invention applies an ensemble classification algorithm from machine learning to DC microgrids, solving the problems of slow detection speed and low accuracy of existing detection methods.
Owner:XIAN UNIV OF TECH

Diversity-aware weighted majority vote classifier for decision making on imbalanced datasets

An ensemble learning based method is for a binary classification on an imbalanced dataset. The imbalanced dataset has a minority class comprising positive samples and a majority class comprising negative samples. The method includes: generatively oversampling the imbalanced dataset by synthetically generating minority class examples, thereby generating a generated dataset; using the generated dataset to generate subsamples, and learning a base classifier on each of the subsamples to determine a plurality of base classifiers; and learning a weighted majority vote classifier by combining outputs of the base classifiers. Each of the base classifiers is assigned a weight in such a way that a diversity between the base classifiers on the positive samples is minimized.
Owner:NEC CORP

A transformer fault diagnosis method based on small sample unbalanced data set

The application discloses a transformer fault diagnosis method based on a small sample unbalanced data set, comprising the following steps: S1, acquiring a transformer dissolved gas analysis data set and a comprehensive feature set; S2, constructing a transformer fault diagnosis model based on a small sample unbalanced data set; S3, according to the gas analysis data set and the comprehensive feature set, performing model training on the constructed transformer fault diagnosis model to obtain an optimal fault diagnosis model; and realizing transformer fault diagnosis based on the small sample unbalanced data set according to the optimal fault diagnosis model. The application solves the problems that the current traditional method cannot accurately diagnose transformer faults due to the technical problems such as complex fault mode recognition, specific class accurate diagnosis, insufficient model generalization ability and difficulty in fusion parameter optimization under a small sample unbalanced fault data set in the prior art.
Owner:SHENYANG AGRI UNIV

Multi-label smell description prediction method

The invention discloses a multi-label smell description prediction method, and relates to the field of compound smell prediction, and the method comprises the steps: obtaining compound identification information, molecular structure descriptors and smell label data, and constructing a multi-label smell data set; generating a molecular structure feature vector through a molecular fingerprint coding technology, and extracting a multi-dimensional descriptor reflecting the physicochemical properties of molecules; compressing the molecular fingerprint features to a low-dimensional space through a dimension reduction algorithm; performing unbalanced data processing on the training set, fusing the dimension-reduced molecular fingerprints with the molecular descriptors to form a joint feature matrix, and configuring a class weight balance mechanism and overfitting suppression parameters by adopting a multi-label classification architecture; independently optimizing a probability threshold for each odor label based on the verification set; and outputting a multi-odor label combination prediction result according to the target molecule identification information. According to the scheme, the multi-odor characteristics of the compound can be accurately depicted, and the combined recognition accuracy of the compound odor is remarkably improved.
Owner:RES CENT FOR ECO ENVIRONMENTAL SCI THE CHINESE ACAD OF SCI

Noise-free loss distribution migration-based credit data synthesis oversampling method and system

ActiveCN121167241AData setAlgorithm
The invention discloses a credit data synthesis oversampling method and system based on noise-free loss distribution migration, and relates to the technical field of unbalanced classification in artificial intelligence and data mining, and the method comprises the steps: obtaining an original credit unbalanced data set; noise label sample filtering is carried out on the data set, the prediction probability of each sample in the data set is predicted through a model, and the noise-free loss value of each sample is calculated; dividing a plurality of loss intervals according to the value range of the noise-free loss value, distributing each sample to the corresponding loss interval, and determining noise-free loss distribution of majority-class and minority-class samples; migration is carried out based on noise-free loss distribution, synthetic sample distribution is determined, root samples and auxiliary samples are screened out, and minority class pseudo samples are synthesized through linear interpolation; and adding the minority class pseudo samples into the original credit unbalanced data set to obtain a class balanced credit data set. According to the optimized balanced data set, the recognition precision of majority-class samples and minority-class samples can be effectively improved.
Owner:SHANDONG CREDIT INFORMATION CO LTD

Method for training gait recognition model of unbalanced data acquired in real world

The invention discloses a method for training a gait recognition model of unbalanced data acquired in the real world, and relates to the technical field of gait recognition data analysis, and the method comprises the steps: obtaining experiment balance gait data and real-time gait data of a corresponding time point; according to the method, gait category real-time analysis and labeling are carried out on a balance data set and a real-time data set, a feature vector feature center point value under each gait type is calculated, and a mean square error is used for optimizing an error of the center point value between experimental balance gait data and real-time gait data at a corresponding time point; setting a judging and comparing program of the real-time feature vector and the balance feature vector, calculating the probability of belonging to the same class by adopting cosine similarity, taking the real-time feature vector after projection dimension reduction as input sample data of an initial gait recognition model, taking the balance feature vector of a balance data set as a supervision label, training the gait recognition model, and obtaining a real-time gait recognition model; therefore, accurate recognition of real-time gaits is realized.
Owner:上海启眼科技有限公司

A power transmission line insulator fault detection method and system based on semi-supervised ensemble learning and edge computing

This invention proposes a method and system for fault detection of power transmission line insulators based on semi-supervised ensemble learning and edge computing, comprising: Step S1: Constructing a hierarchical imbalanced insulator image dataset; constructing a high imbalance dataset in conjunction with actual power inspection scenarios; Step S2: Training a lightweight two-stage deep learning model on the host side; including training a lightweight two-stage model that meets the requirements for edge deployment on the development board; Step S3: Designing a semi-supervised ensemble learning fault diagnosis framework based on Fourier feature enhancement, achieving dual utilization of labels and data distribution through distance threshold collaborative decision-making; Step S4: Model conversion and lightweight deployment on edge devices, including implementing model inference on the development board; Step S5: Integrated fault detection and real-time response at the edge, including realizing an integrated process from image acquisition to target detection to fault diagnosis to alarm output.
Owner:STATE GRID FUJIAN ELECTRIC POWER RES INST +1

Target model training method and device for individualized curative effect evaluation

The embodiment of the invention discloses a target model training method and device for individualized curative effect evaluation, and the method comprises the steps: obtaining sample observation data, fitting the relation between expected outcome data and target body feature data, and obtaining an individual outcome baseline prediction value; fitting a relation between the treatment state data and the target body feature data to obtain an individual treatment probability prediction value; determining a first difference value between the individual outcome baseline predicted value and the actual outcome of the individual; determining a second difference between the individual treatment probability prediction value and the actual treatment state of the individual; estimating an individualized treatment effect; constructing a target equivalent relational expression; training and learning the initial model to enable the target loss function to be the value of the individualized treatment effect to be the minimum, and obtaining a target model. According to the scheme, based on a causal inference theory and a machine learning modeling strategy, an existing causal machine learning method is improved so as to adapt to challenges brought by sample imbalance data, and scientificity and practicability of individualized intervention decision are improved.
Owner:NANJING UNIV OF TRADITIONAL CHINESE MEDICINE +1

A method and system for handling imbalanced data based on nearest neighbor constraints and consistency screening.

This invention relates to the field of data processing technology, and in particular to a method and system for processing imbalanced data based on nearest neighbor constraints and consistency screening. The method includes acquiring industrial fault detection data; performing data preprocessing based on the acquired industrial fault detection data; constructing a minority class MNN skeleton based on DenMune clustering and performing two-layer noise removal; local adaptive oversampling based on MNN sparsity double-layer weights and MNN neighborhood constraint interpolation; dual-threshold screening based on source sub-cluster consistency and original sample dominance; and using a trained model to predict and classify fault samples. This effectively suppresses the introduction of new noise during oversampling, making the final training dataset more stable near the local structure and decision boundary, which is beneficial for improving the stability and robustness of the industrial fault detection model in imbalanced scenarios.
Owner:YANTAI UNIV

Railway signal fault text enhancement method, system and equipment based on improved EDA

The invention discloses a railway signal fault text enhancement method, system and device based on improved EDA, and the method comprises the following steps: constructing a railway field entity dictionary, carrying out the proper noun recognition of a preprocessed to-be-enhanced railway signal fault text, and outputting a structured text; on the basis of a railway signal fault related text data set, a similar word bank of railway signal fault text related vocabularies is generated by adopting a Word2Vec model and used for replacing non-professional vocabularies in railway signal fault text statements, and enhancement operation is performed on the basis of a structured text and the similar word bank of the railway signal fault text related vocabularies by adopting an EDA improvement method. Enhancing the railway signal fault text to be enhanced; and screening the enhanced railway signal fault text. According to the method, the integrity and accuracy of domain terms are ensured through an entity fixing mechanism, an enhanced sample conforming to a real scene is generated in combination with a semantic constraint editing strategy, and the problem of unbalanced data distribution is effectively solved.
Owner:YANSHAN UNIV

Adaptive boundary oversampling method and system based on DBSCAN and Gini coefficient

The invention provides a self-adaptive boundary oversampling method and system based on a DBSCAN and a Gini coefficient, and belongs to the technical field of data processing and machine learning. The method comprises the steps that firstly, density clustering is conducted on a whole data set through a DBSCAN algorithm, and a natural cluster structure and noise are recognized; constructing a local neighborhood unit, accurately detecting a decision boundary region by using a Gini coefficient, and determining a boundary seed point; then, according to the scarcity degree of the minority class clusters, self-adaptive oversampling weights and the sampling number are distributed; and finally, directionally generating a synthetic sample in the cluster based on the boundary seed points to realize data set equalization. According to the method, the boundary identification fine granularity can be improved, the scarce cluster boundary is ensured to be intensively enhanced, invalid samples are prevented from being generated in most abdominal regions or blank regions, the performance of an unbalanced data classification model is remarkably improved, and the method can be widely applied to scenes depending on unbalanced data modeling, such as credit risk assessment, fault diagnosis and medical diagnosis.
Owner:ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE

Offshore wind turbine tower safety consequence minimization risk assessment method

PendingCN122451687AData setRisk rating
The application discloses a method for minimizing the risk of safety consequences for offshore wind power towers. The method includes: obtaining a multi-channel time sequence monitoring signal of the wind power tower; performing a working condition awareness data standardization process on the multi-channel time sequence monitoring signal based on the working condition interval of the unit operation, stripping the normal working condition fluctuations, and obtaining a monitoring data set; extracting risk assessment features representing the structural stress state based on the monitoring data set and the original monitoring signal, the features including cross-channel structural dynamic response proxy features; inputting the risk assessment features into a pre-trained risk level classification model, using the multi-classification cost-sensitive decision mechanism of the model to determine the target level with the minimum expected safety consequences among the candidate risk levels, and outputting the operation safety risk level. The application can solve the problems of easy missed report of high-risk state, multi-classification voting degradation and working condition interference under extremely unbalanced data.
Owner:HUANENG RUDONG BAXIANJIAO OFFSHORE WIND POWER GENERATION CO LTD +2

A power distribution network fault abnormality monitoring and early warning method based on multi-source operation data fusion

PendingCN122266133AImprove coverage integrityReduce abnormal omissionsAlarmsFault locationLow voltageElectric power system
The application provides a power distribution network fault abnormality monitoring and early warning method based on multi-source operation data fusion, and relates to the technical field of power system distribution network fault monitoring and early warning. The application obtains distribution network equipment operation data, defect abnormality data, equipment abnormality data, heavy overload data, low voltage data and three-phase imbalance data in a marketing system, a production system, a safety monitoring system and a dispatching system, forms multi-source abnormality data, classifies and aggregates the multi-source abnormality data according to line area, abnormality type and collection time, and performs same-source deduplication and cross-source merging to obtain fused abnormality data, analyzes and processes the fused abnormality data, identifies data representing distribution network fault abnormality and generates a fault abnormality event, and then determines a responsible unit and a line area person in charge according to the corresponding relationship between the line area and the person in charge and dispatches them, so that the monitoring and early warning of the distribution network fault abnormality are realized, and the pertinence and accuracy of the monitoring and early warning are improved.
Owner:BAZHOU POWER SUPPLY CO OF STATE GRID XINJIANG ELECTRIC POWER CO LTD

Disease grading method and device based on spatial uncertainty perception and knowledge fusion

The invention discloses a disease grading method and device based on spatial uncertainty perception and knowledge fusion, and belongs to the field of medical data processing technology and computer deep learning. According to the method, a balanced data set is constructed based on an unbalanced data set, a pre-trained teacher model is utilized to carry out fine tuning on the two types of data sets, a spatial uncertainty perception decoupling distillation module and a representation decomposition learning module are constructed, and local prediction uncertainty of the teacher model is quantified, so that the spatial uncertainty perception decoupling distillation module and the representation decomposition learning module are obtained. Knowledge transfer weight is dynamically adjusted, deviation propagation is reduced, feature learning is divided into low-layer structure feature alignment and high-layer semantic discrimination feature extraction through a representation decomposition learning module, the capture ability of the model for local lesions is enhanced, a student model is trained, and disease classification is output. The superiority of the method is verified on a medical image data set, the disease grading accuracy and generalization ability in a class imbalance scene are remarkably improved, and reliable support is provided for clinical auxiliary diagnosis.
Owner:RESEARCH INSTITUTE OF TRANSVASCULAR IMPLANTATION EQUIPMENT ZHEJIANG MEDICAL SECOND HOSPITAL BINJIANG DISTRICT HANGZHOU

A multi-source data dynamic risk early warning method and system based on ST-GAN

The application discloses a kind of based on ST-GAN's multi-source data dynamic risk early warning method and system, the method includes the following steps: S1. constructing the spatiotemporal generation confrontation network suitable for spatiotemporal data characteristics, utilize the network to generate extreme precipitation data, realize spatiotemporal unbalanced data reduction, obtain the equalization data of final output;S2. the equalization data of final output with high-dimensional space variable is carried out multi-source data fusion, and nonlinear spatiotemporal information conversion equation, and by local linearization obtains linear approximation model, to predict future time series;S3. under the present situation that extreme rainfall data amount is relatively insufficient, neural network based on dual learning theory accurately learns the parameter of nonlinear spatiotemporal conversion, estimates extreme weather event.The application is through spatiotemporal generation confrontation network (ST-GAN) and multi-source data fusion engine, significantly improves the precision and efficiency of natural disaster warning.
Owner:SI CHUAN KE RUI RUAN JIAN YOU XIAN ZE REN GONG SI

A hyperspectral remote sensing image classification-oriented training sample selection method

The application discloses a kind of training sample selection methods for hyperspectral remote sensing image classification, belong to image classification technical field.The following steps are included: given unlabelled dataset, utilize pre-training GSCVIT model to extract classification features and multi-head attention weight, calculate spatial attention entropy and be spliced into enhanced features with classification features, through K-Center greedy algorithm screening core set sample subset, with the aid of group sampling loader loading data, using dynamic distribution balance loss (DDB Loss) training model to optimize classification performance.The application is verified on four datasets, and the results show that the method can effectively select key samples, dynamically optimize class distribution, significantly improve the classification accuracy and stability of the model under unbalanced data, and enhance the recognition ability of minority classes and weak class targets.Solve the problems of high labeling cost, class imbalance and complex feature expression of hyperspectral remote sensing image.
Owner:JILIN UNIVERSITY

Ship navigation risk early warning method based on unbalanced marine weather data enhancement

The application discloses a ship navigation risk early warning method based on unbalanced offshore meteorological data enhancement, which firstly determines the ship navigation risk grade, and divides the unbalanced data set of target sea area meteorological disaster-causing elements corresponding thereto. Secondly, an adversarial learning model composed of a deep neural network generator and an evidence reasoning discriminator is designed to balance the unbalanced data set under different risk grades, and a Gaussian distribution model is established to describe the feature distribution of meteorological disaster-causing elements under different risk grades. Then, the risk grade reliability distribution of meteorological disaster-causing elements is calculated through the Gaussian distribution model. Finally, the weighted average method is adopted to fuse the risk grade reliability distribution, and the risk mode with the highest reliability after fusion is selected as the navigation risk grade of the current target sea area. The application can generate a small number of meteorological data conforming to the real distribution, effectively solve the data imbalance problem, and improve the extreme weather risk warning capability.
Owner:HANGZHOU DIANZI UNIV

Network intrusion detection method and device, equipment, storage medium and program product

The invention discloses a network intrusion detection method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring network traffic data to be processed; detecting the network flow data by adopting a pre-trained network intrusion detection model to obtain a network intrusion result; wherein the network intrusion detection model is obtained by training a centralized multi-core multi-class support vector machine through a pre-constructed sample set, and the sample set is obtained by performing data enhancement on abnormal samples subjected to network intrusion through a synthetic minority class oversampling algorithm based on distance weight; the problem of unbalanced data sets can be effectively solved, meanwhile, the selection problem of kernel functions can be avoided, and the network intrusion detection performance of the model is improved.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Self-adaptive collaborative oversampling and unbalanced multi-classification method and system

PendingCN121145021AData setAdaBoost
The invention discloses a self-adaptive collaborative oversampling and unbalanced multi-classification method and system, and the method comprises the steps: selecting a minority class sample which is difficult to classify as a target sample for synthesis through the weight generated by AdaBoost before each iteration, and can avoid the synthesis of an invalid sample. In addition, a sample synthesis mode is changed, the phenomenon of uneven local interpolation is avoided, meanwhile, the interpolation position and range of the synthesized sample are adjusted in a self-adaptive mode along with the increase of iteration, and the diversity of minority class samples is increased. And finally integrating T classifier results so as to improve the multi-classification accuracy of the model for the unbalanced data set. Compared with the prior art, the method has the advantages of wide applicability, high classification accuracy of minority class boundary samples and the like.
Owner:ZHEJIANG LULE INTELLIGENT TECHNOLOGY CO LTD