Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

111 results about "Class prediction" patented technology

Parasite ovum microscopic image detection method and system based on polymorphic prior

The invention relates to the technical field of medical image processing and computer vision, in particular to a parasitic ovum microscopic image detection method and system based on polymorphic prior, and the method comprises the steps: obtaining a to-be-detected microscopic image, and carrying out the feature extraction of the to-be-detected microscopic image through a convolutional neural network, and obtaining an initial feature map; constructing a polymorphic convolution kernel library based on preset biological morphological characteristics of the parasitic ova; performing deep convolution and feature fusion operation on the initial feature map by using a polymorphic convolution kernel library to generate a space attention map; performing feature enhancement processing on the initial feature map by using the spatial attention map to obtain an enhanced feature map; and performing bounding box regression and category prediction on the enhanced feature map to obtain a parasitic ovum detection result. According to the method, morphological priori and attention mechanisms are introduced, so that the problems of egg form similarity, background interference and the like are solved, accurate and robust automatic detection is realized, and the clinical diagnosis efficiency is remarkably improved.
Owner:SHANGHAI INSTITUTE OF INFECTIOUS DISEASE & BIOSECURITY

Unmanned aerial vehicle target detection method based on long-tail distribution optimization

The invention provides an unmanned aerial vehicle target detection method based on long-tail distribution optimization. The method comprises the steps of performing data fusion on acquired multi-modal data; inputting the fused data into a feature extraction network for feature extraction, and performing adaptive reweighting of channels and spaces during feature extraction; dynamically adjusting an anchor box according to the extracted feature map by adopting a regional proposal network to generate a candidate box; calculating a joint confidence coefficient according to the candidate box and the feature map so as to optimize a sampling strategy of a category bias sampler; respectively processing the sampled head feature frame and the tail feature frame to obtain a corresponding category prediction probability and a bounding box offset, and decoding a prediction depth value by adopting a depth decoder according to the candidate frame and the feature map; according to the category prediction probability, the bounding box offset and the prediction depth value, constructing a total loss function to perform model training; outputting a detection result by adopting the trained target detection model; therefore, the detection sensitivity of low-frequency and high-risk targets in the construction site is remarkably improved.
Owner:XIAMEN UNIV

Small target vehicle detection method and related equipment

The invention discloses a small target vehicle detection method and related equipment. The method comprises the following steps: acquiring image data; inputting the obtained image data into a trained small target vehicle detection model, and outputting a detection result; the small target vehicle detection model comprises a backbone network, a neck network and a detection head; the backbone network fuses a convolution module and an attention mechanism on the basis of a YOLOv8n network, and is used for carrying out feature extraction on an image and providing feature information for a subsequent neck network; the neck network is additionally provided with an up-sampling module on the basis of a YOLOv8n network, feature map size alignment is carried out through an up-sampling and down-sampling bidirectional path, and feature map information is fully fused; the detection head is used for dividing feature information transmitted by the neck network into two paths and processing the two paths respectively, and completing two tasks of bounding box prediction and category prediction at the same time. According to the method, the feature extraction and detection precision of the model on the small target vehicle is effectively improved, and the strict perception performance requirement in an automatic driving scene is met.
Owner:SOUTH CHINA UNIV OF TECH

Classification using multimodal large language models

Methods, systems, and apparatus for classification. In one aspect, a method includes receiving an input and a request to classify the input into one of a plurality of classes, processing the input using a multimodal model to generate (i) a description of the input and (ii) a class prediction, processing the description of the input and the class prediction using a text encoder embedding neural network to generate a (i) text description feature embedding and (ii) a prediction feature embedding, generating, from at least the description feature embedding and the prediction feature embedding, a query feature embedding representing the input, and classifying the input into one of the plurality of classes using the query embedding.
Owner:GOOGLE LLC

A vehicle re-identification method and device guided by camera topology map

The present invention discloses a vehicle re-identification method and device guided by a camera topology map. The method comprises: constructing a training set to obtain vehicle feature representations; constructing a camera topology map based on the vehicle feature representations; constructing a topological relationship between the feature representations of any two vehicles based on the camera topology map and inputting the relationship into a graph convolutional network to obtain final aggregated features; fusing the final aggregated features with the vehicle feature representations, and inputting the fusion result into a fully connected layer for class prediction; constructing a target loss function, training the graph convolutional network, and stopping the training until the target loss function value is minimized to obtain a trained graph convolutional network; and performing vehicle re-identification using the trained graph convolutional network. The present invention has the advantages of improving the accuracy of re-identification.
Owner:ANHUI NORMAL UNIV

Quantum machine learning method for multi-class classification

The present invention relates to a quantum machine learning method for multi-class classification, and the method comprises the steps of: applying a Quantum Convolution Neural Network (QCNN) quantum circuit to input data having q qubits, and outputting a feature vector based on Pauli-Z measurement; and applying a Quantum Neural Network (QNN) quantum circuit to the feature vector, and outputting a multi-class prediction vector with scalability increased compared to q qubits based on basis measurement.
Owner:KOREA UNIV RES & BUSINESS FOUND

EEG channel selection method and system based on double-branch collaborative guide learning

The invention discloses an EEG channel selection method and system based on double-branch collaborative guide learning. The method comprises the steps that original electroencephalogram signals are collected; the original electroencephalogram signals are input into the channel importance evaluation module through the signal input module, the smooth score of each channel is calculated, and a set number of optimal channels are selected; the signal feature extraction module extracts the space-time features of the optimal channel and pools and projects the space-time features to obtain final embedded features; the embedded features are input into the double branches, global average pooling is carried out on the embedded features, and prediction label probability distribution is obtained through multi-class prediction; reconstructing the complete electroencephalogram signals according to the embedded features to obtain reconstructed electroencephalogram signals; constructing a total loss function, minimizing the total loss, and adjusting the weight of the electroencephalogram channel selection model; and performing channel selection by using the trained electroencephalogram channel selection model. According to the method, through a collaborative learning mechanism, on the basis of ensuring the performance, the number of channels required by electroencephalogram signals is remarkably reduced.
Owner:NAT UNIV OF DEFENSE TECH

SDTW-IPAM-based short-term power load prediction method and system, electronic equipment and storage medium

The invention belongs to an SDTW-IPAM-based short-term power load prediction method and system, electronic equipment and a storage medium. The method comprises the steps of data preprocessing, clustering analysis, single-class prediction model construction and future load prediction. The clustering analysis comprises setting of upper and lower limits of a cluster number, construction of a distance matrix, calculation of a Gap value and a standard error, selection of an optimal cluster number, generation of an initial center, clustering processing and normalization processing; according to the method, a distance measurement method is introduced into load clustering, so that the dynamic time sequence characteristics of a load curve can be described more accurately, the local time deformation of the load curve can be effectively identified, and the distinguishing capability of similar load days under the influence of weather, events or potential new energy fluctuation is improved; a clustering algorithm is improved by using statistics and an initialization strategy, so that the stability and the accuracy of performing mode recognition on complex load data are improved, and the load can be effectively divided into typical modes with different dynamic characteristics.
Owner:XINJIANG UNIVERSITY

Automobile part defect identification method and system, computer equipment and storage medium

The invention discloses an automobile part defect identification method and system, computer equipment and a storage medium, and aims to solve the problem that a training sample is limited in industrial defect detection by performing discrete Fourier transform on an original image to obtain a high-frequency structural feature image. The high-frequency structural features of the target in the original image are fused into the original features of the defect, and the feature expression of the target is enriched, so that the deep learning model can fully learn the feature expression of the defect target with fewer training samples, thereby improving the classification performance of the known category, and improving the classification efficiency. The obtained defect feature vector is subjected to known class prediction through a full-connection classification network in the automobile part defect identification method based on high-frequency structure feature fusion enhancement and class mutual information constraint, and whether input sample data belongs to an unknown class is judged; and finally, accurate defect detection and unknown defect identification in industrial defect detection are realized.
Owner:HUNAN UNIV

Universal AIGC image detection method and device based on edge enhancement, electronic equipment and storage medium

The invention provides a universal AIGC image detection method and device based on edge enhancement, electronic equipment and a storage medium, and the method comprises the steps: carrying out the edge perception of an AIGC image, and obtaining an edge image; performing edge artifact enhancement on the AIGC image by using the edge image; wherein the edge artifacts are features generated through up-sampling processing in the process of generating the AIGC image; performing local pixel difference extraction operation in the AIGC image after edge artifact enhancement to obtain a high-frequency information feature map; and inputting the high-frequency information feature map into the trained classification network for class prediction to obtain the authenticity probability of the AIGC image. According to the method, the edge artifact features left in the general up-sampling process of the generative models are used for detection, priori knowledge of the generative models does not need to be obtained, the method does not depend on a specific generative model, the effect of uniformly detecting the generated images of multiple generative models can be achieved, and the cross-domain generalization problem is solved.
Owner:BEIZHI TECHNOLOGY (ANJI) CO LTD

Intestinal polyp identification method based on image and multi-modal data

The invention discloses an intestinal polyp identification method based on images and multi-modal data. The method comprises the following steps: acquiring pathological images of a plurality of microscopes, objective magnification information corresponding to the pathological images, and text information corresponding to enteroscopes; performing data preprocessing on each frame of acquired pathological image and text information to obtain a training data set; a multi-scale multi-modal feature classification network is constructed, the training data set is used for training, and a trained multi-scale multi-modal feature fusion classification network is obtained; the method comprises the following steps: collecting a microscope pathological image of a patient, obtaining corresponding enteroscope text information, respectively carrying out data preprocessing, and inputting into a multi-scale multi-modal feature fusion classification network to obtain disease category prediction of a model. According to the method, pathological images and enteroscope text information are fully fused, accurate classification and prediction of disease types are finally achieved, the effect and efficiency of multi-modal diagnosis of intestinal diseases are improved, and therefore the accuracy of patient diagnosis is improved.
Owner:ZHEJIANG UNIV

Traffic target detection method, system and model architecture under severe weather condition

The invention belongs to the technical field of traffic target detection, and discloses a traffic target detection method, system and model architecture under severe weather conditions. The method comprises the following steps: preprocessing an input traffic scene image; performing multi-scale feature extraction on the preprocessed image, and obtaining multi-scale feature representation through AMSConv operation; carrying out bidirectional fusion on the extracted multi-scale features to obtain fusion features; the fusion features are processed, and category prediction, bounding box coordinates and confidence of the target are generated; and optimizing bounding box regression based on InfoAware-IoU measurement, and dynamically weighting a high information sample. Through multi-scale feature extraction, efficient detection head design and information perception regression optimization, the problem of visual degradation under severe weather conditions is effectively solved, and the accuracy and real-time performance of traffic target detection are improved.
Owner:TECH TRAFFIC ENG GRP CO LTD +2

A visual language model test adaptation method based on logarithmic calibration and consistency cache

The application relates to a visual language model test adaptation method based on logarithmic calibration and consistency cache, and relates to the technical fields of computer vision, pattern recognition, machine learning and artificial intelligence. The method aims at the problems of class prediction bias, insufficient cache sample coverage and low utilization rate of boundary samples of a visual language model in a target domain test stage, and constructs an online adaptation framework composed of an image encoder, a text encoder, a dynamic logarithmic calibration module, a consistency guide exploration cache module and a cross-modal joint optimization module. The method improves the identification opportunities of difficult classes and low-frequency classes through a dynamic logarithmic calibration mechanism, and improves the overall class balance and classification stability. Through the consistency guide exploration cache mechanism, the coverage range of the cache to the real distribution of the target domain is expanded, and the perception ability of the model to the decision boundary region is enhanced, so that the adaptation effect in a complex distribution offset scene is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Multi-scale feature representation enhancement method and medium for underwater sonar target detection

The invention discloses a multi-scale feature representation enhancement method and medium for underwater sonar target detection, and the method comprises the steps: 1, constructing a lightweight composite feature extraction backbone network of a target detection model, and extracting a three-level scale feature map; 2, obtaining a five-level scale feature map through up-sampling and down-sampling, fusing semantic features of a top layer to a lower layer, and enhancing context semantic information; 3, a multi-dimensional dynamic attention module is added through feature fusion and input, and attention is enhanced for key features from three dimensions of scale, space and channel; 4, a detection module carries out dense prediction on each layer of features fused by the features of the multi-dimensional dynamic attention module, the dense prediction comprises a frame regression branch, a category prediction branch and positioning quality prediction, and a positioning quality detection branch is improved by adopting central positioning confidence; and 6, total loss is calculated, model parameters are updated through back propagation, a target detection model is obtained after training is ended, and the scheme obtains good detection precision in an underwater sonar target detection task.
Owner:YICHANG TESTING TECHNIQUE RESEARCH INSTITUTE

Multi-class attitude estimation method and system based on shared key point adaptive matching

The invention provides a multi-class attitude estimation method and system based on shared key point adaptive matching, and the method comprises the steps: constructing a large-scale multi-class attitude data set for model training and evaluation, carrying out the processing of the data set, obtaining a training image subset and a test image subset, respectively inputting the training image subsets into a query model comprising an image feature extraction module, a structure prototype classification module, a shared key point prediction module, an adaptive matching module and a model optimization module, training the query model through mutual cooperation of the modules, and optimizing model parameters; and applying the optimized query model to a test image to obtain a structure prototype category prediction result and a key point estimation result of the test image. According to the method, cross-category universal features are learned by sharing key point embedding, and unified and high-precision attitude estimation of mass categories of objects with variable structures is realized by using a prototype exclusive matching mechanism of momentum updating.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

An aerial image rotating target detection method based on annular smooth label

The application relates to a kind of aerial image rotating target detection methods based on annular smooth label, belong to image processing field.The method comprises the following steps: S1: the format of image label is handled using annular smooth label (CSL), so that the angle information in the label has periodicity;S2: for the problem of difficult feature extraction caused by occlusion, small target and other reasons, a global context module is introduced into the feature extraction module of YOLOv5s network, which is improved and a rotating target detection model is built;S3: use GSConv convolution to replace the standard convolution to optimize the Neck module of the network and generate feature maps of different sizes;S4: class prediction and rotating envelope box prediction containing angle information are carried out;S5: calculate classification loss, confidence loss, positioning loss and angle loss;S6: use the improved detection model to train the preprocessed data set;S7: input test set, obtain detection result.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Data classification and recognition method and apparatus, device, and medium

A data classification and recognition method includes: obtaining a first data set and a second data set, the second data set including second data, samples in the second data being labeled; performing training using first data in an unsupervised training mode and using the second data in a supervised training mode to obtain a first classification model; obtaining a second classification model; performing distillation training on a model parameter of the second classification model to obtain a data classification model; and performing class prediction on target data by using the data classification model.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

YOLO-based lightweight attention-guided digestive endoscopy abnormal region detection method

The invention discloses a YOLO-based lightweight attention-guided digestive endoscopy abnormal region detection method. The method comprises the following steps: preprocessing data by adopting an image enhancement strategy, and inputting the preprocessed data into a backbone network to extract features; and the extracted features are subjected to up-sampling, fusion and enhancement on the neck layer and then are input into the detection layer to complete position regression and category prediction of the target area. According to the method, an MDR-YOLO model is developed and is used for abnormal region detection of a digestive endoscopy image, and a proposed MSC2f module comprises a multi-scale parallel convolution attention module used for extracting multi-scale feature information; the provided D4f module comprises a deep semantic expression structure constructed by a multi-level convolution residual path, so that high-quality feature fusion and enhancement are realized; a proposed RSC2f module comprises a channel attention residual module which is used for improving the ability of the model to extract all-channel features. The experimental result verifies the effectiveness of the MDR-YOLO model in the detection of the abnormal region in the digestive endoscopy image, and especially shows higher accuracy and robustness in the detection of a small target.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Endoscopic ultrasound-based lesion invasion depth evaluation system and method, and storage medium

Disclosed in the present invention is an endoscopic ultrasound-based lesion invasion depth evaluation system, comprising: an image acquisition module for acquiring an endoscopic ultrasound image, a model creation module for creating a convolutional neural network model, and a model training module. The convolutional neural network model comprises: a feature extraction module, used for performing layer-by-layer feature extraction across at least three layers on the endoscopic ultrasound image, so as to obtain encoded feature maps; a feature fusion module, used for performing feature fusion on at least three groups of encoded feature maps, so as to obtain a first fused feature map; a decoder module, used for using an attention mechanism to process the encoded feature maps used in the feature fusion module and the feature map obtained by the feature fusion module, so as to obtain multiple decoded feature maps, and used for performing image segmentation on a finally obtained decoded feature map, so as to obtain a lesion region; and a class prediction module, used for performing classification with respect to the lesion invasion level on the basis of the first fused feature map and the decoded feature maps. The present invention can accurately predict lesion regions and lesion invasion depth types.
Owner:TIANJIN YUJIN ARTIFICIAL INTELLIGENCE MEDICAL TECH CO LTD

A method and system for assisting in the diagnosis of high altitude hypertension based on electrocardiogram

PendingCN122350727AFeature extractionMedicine
This invention relates to the field of cardiovascular disease auxiliary diagnosis technology, specifically to a method and system for auxiliary diagnosis of high-altitude hypertension based on electrocardiogram (ECG). It uses 12-lead ECG images as input, constructs a multi-branch adaptive classification model that integrates regional priors, sequentially extracts shallow general features through a shared feature extraction module, combines independent feature branches from multiple leads with a paired spatial attention module to capture relative potential differences between leads, and a multi-path adaptive fusion module dynamically fuses differential features between high-altitude and plain areas using regional labels as gating signals. Finally, a hybrid classifier composed of capsule networks and DenseNet cascades the features to complete a four-class prediction: high-altitude hypertension, high-altitude normal, plain hypertension, and plain normal. This invention effectively overcomes the problem of cross-regional data domain shift, improves the ability to capture weak pathological features in ECG images, and has strong clinical transferability and robustness, making it suitable for auxiliary screening of hypertension across high-altitude and plain areas.
Owner:DALIAN UNIV

Rock Class Prediction Method Based on Multi-Teacher Knowledge Distillation and Normalized Attention

The present invention discloses a rock category prediction method based on multi-teacher knowledge distillation and normalized attention, including: S1: Collect rock sample images, establish a rock data set, and preprocess the original rock sample image data; S2: Insert a normalization-based attention module on the basis of the original network structure to construct a lightweight neural network; S3: Adopt the multi-teacher knowledge distillation method, that is, use the self-attention module represented by Swin-Transformer and the convolutional model represented by ResNet to train a lightweight student model; S4: Divide the data set into a training set and a test set, and record the classification results and various performance indicators. The present invention uses MobileNetV3 as the student model, combines the advantages of the CNN and Self-Attetion paradigms, trains a lightweight model with better performance, obtains the scaling factor through batch normalization, and reflects the change size of each channel through the scaling factor, which also represents the importance of the channel, enabling the implementation of the attention mechanism without additional parameters.
Owner:HOHAI UNIV

A small sample action recognition method based on a graph neural network

A small sample action recognition method based on a graph neural network, comprising the following steps: S1, acquiring a video and extracting video features; S2, remodeling and enhancing the video features to obtain video timing features; S3, performing average pooling operation on all video timing features to form corresponding node features, and using the node features to construct edge features, the node features and the edge features are input into a pre-trained graph network for feature propagation and update, and task-oriented features of a query video and a class support set video are calculated; S4, performing class matching on the task-oriented features to obtain a class prediction value of the query video; S5, identifying the action of the actor in the query video through the class prediction value, wherein the method can accurately and quickly identify the action of the actor in the actual scene with only a small amount of labeled training data.
Owner:ZHEJIANG UNIV

Anti-background interference diatom drowning diagnosis detection method and device

The invention relates to a drowning diagnosis diatom detection method and device resistant to background interference. The method comprises the steps that S1, diatom image data are collected, an image and a format are processed, a double-decoupling gating convolution fusion unit DDG-Conv is constructed, an expert kernel is generated through frequency domain decoupling, and high and low frequency feature fusion is achieved by combining feature translation to expand a receptive field and gating; s2, constructing a fine-grained hybrid encoder FGH-Encoder, extracting multi-scale features by a backbone network, performing global modeling, constructing a feature fusion path by using DDG-Conv, and outputting full-scale encoding features; and S3, constructing a micro-scale refining decoder MSR-Decoder to analyze and predict the full-scale coding features, and outputting a target bounding box and a category prediction result. According to the invention, through frequency domain decoupling, spatial feature translation and gating fusion, the problems of weak feature expression in small target diatom detection and poor detection effect under a complex background are effectively solved.
Owner:GUANGDONG UNIV OF TECH +1

Inter-group conformal scoring fairness with set size calibration

Approaches that intend to reduce disparate impact, particularly those for providing equal coverage sets in conformal prediction, can in fact increase disparate impact for human-in-the-loop systems. To improve these systems, rather than optimizing selection processes for a class prediction set for a confidence level (e.g., a percentage confidence that the correct class is in the class prediction set), a selection process is determined that reduces the set size difference across groups.
Owner:THE TORONTO DOMINION BANK

An open-vocabulary-based photovoltaic panel damage detection method, system and medium

PendingCN122434878AData setEngineering
The application discloses a photovoltaic panel damage detection method and system based on an open vocabulary, and a medium, solves the problems of high model training cost and long training time in the prior art, has the beneficial effects of reducing training cost and improving recognition accuracy, and the specific scheme is as follows: a photovoltaic panel damage detection method based on an open vocabulary, comprising constructing a training data set; performing knowledge distillation pre-training, using the pre-trained open vocabulary target detection model as the teacher model, using the detection model to be trained as the student model, distilling the intermediate layer features and class prediction distribution of the teacher model respectively, and the student model learns basic semantic understanding ability through the first pre-training data set; freezing the knowledge distillation module of the student model, collecting images of photovoltaic panels to be detected and preprocessing, and inputting the fused multi-scale feature map and the updated text embedding vector into the detection output module.
Owner:SHANDONG JIANZHU UNIV

Classification model training method based on hierarchical knowledge migration

The embodiment of the invention provides a classification model training method based on hierarchical knowledge migration. The method comprises the steps of obtaining a universal baseline model with a semantic understanding capability; according to the hierarchical classification structure of the target domain, an industry domain model containing a plurality of sub-classification output heads is constructed, and initialization is carried out by using a general baseline model parameter; constructing a hierarchical constraint loss function, and constraining the parent class prediction probability to be greater than or equal to the child class prediction probability; obtaining domain classification training samples, and respectively inputting the domain classification training samples into the general baseline model and the industry domain model to obtain first and second prediction category probability distributions; calculating domain cross entropy loss based on the second prediction distribution and the real label, calculating knowledge distillation loss based on the difference between the two distributions, and constructing a target loss function in combination with hierarchical constraint loss; and taking minimization of the target loss function as a target training industry domain model. According to the method, the model training efficiency is effectively improved, efficient migration of domain knowledge is realized, and the hierarchical classification accuracy is remarkably improved.
Owner:ZHONGDIAN DATA IND CO LTD +1

Adaptive feature enhancement method, device, equipment and readable storage medium

The application provides a self-adaptive feature enhancement method, device and equipment and a readable storage medium. The method comprises the following steps: performing multi-dimensional feature extraction on input image data to obtain a multi-dimensional feature map, performing fusion enhancement on the multi-dimensional feature map to generate a multi-scale feature map; inputting the multi-scale feature map into a fountain feature enhancement module, performing dimension reduction convolution module calculation at each scale to obtain the class prediction confidence of each detection unit; determining an overall loss function based on the class prediction confidence of each detection unit, and determining the detection result of the input image data according to the overall loss function. The application can effectively improve the problem of insufficient high-dimensional semantic feature point feature cohesion and make the sub-optimal detection unit pay more attention to the target region features in the receptive field and enhance the feature expression, thereby improving the image target detection accuracy.
Owner:AEROSPACE INFORMATION RES INST CAS

Pulse time driven network performance optimization system and method

The invention discloses a pulse time-driven network performance optimization system and method, and the system comprises a sparse scrambling coding module which is used for carrying out the sparse scrambling coding of input data, so as to convert the input data into a pulse sequence with the determined time distribution in the coding time, and obtain a coded pulse sequence; the decoding module is used for inputting the coded pulse sequence into the pulse neural network for reasoning, and continuously monitoring pulse activity of each category of channel of an output layer of the pulse neural network in the reasoning process; when the first pulse of any category of channel is detected, generating category prediction based on the pulse time, the pulse count or the combination thereof of the first pulse; and the coding and decoding driving training module is used for carrying out combined optimization of weights and thresholds on the network or carrying out conversion from an artificial neural network and carrying out threshold calibration on the network based on category prediction. Therefore, by adopting the embodiment of the invention, the recognition robustness of the network on the data set can be obviously improved.
Owner:PEKING UNIV

Weld defect real-time detection method and system based on lightweight neural network

The invention discloses a real-time weld defect detection method and system based on a lightweight neural network. The method comprises the steps of image acquisition, feature extraction, feature fusion, target detection and result output: acquiring a weld image through industrial equipment and performing standardization processing; a StarNet lightweight backbone network is adopted, and star operation is taken as a core to extract multi-scale features; partial convolution, local attention extraction and a global attention mechanism are fused through a C3k2-FasterFusion module, noise is suppressed, and feature representation is enhanced; through shared convolutional layer design of a LeanHead lightweight detection head, defect position regression and category prediction are realized, and a defect position frame and a classification label are output. Through collaborative optimization of the three core modules, model parameters and calculation loads are greatly simplified while high detection precision is guaranteed, and the method has the advantages of being light in weight and high in real-time performance, can be flexibly deployed on various edge calculation devices and is suitable for online real-time detection of weld defects, and the production quality control efficiency is effectively improved.
Owner:SHANDONG XIANGNENG INTELLIGENT EQUIP TECH CO LTD

A method for recognizing sports video actions based on an action granularity grouping structure

The present invention belongs to the field of computer vision and video action recognition, and discloses a method for recognizing sports video actions based on an action granularity grouping structure. A hierarchical grouping structure based on action granularity is proposed, and a lightweight multi-scale spatio-temporal modeling and information fusion mechanism is designed. The steps are as follows: video frame extraction, segmented random frame sampling, video frame preprocessing, selection of a backbone network, insertion of an action granularity grouping module in the backbone network to achieve multi-scale spatio-temporal feature aggregation, use of a fully connected layer and a softmax layer for class prediction, use of cross-entropy loss to train the action classes, and training and verification. By using the present invention, multi-granularity action information can be effectively extracted, which is applicable to the recognition of sports video actions containing multi-level categories, and significantly improves the accuracy of sports video action recognition. As a method for recognizing sports video actions based on an action granularity grouping structure, the present invention can be widely applied to the field of sports video action recognition.
Owner:DALIAN UNIV OF TECH