Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Visual Pattern Recognition" patented technology

Cascade attention-based detection model training and target detection method and system

ActiveCN117036770BCharacter and pattern recognitionFeature extractionVisual Pattern Recognition
The application provides a kind of detection model training and target detection method and system based on cascaded attention, belongs to computer vision, pattern recognition and artificial intelligence technical field, obtains data and pre-processes after, obtains global feature by feature extraction network, the feature of different stages of network is used as output, cascaded attention pyramid network, it uses the feature of above-mentioned different stages as input, can fully exploit the saliency information under the network layer, and gradually progressively filter out the current each layer redundant information using these information.By means of feature pyramid structure, the most discriminative information is passed down, so that the remaining scale feature representation space is more discriminative, providing high-quality features for the region generation network, generating more accurate candidate regions, and improving the convergence stability of the region of interest pooling network;By using the cascaded attention method to explore the hierarchical relationship of each layer of the network, more discriminative feature expression is obtained, and the model detection precision is improved.
Owner:BEIJING JIAOTONG UNIV

Intraoperative gauze counting and tracking method based on visual perception

The invention provides an intraoperative gauze counting and tracking method based on visual perception, and relates to the technical field of visual pattern recognition. A collaborative perception framework integrating multispectral physical fingerprint recognition, visual space-time tracking and adaptive entangled state filtering is constructed; the core problem of target identity confirmation and persistent tracking in the prior art is solved. High-robustness fluorescent fingerprints are introduced to serve as identity anchor points, deep coupling and mutual correction of physical identities and spatial-temporal trajectories are achieved through an adaptive fusion algorithm, and finally absolute identity confirmation and high-precision and high-robustness continuous tracking of each target are achieved. According to the method, the recognition accuracy in a seriously polluted and shielded environment is improved, error accumulation in a long-term tracking process is effectively inhibited, and the self-adaptive capability and reliability of the whole system in a dynamic complex scene are enhanced.
Owner:THE SEVENTH MEDICAL CENTER OF PLA GENERAL HOSPITAL

A visual language model test adaptation method based on logarithmic calibration and consistency cache

The application relates to a visual language model test adaptation method based on logarithmic calibration and consistency cache, and relates to the technical fields of computer vision, pattern recognition, machine learning and artificial intelligence. The method aims at the problems of class prediction bias, insufficient cache sample coverage and low utilization rate of boundary samples of a visual language model in a target domain test stage, and constructs an online adaptation framework composed of an image encoder, a text encoder, a dynamic logarithmic calibration module, a consistency guide exploration cache module and a cross-modal joint optimization module. The method improves the identification opportunities of difficult classes and low-frequency classes through a dynamic logarithmic calibration mechanism, and improves the overall class balance and classification stability. Through the consistency guide exploration cache mechanism, the coverage range of the cache to the real distribution of the target domain is expanded, and the perception ability of the model to the decision boundary region is enhanced, so that the adaptation effect in a complex distribution offset scene is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Ship power equipment state monitoring and fault diagnosis method and system based on multi-source data visualization

The invention provides a ship power system intelligent monitoring and early fault diagnosis method and system based on multi-modal image recognition. The method comprises the following steps: S1, constructing a three-dimensional digital twin base model of monitored power equipment; s2, multiple types of operation time sequence signals of the monitored power equipment are collected; s3, converting the acquired at least one operation time sequence signal into a two-dimensional feature image; s4, mapping and fusing the two-dimensional feature image and a parameterized data stream generated based on the operation time sequence signal to a corresponding spatial position of the three-dimensional digital twin base model, and generating a fused visual training picture used by a machine learning model; and S5, training a visual AI model by using the generated fusion visual training picture, and performing state recognition and fault diagnosis on the fusion visual training picture generated in real time by using the trained model. According to the invention, a complex equipment state monitoring problem is converted into a visual mode identification problem, and ship power equipment monitoring and early, accurate and automatic fault diagnosis are realized.
Owner:NO 703 RES INST OF CHINA SHIPBUILDING IND CORP

A visual perception-based method for intraoperative gauze counting and tracking

This invention provides a visual perception-based method for intraoperative gauze counting and tracking, belonging to the field of visual pattern recognition technology. By constructing a collaborative perception framework integrating multispectral physical fingerprint recognition, visual spatiotemporal tracking, and adaptive entangled state filtering, this invention solves the core challenges of existing technologies in target identification and persistent tracking. By introducing highly robust fluorescent fingerprints as identity anchors and utilizing an adaptive fusion algorithm to achieve deep coupling and mutual correction between physical identity and spatiotemporal trajectory, absolute identification of each target and high-precision, highly robust continuous tracking are ultimately achieved. This improves the identification accuracy in severely polluted and occluded environments, effectively suppresses error accumulation during long-term tracking, and enhances the adaptive capability and reliability of the entire system in dynamic and complex scenarios.
Owner:THE SEVENTH MEDICAL CENTER OF PLA GENERAL HOSPITAL

Video face forgery detection method and system based on time-space domain collaborative information bottleneck

The invention discloses a video face forgery detection method and system based on time-space domain collaborative information bottleneck, and relates to the technical fields of computer vision, pattern recognition, machine learning and the like. The method comprises the steps of collecting a real video and a false video, designing a target function of reconstruction loss by using a video face counterfeiting detection neural network based on a time-space domain collaborative information bottleneck, and training a video face counterfeiting detection model. And during forgery detection, inputting a plurality of face image frames into the trained video face forgery detection model, and judging whether the input video is a forgery video or not. According to the invention, through information optimization based on information bottleneck, the model generalization is improved, and a plurality of randomly combined counterfeiting modes such as identity modification, expression replay, local content tampering and the like can be coped with.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

A multi-prototype representation enhanced visual pattern recognition method for robots

The application provides a robot-oriented multi-prototype characterization enhanced visual pattern recognition method, which can be applied to the fields of artificial intelligence, image processing and signal processing technology. The method comprises the following steps: using a prototype integrated classifier to perform feature extraction on an image training data set for a robot visual pattern recognition task, and obtaining a predicted classification result of a sample image by calculating the pattern similarity between the extracted sample image features and each prototype in a multi-prototype set and by mapping the sample image features; using a label-aware multi-prototype updater to perform target prototype recognition on image data collected in real time by the robot based on the visual pattern recognition task, obtaining a prototype to be updated from the multi-prototype set, performing prototype movement quantitative updating on the prototype to be updated based on the predicted classification result of the sample image, obtaining an updated prototype, and obtaining a classification result of the image data based on the updated prototype.
Owner:SUZHOU INST FOR ADVANCED STUDY USTC +1

Robot-oriented multi-prototype representation enhanced visual pattern recognition method

The invention provides a robot-oriented multi-prototype representation enhanced visual pattern recognition method, which can be applied to the technical fields of artificial intelligence, image processing and signal processing. The method comprises the following steps of: performing feature extraction on an image training data set for a robot visual pattern recognition task by utilizing a prototype integrated classifier, and calculating a pattern similarity between extracted sample image features and each prototype in a multi-prototype set and mapping the sample image features to obtain a multi-prototype pattern recognition task; obtaining a prediction classification result of the sample image; the method comprises the following steps: performing target prototype recognition on image data acquired by a robot in real time based on a visual pattern recognition task by using a label sensing multi-prototype updater, obtaining a to-be-updated prototype from a multi-prototype set, performing prototype movement quantitative updating on the to-be-updated prototype based on a prediction classification result of a sample image, and obtaining an updated prototype; and obtaining a classification result of the image data based on the updated prototype.
Owner:SUZHOU INST FOR ADVANCED STUDY USTC +1

A robust event-driven gait recognition method, system, device and storage medium based on event stream

This invention discloses a robust event-driven gait recognition method, system, device, and storage medium based on event flow, relating to the fields of event vision, computer vision, pattern recognition, and intelligent security technologies. The method includes renormalizing spatial displacement, temporal displacement, and edge length according to a unified scale and robustness scale, recalculating edge attributes after each pooling, introducing a motion intensity index to assist in determining edge reliability, employing continuous reweighting instead of direct edge deletion, and using motion consistency, radius validity, orientation validity, and entropy constraints to weakly supervise edge confidence. A graph convolutional backbone network is used to extract spatial graph features for each time slice, and temporal relationships are jointly modeled through difference and similarity branches. The additive angular interval loss and dynamic center loss are jointly optimized. This invention improves message passing stability, enhances anti-disturbance robustness, overcomes intra-class fluctuations such as cross-viewpoint and low-light conditions, and possesses excellent potential for edge device deployment.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Intelligent flower classification method based on lightweight cross-layer network structure

The invention relates to the technical field of computer vision, pattern recognition and artificial intelligence, in particular to an intelligent flower classification method based on a lightweight cross-layer network structure for constructing the lightweight cross-layer network structure which comprises an encoder, a bridging layer and a decoder. In the bridging layer, flattening and position coding are firstly carried out on the deepest layer features in the encoder, and then the deepest layer features are sent into two layers of lightweight multi-head self-attention Transform Encoder to generate a global semantic sequence T; the decoder recovers the spatial resolution through the up-sampling module in sequence; after each level of up-sampling, a cross-layer GCN-ViT fusion module is firstly entered to obtain a fusion feature Ffuse, and then the fusion feature Ffuse is handed over to a decoding convolution dk for continuous processing; and finally, outputting a classification result through 1 * 1 convolution + Softmax. According to the method, the lightweight cross-layer network model is constructed, the calculation efficiency and deployment flexibility of flower identification are effectively improved, and the visual interaction experience of a user on the experiment process and result is remarkably enhanced.
Owner:CHANGZHOU UNIV

Method and system for judging emotion of passenger in passenger vehicle based on video monitoring

The invention provides a passenger vehicle passenger emotion judgment method and system based on video monitoring, and belongs to the technical field of computer vision, mode recognition and artificial intelligence, and the system comprises a video / audio collection module, a modal confidence estimation sub-module, a modal time alignment module, a feature extraction module, and a modal fusion and emotion judgment module. A deep neural network is adopted for fusion, and interaction and weighted fusion between different modal features are achieved; and finally, inputting the fused feature vector into a classifier, and outputting the emotion category of the passenger and the corresponding confidence coefficient by the classifier. According to the method, the facial expressions and the body language features of the passengers are comprehensively analyzed through the multi-modal fusion technology, and the voice features can be selectively combined, so that real-time and accurate judgment on the emotion of the passengers in the passenger vehicle is realized.
Owner:NANKAI UNIV +1

Multi-scale feature alignment method based on Wasserstein distance

The invention provides a multi-scale feature alignment method based on a Wasserstein distance, and the method comprises the steps: constructing a feature space mapping relation, calculating the Wasserstein distance through employing an improved Sinkhorn iterative algorithm, and carrying out the feature alignment optimization in combination with a self-adaptive step length strategy and entropy regularization constraint. By adopting the method, compared with the prior art, the feature alignment precision is improved by 85.1%, the processing speed is improved by 73.1%, and the alignment precision of 95.6% is still kept under 20% noise interference. The method has the technical effects of high calculation efficiency, excellent alignment precision and strong anti-noise capability, and is suitable for feature alignment tasks in the fields of computer vision, pattern recognition and the like.
Owner:GUIZHOU QIANZHI INFORMATION

Pantograph and catenary arc visual detection method

The application provides a pantograph and catenary arc visual detection method, and belongs to the technical field of computer vision pattern recognition and target detection. A four-stage feature down-sampling module built by using a self-attention mechanism models the global context of an image at each stage, captures the global features of a target and models long-range semantic dependency relationships, and aggregates features and position information from the entire input domain. Meanwhile, the encoder and the decoder are connected through a skip layer connection mode, the low-dimensional shallow semantic information at each stage in the down-sampling is fused with the high-dimensional deep semantic information at each stage in the up-sampling, the detection precision is improved, and the data requirement is significantly reduced. The multi-dimensional global feature fusion network is trained by using a training data set and a stochastic gradient descent method, and the serialized image features are first gradually extracted by four cascaded self-attention modules in the encoder to obtain high-dimensional deep features, and four features with different dimensions are generated. The method is used for pantograph and catenary arc visual automatic detection.
Owner:CHENGDU GUOJIA ELECTRICAL ENG CO LTD