Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

74 results about "Max pooling" patented technology

Max pooling is an operation of taking a tile with a size for example : 2*2 and then taking the maximum value from the values of this tile and moving to another tile not covered and doing the same. good luck.

Lightweight bearing surface defect small target detection method based on improved YOLOv8

The invention provides a lightweight bearing surface defect small target detection method based on improved YOLOv8. According to the method, based on a YOLOv8n network architecture, a dynamic detection head is constructed, an ADown module is introduced, a small target detection layer is additionally arranged, and finally an improved YOLOv8n-DAS model is established; the method comprises the following steps: firstly, constructing a dynamic detection head, fusing scale, space and task perception attention mechanisms, intensifying feature expression ability in all directions, and accurately capturing complex tiny defect features on the surface of a bearing; secondly, the ADown module highlights edge defect information, simplifies the model parameter scale and reduces consumption of computing resources by combining traditional convolution, average pooling and maximum pooling operations; finally, small target detection layers are arranged at the neck and the head of the network, the small target detection layers are added, an extra small target detection head is arranged, shallow details and deep semantic information are deeply integrated, key details are reserved to the maximum extent, and the recognition capacity of small targets is improved. According to the method, the detection performance of the model on irregular and tiny defects and the detection capability of the model on small target defects are remarkably improved, the parameter quantity and the calculation complexity of the model are reduced, and deployment and application on lightweight equipment with limited resources are facilitated. And the method can be expanded to defect detection of industrial parts such as gears and blades by replacing training data, has the characteristics of light weight and high universality, and meets multi-scene deployment requirements.
Owner:EAST CHINA UNIV OF TECH

Artificial intelligence-based teenager mental health development track prediction system

The invention relates to the technical field of artificial intelligence prediction, in particular to an artificial intelligence-based teenager psychological health development trajectory prediction system, which comprises a data acquisition module, an emotion recognition module, a trajectory prediction module, a risk assessment module and an intelligent feedback module, and provides a feature extraction method based on a point cloud structure. Time sequence feature modeling is carried out by combining K nearest neighbor with Transform, and global feature extraction is carried out by using position sensing optimization and Max Pooling, so that key features of the mental health state of the teenagers are extracted more comprehensively; according to the method, a clustering-based bidirectional LSTM is combined with a GRU prediction model, firstly, DBSCAN is used for clustering, different mental health state modes are defined, Bayesian optimization and Bayesian filtering are adopted for improving the prediction precision, and finally more accurate mental health development trajectory prediction is provided.
Owner:武夷学院

Mechanical fault diagnosis method and system based on deep learning

The invention relates to a mechanical fault diagnosis method and system based on deep learning. The method comprises the following steps: converting a multi-source time domain signal into a time frequency image through continuous wavelet transform, extracting features by using a primary feature encoder, and extracting cross-source common features through adversarial training of a shared feature discriminator; therefore, a gated multi-scale encoder is guided to enhance common feature expression, and deep fusion of multi-source features is realized through a cross multi-head attention network. Global average pooling and maximum pooling are synchronously carried out on the fused features to give consideration to overall and local information, and a comprehensive feature vector is formed; and finally, by means of a double-branch diagnosis network, the training loss of the multi-class network is dynamically weighted according to the prior probability output by the binary network, so that multi-source information is effectively fused under the condition of data imbalance, and the accuracy and robustness of fault classification are remarkably improved.
Owner:NAVAL UNIV OF ENG PLA

Abnormal driving behavior identification method based on multi-modal sensor data fusion

The invention belongs to the technical field of abnormal driving recognition, and particularly relates to an abnormal driving behavior recognition method based on multi-modal sensor data fusion, and the method comprises the steps: obtaining a road image in front of a vehicle and vehicle speed data; the method comprises the following steps: respectively converting vehicle speed data by applying short-time Fourier transform (STFT), a Markov transition field (MTF), a Gramb angle sum field (GASF) and a Gramb angle difference field (GADF), and splicing conversion results to obtain a speed spectrogram; constructing a feature extraction sub-network based on an asymmetric space attention convolution block, an asymmetric channel attention convolution block, a stride convolution layer, a maximum pooling layer and a full connection layer; respectively carrying out feature extraction on the road picture and the speed spectrogram; and carrying out double-flow feature fusion on the features extracted from the road image and the speed spectrogram, carrying out identification by using a classifier based on the obtained fusion features, and outputting a classification prediction result of the driving behavior. According to the invention, the accuracy and robustness of abnormal driving behavior recognition are improved.
Owner:SHANDONG INST OF BUSINESS & TECH

Underwater image enhancement method and system based on combination of multiple convolution models and attention mechanism

The invention belongs to an image processing technology, and particularly relates to an underwater image enhancement method and system based on combination of multiple convolution models and an attention mechanism, and the technology is formed by combination of a DIAM model and a Shallow-UWnet improved model. The DIAM model is mainly used for carrying out color restoration in a mode of complementing hue weakening and non-uniformity of the underwater image, and can solve the problems of underwater noise, atomization and the like. According to the Shell-UWnet improved model, on the basis of Shell-UWnet, a highest pooling layer and an Inc multi-convolution module are added, the highest pooling layer effectively reduces network parameters and calculation overhead through dimension reduction operation, the operation efficiency is improved, the key feature extraction capacity is enhanced, detail loss is avoided, Inc can extract features from different scales, and the algorithm is simple and convenient to operate. And the feature capture capability of the model is enhanced.
Owner:QINGDAO UNIV OF SCI & TECH

High-precision lightweight method for ship target detection of synthetic aperture radar

The invention relates to the technical field of ship detection, in particular to a high-precision lightweight method for ship target detection of a synthetic aperture radar. The high-precision lightweight method for synthetic aperture radar ship target detection comprises the following steps: constructing a lightweight backbone network: based on a YOLOv8 algorithm, introducing a GCADDown module and an ERC2f module, and the GCADDown module reducing the calculation amount through the combination of average pooling, maximum pooling and Ghost convolution. According to the invention, by reducing detection heads, an existing YOLOv8 algorithm is enabled to better fit a radar ship image target detection task, but precision loss caused by reduction of the detection heads is made up by providing an EDWR module, the loss is made up, a light-weight and high-precision model is constructed, and the detection efficiency is improved. The complexity of the algorithm is further reduced, and meanwhile the high precision of the algorithm can be kept.
Owner:DATA SPACE RES INST

Medical image segmentation system and method combining position and channel double attention

The invention discloses a medical image segmentation system and method combining position and channel double attention, the system comprises a feature encoder and a feature decoder, the ResNet50 part of the feature encoder comprises a standard convolution layer, a group normalization layer, a ReLU activation function, a maximum pooling layer and three different stage blocks; the ReLU activation function and the output features of the first two stage blocks are respectively connected to cascade up-sampling units of different levels of a feature decoder through a double attention block so as to reconstruct a feature map; the output feature of the third stage block is connected to a Transform encoder part of the feature encoder through a double attention block; the double attention block extracts features of sparse codes from the angles of positions and channels; a coding block of the Transform encoder adopts a space reduction attention layer to replace a multi-head self-attention layer. According to the method, the robustness of the model is enhanced, the sensitivity of the model to overfitting is effectively reduced, and the overall performance and generalization ability of the model are improved.
Owner:FUJIAN UNIV OF TECH

High-fidelity point cloud completion method and system based on double-path attention and fractal structure

The invention relates to a high-fidelity point cloud completion method and system based on a double-path attention and fractal structure. The method comprises the following steps: processing an input point cloud by an iterative farthest point sampling algorithm; extracting and constructing a joint feature vector by combining an extended combined multi-layer perceptron with a double-path attention mechanism; reconstructing missing region point clouds in a layered manner by using a pyramid point generator; designing a composite loss function and a dynamic weight scheduling strategy to optimize a reconstruction target; designing an adversarial training framework of a discriminator based on a built-in spectrum normalization and gradient penalty mechanism; and reasoning to generate a point cloud reconstruction result. According to the method, the problem that a traditional point cloud completion method depends on geometric prior and a matching template is solved, the limitation that calibration on semantic information of different dimensions is lacked, excessive dependence on maximum pooling is achieved, and the uniformity of discriminator gradient anomaly and a loss function is difficult to guarantee is relieved; and the geometric detail recovery capability and the visual authenticity in a complex structure scene are obviously improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method for thermal video surveillance based on feature pooling module

The present invention relates to a method for thermal video surveillance based on an Encoder-Decoder-induced feature pooling module. The method is explained as follows: An input thermal image that is to be processed, by a pre-trained ResNet-152 deep learning network, wherein the network comprises several convolutional layers, batch normalization layers, and a rectified linear unit (ReLU) function to extract features at low, mid and high levels, and the in-depth target features; receiving, by a feature pooling module (FPM), from the deep learning network, wherein the feature pooling module comprises of a max pooling layer, a convolutional layers, and various atrous convolutional layers for extracting the target features in higher dimensional features space at multi-scales; obtaining higher dimensional features space by a decoder network, wherein the decoder network is configured to project the higher dimensional features into image space for the generation of a probability mask.
Owner:GHOSH ASHISH

Semantic and structure preserving-based point cloud adaptive downsampling method and device

The invention discloses a semantic and structure preserving-based point cloud adaptive downsampling method and device, and belongs to the field of computer point cloud analysis and feature learning. Firstly, point cloud features are extracted, a local neighborhood is constructed, and semantic features and space coordinates of all points are obtained; counting the number of times of selection of the feature channels in the neighborhood based on maximum pooling, and obtaining a local importance score through normalization; through cross-neighborhood aggregation and in combination with spatial distance attenuation weight, geometric consistency is enhanced, and a global importance score is generated; a lightweight multi-layer perceptron is used for fusing semantic and spatial features to predict a comprehensive importance score, and key points are selected according to the score to form a down-sampling subset; and finally, a teacher-student self-supervised training framework is adopted, and a high-confidence-coefficient pseudo tag is generated through a teacher branch to guide student branch parameter optimization, so that efficient reasoning is realized. According to the method, the semantic information and the geometric structure of the point cloud can be effectively kept while the data volume is remarkably reduced.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A three-dimensional city natural landscape point cloud online processing system based on deep learning

The application discloses a ground three-dimensional laser scanning point cloud collection system and relates to the field of three-dimensional point cloud data processing of the ground; the size of data input is 10 square meters; a KPConv algorithm is improved; the characteristics of each input point are expanded; Max Pooling is used to aggregate the relative position and the Euclidean distance of the field points and the center point; Average Pooling is used to aggregate the field characteristics of each point extracted through the multi-layer KPConv into global characteristics; and a set of point cloud visualization websites is deployed by using WebGL.
Owner:SICHUAN AGRI UNIV

3D point cloud classification segmentation method based on dynamic edge convolution and residual error double attention

The invention relates to the technical field of image processing, in particular to a 3D point cloud classification segmentation method based on dynamic edge convolution and residual double attention, which comprises the following steps: acquiring an image to be processed; a 3D-AGCN network is constructed; feature alignment is carried out by using a Transform encoder, feature groups are obtained through a plurality of EADEC, RDA and AMFE modules, the feature groups are connected in series, and then convolution compression is carried out to obtain shared features; the classification branch performs global maximum pooling and average pooling on the shared features, and outputs classification scores; and the segmentation branches repeat the shared features and the one-hot category vectors to each point and then splice the shared features and the one-hot category vectors, and a segmentation score is output. The method solves the problem that an existing method still needs to be improved in the aspects of effectively fusing local details and global contexts and optimizing model performance.
Owner:CHANGZHOU UNIV

Distraction driving detection method based on depth separable convolution and multi-spectrum attention

The invention provides a distracted driving detection method based on depth separable convolution and multi-spectrum attention, and the method comprises the following steps: firstly, carrying out the efficient feature extraction of an input driver image through a feature extraction module based on depth separable convolution, so as to reduce the calculation complexity of a model and the number of parameters; secondly, a multi-spectrum attention mechanism is designed and integrated, and the capturing capability of the model on key features is enhanced through self-adaptive attention on different spectrum information, so that the accuracy of distraction driving detection is improved; secondly, carrying out diversified processing on training data by adopting a data enhancement technology so as to improve the generalization ability of the model; and finally, constructing a deep neural network architecture containing a maximum pooling layer and a classifier, and realizing real-time detection and classification of distraction driving behaviors. Compared with the prior art, the method has the following advantages: 1) by introducing the depth separable convolution, the calculation amount and the parameter scale of the model are significantly reduced, and the training and reasoning efficiency is improved; 2) a multi-spectrum attention mechanism enhances the attention capability of the model on key features, and improves the accuracy of distraction driving detection; 3) an optimized data enhancement strategy improves the generalization performance of the model and reduces the risk of overfitting; and 4) the whole model is simple in architecture and convenient to deploy and expand in practical application.
Owner:GUILIN UNIV OF ELECTRONIC TECH

UAV Detection Method Based on Residual Network Multi-View Feature Fusion

This invention discloses a UAV detection method based on multi-view feature fusion using residual networks, comprising: constructing multi-view data: obtaining the time-domain plot, short-time Fourier transform plot, continuous wavelet transform plot, and Wegener-Will distribution plot of the measured signal; constructing a ResNet34, including an input structure, an intermediate structure, and an output structure; the input structure processes the input data through convolution and max pooling operations; the intermediate structure consists of four similar structural layers, each consisting of multiple residual blocks, each residual block containing three convolutional layers and a shortcut connection; starting from the second structural layer, the initial residual block of each structural layer also contains an up-dimensional sampling structure; constructing a multi-view feature fusion network model based on residual networks, and outputting the UAV detection accuracy after multi-view feature fusion at the output end. This invention improves the UAV detection efficiency by fusing features from the multi-view data of the signal.
Owner:GUILIN UNIV OF ELECTRONIC TECH

A method for reconstructing building point clouds based on an improved KNN-DGCNN model

This invention discloses a method for reconstructing building point clouds based on a DGCNN model with an improved KNN algorithm. The method includes: normalizing the original building point cloud to be reconstructed to obtain normalized point cloud data; constructing a DGCNN network based on the improved KNN algorithm and training the DGCNN network to obtain a trained DGCNN model. The DGCNN network based on the improved KNN algorithm includes a spatial transformation layer, four graph convolutional layers, a max pooling layer, a first multilayer perceptron, and a second multilayer perceptron connected sequentially; and inputting the normalized point cloud data into the trained DGCNN model to obtain the corresponding prediction results. This invention utilizes the local update mechanism of the KD tree to efficiently and dynamically adjust the adjacency graph during network training, avoiding the high computational cost of reconstructing the entire search tree.
Owner:WUHU RES INST OF XIAN UNIV OF ELECTRONIC SCI & TECH +1

A real-time high-resolution portrait matting method based on deep neural network

The application discloses a kind of real-time high-resolution portrait matting methods based on deep neural network, including obtaining training dataset, and marking generation training groundtruth alpha matte;Training data set is data enhanced;Network model is trained in step phase;Using trained network carries out matting.Through embedding ConvLSTM module in network configuration, using Max Pooling Indices, high-definition detail optimization is carried out using PRM, semantic segmentation task is added, and the core technology of high-precision real-time portrait matting is created, simultaneously, data set and data enhancement method are innovated, and training is carried out in stages, from simple to complex, from rough to fine, the training effect of algorithm is strengthened, the innovation and application of the three aspects interact, mutually unified, the performance and practicality of algorithm are comprehensively improved, and powerful technical support is provided for high-precision real-time portrait matting application.
Owner:SHENZHEN CHAOYUAN CREATION TECH CO LTD

An Image Segmentation Method Based on the Global Perception U-Net Model

The present invention discloses an image segmentation method based on a global perception U-Net model, specifically relating to the technical field of image segmentation. The technical key points are as follows: The global perception U-Net model includes an encoder, a decoder, a CFGC module, and a CSCP module. The first layer of the encoder adopts a first convolution module and a max pooling module connected in sequence. The remaining four layers of the encoder all adopt a CFGC module and a max pooling module connected in sequence. The last layer of the decoder adopts a second convolution module. The remaining four layers of the decoder all adopt a transposed convolution module, a CSCP module, and a CFGC module connected in sequence. The first four layers of the encoder are respectively connected to the first four layers of the decoder by skip connections; The CFGC module includes H a dimensional feature selection unit and W a dimensional feature selection unit, which is used to select feature information in the H and W dimensions at different scales, and based on H a dimensional feature selection unit and W a dimensional feature selection unit to perform feature enhancement on the global features in the form of matrix multiplication.
Owner:SOUTHWEAT UNIV OF SCI & TECH

A Deep Pulse Neural Network-Based ECG Classification Method Based on Attention and Integer Training Pulse Inference

This invention provides a deep spiking neural network method for ECG classification based on attention and integer training pulse inference, comprising the following steps: acquiring raw ECG signal data and preprocessing the raw ECG signal data; performing three rounds of convolution and corresponding max pooling on the ECG data features; performing block-based local self-attention processing on the ECG data features to obtain feature associations in local regions of the ECG data features; performing global self-attention processing on the ECG data features to obtain feature associations across the entire sequence of the ECG data features; performing two rounds of convolution and corresponding max pooling on the ECG data features; classifying the integrated ECG data features using a classification head, and outputting the ECG data classification result. This invention can effectively mine long-range dependency information of ECG data and retain shallow feature information through residual fusion, avoiding the gradient decay problem in deep networks, thereby enabling deep spiking neural network learning.
Owner:SHENZHEN INST OF ADVANCED TECH

3D Object Detection Method with Multimodal Input and Spatial Partitioning

The present invention discloses a three-dimensional object detection method for multimodal input and spatial partitioning, and proposes the following targeted strategies: using the original point cloud data and the RGB three-channel color image as multimodal input; partitioning the space of the original point cloud data, indexing the point cloud grouping row by row and column by column, randomly sampling K points, extracting features, and using the max pooling layer to reduce the dimension of the obtained K local-global feature vectors; slicing the RGB three-channel color image, indexing the slices row by row, and inputting them into the two-dimensional feature extractor VGG16 to only extract the shallow-layer relevant features of the texture color of the 8th layer, obtaining K color texture feature vectors; fusing the local-global feature vectors and the color texture feature vectors to obtain the fused feature vectors; passing through the fully connected layer, outputting the prediction results, and drawing the BBox according to the confidence level to complete the post-processing task. The present invention reduces the computational amount and improves the accuracy of classification and detection.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Max pooling method and apparatus for protecting privacy data

The embodiment of the specification provides a maximum pooling processing method and device for protecting privacy data. The method comprises the following steps: based on a local slice of an input matrix, a local slice of a first comparison matrix is determined by multiple parties, which indicates a comparison result of horizontally adjacent elements in the input matrix; based on the local slice of the first comparison matrix, a local slice of an intermediate result matrix is determined by multiple parties, an element of the intermediate result matrix being a larger value of horizontally adjacent elements in the input matrix; based on the local slice of the intermediate result matrix, a local slice of a second comparison matrix is determined by multiple parties, which is used for indicating a comparison result of vertically adjacent elements in the intermediate result matrix; and based on the local slice of the second comparison matrix, a local slice of a pooling result matrix is determined by multiple parties, an element of the pooling result matrix being a larger value of vertically adjacent elements in the intermediate result matrix. The communication overhead can be reduced, and the overall calculation efficiency can be improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Point cloud semantic segmentation method based on stage information fusion transformer and grouping normalization

The invention discloses a point cloud semantic segmentation method based on stage information fusion transformer and grouping normalization, and aims to improve the accuracy of point cloud semantic segmentation and reduce the operation cost. A traditional U-net network is adopted to divide the process into a coding stage and a decoding stage for five times, coding firstly adopts core point convolution to extract initial features of point clouds, then downsampling is carried out, features before and after sampling and grouping are utilized to carry out mutual normalization, then the features are linearly combined, and finally, the initial features of the point clouds are extracted; and finally, carrying out adaptive combination of maximum pooling and learnable weight pooling. And then transform attention operation is carried out, information of different coding stages is fused into an attention result, operation opposite to coding is carried out during decoding, attention calculation is carried out after up-sampling is carried out, information fusion is added, and finally segmentation is carried out.
Owner:HEBEI UNIV OF TECH

A bridge bending detection method based on a lightweight multi-scale sparse gating network

This invention discloses a bridge bending detection method based on a lightweight multi-scale sparse gating network, relating to the field of bridge structural health monitoring. The method includes: acquiring multimode fiber speckle images corresponding to different bending states of the bridge and preprocessing them to obtain standardized speckle images; obtaining an initial feature map based on a lightweight multi-scale sparse gating network through initial convolutional layers and max pooling layers; extracting and fusing bending-sensitive features through a multi-level multi-scale feature fusion module to obtain a multi-scale fused feature map; using a learnable sparse gating module for feature selection to obtain a sparse enhanced feature map; modeling global dependencies through a global context enhancement module to obtain a globally enhanced feature map; and inputting the globally enhanced feature map into a dual-task prediction module to output the corresponding result. This method achieves synchronous, lightweight, and high-precision detection of bridge bending degree and location, effectively decoupling the problem of multi-parameter cross-sensitivity.
Owner:TIANJIN UNIVERSITY OF TECHNOLOGY

Method and apparatus for max pooling of convolutional neural networks

The present application relates to the field of computer vision and artificial intelligence, and specifically to the processing of the pooling layer of the convolutional neural network, and proposes a method and device for large-core pooling. The method mainly includes the HBLK / WBLK block mode of mapping the pooling output to the input, including: performing the operation of the pooling core kh / kw length characteristic data on the H / W direction dimension, saving the temporary results of the direction in the internal cache SRAM, storing the temporary data in the maximum output mode, and when the internal cache size range is exceeded, the H / W direction block of cblk is performed; the W and H directions represent the width and height directions of the feature map; kw / kh are the sizes of the W / H direction pooling core respectively. The device includes a pooling top layer control device, an input / output control device, an HBLK unit control device, an HBLK unit control device, and a pooling operation device. The present application can realize the operation of the pooling core with unlimited size, the hardware device proposed can achieve the balance of performance, bandwidth, power consumption and area, and can effectively solve the technical problems such as small pooling core size and high redundancy of tensor characteristic data operation in the prior art.
Owner:EEASY TECH CO LTD

A pipe network connectivity assessment method and device based on a DS-PIGNN and related equipment

This application relates to the field of urban pipeline network analysis technology, and particularly to a pipeline network connectivity assessment method, device, and related equipment based on DS-PIGNN. The method includes: acquiring pipeline network data, constructing a macroscopic topology map, and collecting the axial microscopic physical field sequence of each pipeline; extracting the features of the microscopic sequence of each pipeline using a one-dimensional residual convolutional neural network, and obtaining a microscopic health state vector through global max pooling; injecting this vector into the edge features of the macroscopic topology map to drive a graph neural network trained with a physical constraint loss function (including Kirchhoff flow conservation), which dynamically calculates the attention weights between nodes based on the hydraulic features of the nodes and the microscopic health state vector; and performing information aggregation based on this dynamic weight, simultaneously outputting the pipeline failure probability and node connectivity reliability. This application achieves dynamic and coupled analysis of microscopic physical damage details and macroscopic network cascade failures, and ensures the physical reliability of the assessment results while maintaining computational efficiency.
Owner:CHINA THREE GORGES CORPORATION

A method and system for extracting feature points from point clouds based on self-attention mechanism

This invention discloses a method and system for extracting feature points from point clouds based on a self-attention mechanism. The method includes: acquiring point cloud slices of a point cloud model; inputting the point cloud slices into a neural network to obtain multi-channel feature neighborhoods; performing MLP calculations on the multi-channel feature neighborhoods and then applying a self-attention mechanism to the calculation results to obtain global features; and sequentially performing max pooling, MLP, and FNN calculations on the global features to obtain the probability that the center point of the point cloud slice is a feature point. This invention obtains multi-channel feature neighborhoods from point cloud slices. In addition to the spatial location information of the point cloud, the multi-channel feature neighborhoods also include Euclidean distance information and center point neighborhood information, thus obtaining more semantic information. Combined with self-attention mechanism calculations and post-processing, the dimensionality of the output is reduced, decreasing the computational load of subsequent feature mapping; this allows for convenient and efficient acquisition of feature points.
Owner:NANJING UNIV OF POSTS & TELECOMM

A Neural Network Image Recognition Method and Electronic Device Based on Pulse Statistics

The present invention discloses a neural network image recognition method and an electronic device based on pulse statistics, including: training a CNN network using training image data to obtain a trained CNN network, and extracting network parameters of the convolutional layer and the fully connected layer in the CNN network; converting pixel values of an image to be recognized into a pulse sequence by using a pulse excitation frequency encoding method; using an SNN network to recognize the pulse sequence to obtain an image recognition result; wherein, the convolutional layer and the fully connected layer in the SNN network adopt the same network structure and network parameters as the convolutional layer and the fully connected layer in the CNN network; the max pooling layer and the Softmax layer in the SNN network adopt the same network structure as the max pooling layer and the Softmax layer in the CNN network, and the output corresponding to the max pooling layer and the Softmax layer in the SNN network is realized by using a pulse number statistics method. The present invention can meet the image recognition requirements of existing portable devices with lower power consumption.
Owner:XIDIAN UNIV

A multi-category building material video counting method, system and counting device

The present invention provides a multi-category building material video counting method, system and counting device. The counting method includes: extracting a to-be-detected image of a video captured by a robot; inputting the to-be-detected image into a YOLOv4 model to extract features of the to-be-detected image; after performing three convolutions on the last feature layer of the backbone feature extraction network, using multi-scale max pooling processing to separate the context features in the to-be-detected image; performing multi-scale prediction on the obtained features, and decoding to obtain the positions of the prediction boxes in the to-be-detected image; inputting all box information into an NMS module to obtain the filtered box information; inputting the box coordinate sequences of the front and rear frames in the frame sequence output by the target detector into a sort tracking module to output the inter-frame target ids. The present invention adopts a neural network method and uses a multi-category multi-object tracking to associate the inter-frame information of the video, overcome target occlusion, and finally calculates the quantity and types of building materials in the entire video through a double-crossing counting algorithm.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN) +2

Coarse-to-fine image matching method based on aggregation attention mechanism

The invention discloses a coarse-to-fine image matching method based on an aggregation attention mechanism, and belongs to the technical field of computer vision. The method comprises the following steps: firstly, extracting multi-scale features of an input image through a lightweight heavy parameterized convolutional neural network; then, an aggregation attention module is used for efficiently converting coarse-grained features, and the module is used for aggregating tokens through deep convolution and maximum pooling and enhancing feature discrimination ability in combination with rotation position coding; then calculating a similarity matrix based on the converted features, and obtaining rough matching point pairs through double softmax operation; and finally, by taking rough matching as guidance, realizing sub-pixel-level accurate positioning on fine-grained features through a two-stage process of mutual nearest neighbor screening and space expectation calculation. According to the method, the calculation efficiency is remarkably improved while the high matching precision is guaranteed, and the method is suitable for scenes such as unmanned aerial vehicle visual positioning and navigation which have strict real-time requirements.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A small target detection method based on PBAF-YOLO progressive boundary perception and multi-scale fusion

This invention provides a small target detection method based on PBAF-YOLO with progressive boundary awareness and multi-scale fusion. The method involves inputting a feature map, calculating neuron energy values ​​based on the spatial mean and unbiased sample variance of each channel's feature values, generating spatial attention weights based on these neuron energy values, and then weighting and enhancing the input feature map to obtain boundary enhancement features. These boundary enhancement features are grouped along the channel dimension, and average pooling and max pooling are performed on each group in the horizontal and vertical directions respectively. These are then fused to generate direction-sensitive spatial attention weights, resulting in multi-scale enhancement features. Multi-scale feature maps output from different stages of the backbone network are reused, and feature fusion is performed through single downsampling and single upsampling paths with skip connections. The high-resolution features are iteratively optimized to generate a multi-scale detection feature map. This multi-scale detection feature map is then input into a high-resolution prediction head adapted for small targets, outputting the small target detection result.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY