Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2767 results about "Backbone network" patented technology

A backbone is a part of computer network that interconnects various pieces of network, providing a path for the exchange of information between different LANs or subnetworks. A backbone can tie together diverse networks in the same building, in different buildings in a campus environment, or over wide areas. Normally, the backbone's capacity is greater than the networks connected to it.

Solar Azimuth Estimation Method and System Based on Multi-Channel Feature Enhancement and Region-Aware Attention

The present invention relates to a solar azimuth estimation method and system based on multi-channel feature enhancement and region-aware attention, belonging to the technical field of intelligent navigation for low-altitude economy unmanned systems. Aiming at the problem of decreased accuracy in solar azimuth estimation based on polarization images under complex cloud cover conditions, the present invention proposes a deep learning framework integrating multi-channel features and direction-aware attention. First, based on polarization light field information acquired by a division-of-focal-plane polarization camera, a three-channel composite input feature composed of a polarization intensity map, an adaptive threshold gradient map, and high-frequency residual edge information is constructed. Second, a ResNet backbone network embedded with a squeeze-and-excitation mechanism is adopted, and a direction-aware polarization attention module is introduced to achieve adaptive fusion of multi-scale features through luminance guidance, deep feature enhancement, and a gradient edge branch.
Owner:HANGZHOU CITY UNIV

Power transmission line foreign matter detection method and system based on multi-modal image fusion

The invention discloses a power transmission line foreign matter detection method and system based on multi-modal image fusion, and relates to the technical field of intelligent operation and maintenance and state monitoring of a power system, a lightweight Ev-Mama architecture is introduced into a backbone network part of YOLOv13, the model keeps relatively low calculation complexity, and meanwhile, the power transmission line foreign matter detection efficiency is improved. And the modeling capability of the method on the long-range dependency relationship and the global semantic information is obviously enhanced. Besides, by using the CDIDF module, the EVCS module and the MHSAA module, on the basis of increasing a small amount of calculation, the scale sensing ability, the space structure modeling ability and the context understanding ability of the model are effectively improved, and the performance bottleneck of a traditional YOLO series network in the aspects of processing small targets, shielding targets and cross-scale information fusion is effectively relieved.
Owner:KUNMING UNIVERSITY

Text abstract generation method and system based on sparse attention acceleration

The invention discloses a text abstract generation method and system based on sparse attention acceleration, and the method comprises the steps: reading long text data, carrying out word segmentation and embedded coding processing, extracting a sequence feature vector, and mapping the sequence feature vector into a query matrix Q, a key matrix K and a value matrix V; constructing an abstract generation network, wherein the abstract generation network comprises a sparse attention calculation module, a feedforward calculation module, a prediction head module and a key value cache module; inputting the query matrix Q, the key matrix K and the value matrix V into an abstract generation network, passing through a backbone network formed by stacking a sparse attention calculation module and a feedforward calculation module for multiple times, processing by a prediction header module, and caching a historical decoding state in real time through a key value caching module to obtain an initial text feature vector; and based on the initial text feature vector, executing an autoregressive decoding process through an abstract generation network, and outputting a final abstract result. According to the method, the decoding process can be accelerated, and the semantic integrity and coherence of the generated abstract are ensured.
Owner:ZHEJIANG UNIV

Remote sensing small sample target detection method based on double-attention guided transfer learning

The invention belongs to the technical field of computer vision and image processing, and discloses a remote sensing small sample target detection method based on double-attention guided transfer learning, and the method comprises the steps: obtaining a remote sensing image data set, and carrying out the preprocessing; taking the preprocessed remote sensing image training set as input, constructing a basic detection model by using a ResNet-101 backbone network, a feature pyramid network and a content awareness upsampling and regional proposal network, and obtaining basic model parameters; basic model parameters are used as input, a DA-FSDET network is trained based on a content awareness strip pyramid and a deformable attention area proposal network, and the trained DA-FSDET network is used to acquire a category detection frame containing small sample categories and confidence. Through cascading and cooperative work of the content awareness stripe pyramid and the deformable attention area proposal network, the detection precision and robustness of the multi-scale target in the remote sensing image are effectively improved.
Owner:ZHONGYUAN ENGINEERING COLLEGE

Underwater sound target identification method and system based on autonomous task perception

The invention provides an underwater acoustic target recognition method based on autonomous task perception, and belongs to the field of underwater acoustic signal processing and artificial intelligence. S3, performing feature extraction on the time-frequency spectrogram by the task type extraction network to obtain a task embedding vector, calculating a similarity score, when the maximum similarity score is greater than a set threshold value, outputting a task representation vector, and entering S3, otherwise, inputting features output by the last Transform layer into a classifier; s3, selecting a router according to the task representation vector, selecting a trained expert network by the router, calculating a door control weight, processing universal acoustic features in parallel by the expert network, fusing output of the expert network, fusing deep features and fused features to obtain enhanced features, and enabling the enhanced features output by the last Transform layer to enter a classifier; the invention further provides a system. The problems that the task category autonomous recognition capability is insufficient and the correlation between tasks is ignored are solved.
Owner:NAT UNIV OF DEFENSE TECH

Drainage pipeline defect detection system and method based on multi-scale feature fusion and shielding perception

The invention discloses a drainage pipeline defect detection system and method based on multi-scale feature fusion and occlusion perception, and the system comprises a data set construction module which is used for constructing a drainage pipeline data set covering multi-defect, multi-scale and multi-occlusion scenes; the model training module is based on establishment of a defect collaborative detection model, and a backbone network of the model training module adopts a feature pyramid sharing convolution module to reinforce the multi-scale detail extraction capability; the neck network introduces an advanced screening path aggregation network and a selective feature fusion module to realize dynamic feature screening and fusion; the detection head is integrated with an MCFEM module, and shielding perception and scale adaptability are enhanced. According to the method, the problems of high omission ratio and poor robustness caused by large defect scale change, serious shielding and complex background in drainage pipeline detection are effectively solved, the small target defect identification precision and the shielding scene detection accuracy are remarkably improved, and the method is suitable for high-precision and light-weight detection of multi-scale defects in a complex shielding environment.
Owner:WUHAN INST OF TECH +1

Robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention

The invention relates to a robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention, and belongs to the technical field of image processing. Aiming at the problems of small target feature loss, semantic gap, background noise interference and the like caused by a fixed convolution kernel scale, one-way feature fusion and a static attention mechanism in an existing unmanned aerial vehicle aerial image target detection method, the method comprises the following steps: constructing a detection model comprising a backbone network, a neck network and a detection head network; a feature rearrangement and extraction module is designed in the backbone network to enhance feature learning, an enhanced double-flow feature fusion pyramid is designed in the neck network to optimize multi-scale feature fusion, and a dynamic multi-scale context attention mechanism is designed in the detection head network to suppress irrelevant background noise. The method effectively improves the accuracy and robustness of small target detection, and achieves a clearer and more stable detection effect in a complex environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Insulator defect detection method based on layered multi-scale adaptive fusion network

The invention provides an insulator defect detection method based on a layered multi-scale adaptive fusion network, and the method comprises the following steps: S1, collecting an insulator image data set in a power transmission line through unmanned plane equipment, and marking the type of an insulator; s2, building an HMSAF-YOLO model, and introducing a cross-stage partial multi-scale channel adaptive fusion CSPMSCAF module and a fast spatial pyramid pooling SPPFLSKA module with large kernel separation kernel attention into a backbone network Backbone so as to carry out multi-scale feature extraction and feature enhancement; combining a pyramid region context module PRCM, a gating feature fusion module GFFM and a cross-scale interpolation fusion module CSIF in a neck network to perform feature fusion; s3, performing model training by using the preprocessed insulator image data set; and S4, carrying out insulator defect detection on the to-be-detected image through the final detection model. According to the method, relatively high detection precision can be realized while high calculation efficiency is ensured when the method is deployed on an unmanned aerial vehicle platform with limited calculation power.
Owner:QUANZHOU ELECTRIC POWER TECH INST OF FUJIAN ELECTRIC POWER +2

Underwater target detection method and system, storage medium and equipment

The invention discloses an underwater target detection method and system, a storage medium and equipment, and relates to the technical field of target detection, and the method comprises the steps: obtaining an image data set of an underwater target; constructing an improved YOLOv8 model which comprises a backbone network, a neck network and a head network; a CA attention mechanism module is inserted into the backbone network; replacing a Neck neck network with a BiFPN neck network based on a feature pyramid, and replacing a CBS module with an ADSAMB module; and newly adding a detection head aiming at the tiny target in the head network. Using the preprocessed image data set to train the improved YOLOv8 model; and inputting a to-be-detected underwater image into the trained improved YOLOv8 model, and outputting a position bounding box, a category label and confidence of an underwater target to complete underwater target detection. According to the method, the accuracy, recall, mAP at 0.5 and mAP at 0.5-0.95 of the improved YOLOv8 model in a complex underwater environment are effectively improved, and the method is particularly excellent in the aspect of tiny target detection.
Owner:NANCHANG UNIV

Wafer probe trace accurate detection method based on deep learning

The invention discloses a wafer probe mark accurate detection method based on deep learning, and belongs to the field of wafer probe mark detection, and the method comprises the steps: constructing a pin mark image denoising preprocessing network, employing a small target feature protection and enhancement strategy based on HSV color space and local contrast joint adjustment, and carrying out the recognition of a small target feature; denoising and contrast optimization are carried out on the needle mark image; a multi-scene training sample is generated through mosaic splicing and mix fusion; a dense small target enhancement module is introduced into the backbone network to enhance needle mark feature expression, and a multi-scale feature fusion module is arranged in the neck network to extract full-scale features; and establishing an anchor frame optimization system adaptive to the minimum needle mark target, and adopting an optimizer and learning rate collaborative optimization training strategy to realize model adaptive convergence. According to the method, high-precision detection and robust identification of the wafer probe mark can be realized under a complex background, and the detection accuracy and stability are remarkably improved.
Owner:WUXI UNIV

Unmanned aerial vehicle aerial image small target detection model construction method

The invention relates to the technical field of image target detection, and discloses an unmanned aerial vehicle aerial image small target detection model construction method comprising the following steps: preparing an aerial image data set, preprocessing the aerial image data set, and generating an aerial image sample set; the method comprises the following steps: establishing a basic model on the basis of a YOLOv8 network model, removing a P5 detection layer in a head network Head of the basic model, introducing a P2 detection layer, replacing a specified position of a Conv module in a backbone network Backbone by adopting an ACMConv feature enhancement module, replacing a specified position of a C2f module in a neck network Neck by adopting a C2fMixStructure mixed structure module, and constructing an improved model by adopting a lightweight enhanced detection head structure; and dividing the aerial image sample set into a training set, a verification set and a test set in proportion, training and verifying the improved model, and generating an unmanned aerial vehicle aerial image small target detection model based on AMLP-YOLOv8. According to the invention, the method has higher perception capability when extracting fine target features, and improves the stability and robustness of small target detection of the aerial image of the unmanned aerial vehicle.
Owner:GUIZHOU NORMAL UNIVERSITY

Infrared unmanned aerial vehicle target detection method based on multi-scale self-enhancement cross-layer fusion

The invention discloses an infrared unmanned aerial vehicle target detection method based on multi-scale self-enhancement cross-layer fusion, and aims to solve the problem of difficulty in small target detection of an infrared unmanned aerial vehicle under a complex background. According to the method, a novel target detection network is constructed, and a lossless down-sampling module and a cascade asymmetric convolution module based on spatial reconstruction are designed in a backbone network of the novel target detection network, so that multi-scale context features are extracted while information loss is reduced. The neck network adopts an enhanced pyramid structure, deep semantics and shallow details are integrated through a cross-layer feature fusion module, and a multi-scale adaptive channel attention module is introduced to dynamically re-calibrate fusion features so as to focus key information. The network adopts four detection heads to output multi-scale prediction, and uses an EIoU loss function to optimize small target positioning precision. The method can effectively improve the feature extraction and detection capability of'weak, small and dark 'infrared unmanned aerial vehicle targets, and has higher robustness and accuracy in a complex scene.
Owner:CHANGCHUN UNIV OF SCI & TECH

Bearing ring surface defect detection method based on improved YOLOv11 network

The invention provides a bearing ring surface defect detection method based on an improved YOLOv11 network. The method comprises the following steps: constructing an improved YOLOv11 network model; wherein in the backbone network and the neck network, an original standard convolution module of the YOLOv11 network architecture is replaced by a receptive wild coordinate attention convolution module; a surface detail fusion module is arranged on each of three feature map paths with different scales output from the neck network to the head network; a positioning loss function is configured to be a Focaler-DIOU loss function; training the model; and obtaining a to-be-detected bearing ring surface image, and inputting the to-be-detected bearing ring surface image into the trained model to obtain surface defect information. According to the method, the feature extraction capability of a network model on micro defects is remarkably improved, the detection performance of the network model on multi-scale and polymorphic defects is enhanced, and the positioning precision and convergence efficiency of the network model on irregular defects are improved.
Owner:ZHEJIANG SCI-TECH UNIV

Small target detection method based on Fourier mixed attention mechanism

The invention discloses a small target detection method based on a Fourier mixed attention mechanism. An improved small target image detection network model based on RT-DETR is researched and designed, and a Basic Block module in a backbone network is replaced by a self-developed Fourier mixed attention enhancement module (FTABlock). The module is composed of a Fourier transform attention module (FTAModule) and a hybrid dynamic convolution feedforward network (DKMixFFN). The FFTModule calculates the attention weight in the frequency domain through Fourier transform, and strengthens small target feature expression in combination with position coding and multi-head attention; the DKMixFFN adopts a dynamic convolution kernel and a multi-scale dynamic convolution layer to realize adaptive modeling of multi-scale features. According to the method, under the synergistic effect of frequency domain attention and dynamic convolution, the small target feature extraction and perception capability of the RT-DETR model is effectively enhanced, and the small target detection precision is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

RGBT target tracking network and method fusing multi-interaction feature enhancement mechanism

The invention discloses an RGBT target tracking network and method fused with a multi-interaction feature enhancement mechanism, relates to the technical field of computer vision and image processing, and aims to solve the problems that multi-modal fusion is insufficient and tracking is easy to drift due to fixed or blind updating of a template. The network adopts an end-to-end tracking framework, a backbone network of the network extracts visible light and infrared template images and searches feature tokens of the images through a convolution token embedding module, and performs intra-modal and inter-modal mixed attention interaction by using a multi-interaction Transform module to realize multi-level feature fusion. A target frame is output by adopting an angular point prediction head, a template updating module is introduced, and a template token is dynamically evaluated and updated through two Transform modules, so that the long-term tracking stability is improved. The network effectively deals with complex environment changes by fusing multi-modal information and a dynamic updating mechanism, and the tracking accuracy and robustness are remarkably improved.
Owner:HEFEI NORMAL UNIV

SMPL-X action-to-text generation method based on global and local feature fusion

ActiveCN121502730ASemantic analysisBiological modelsAlgorithmAction semantics
The invention provides a global and local feature fusion SMPL-X action-to-text generation method, and belongs to the field of artificial intelligence. The method comprises the following steps: preprocessing and coding an input SMPL-X action sequence into a double-flow action feature; through a cross-modal mapping module, the double-flow action features are mapped to a pre-trained large language model through independent projection branches, and global conditions and local action prefix embedding are obtained; through a text generation module, a decoder of a pre-trained large language model is used as a trunk network, text cue word embedding is extracted based on a text instruction given by a user, local action prefix embedding and text cue word embedding are spliced and then input into the decoder, and global conditions are injected into each layer of the decoder through a cross attention mechanism. And generating a description text in an autoregression mode. According to the method, the description text which is consistent with action semantics and has sufficient details can be stably and accurately generated, and the generation stability and the cross-scene applicability are improved when disturbance exists in the action sequence.
Owner:ZHEJIANG UNIV

Multi-modal underwater target detection method and system based on layered feature alignment

The invention relates to the technical field of underwater target detection, in particular to a multi-modal underwater target detection method and system based on layered feature alignment. The method comprises the following steps: acquiring an underwater sonar image and an optical image; sonar and optical features are extracted through a double-flow backbone network; background noise is suppressed and key features are enhanced through a local enhancement module; decomposing the features into low-frequency and high-frequency components through a hierarchical alignment module, and establishing a cross-modal feature corresponding relation by using deformable convolution; multi-modal fusion features are generated through a fusion module; and target detection is carried out through the detection head. According to the method, the problem of feature misalignment caused by the difference between the underwater sonar and the optical image due to the imaging principle is solved, adaptive cross-modal feature alignment and efficient fusion are realized, and the precision and robustness of underwater target detection are remarkably improved.
Owner:GUANGDONG OCEAN UNIVERSITY

Target detection method and device in wide-area complex scene, equipment and storage medium

The invention relates to a target detection method and device in a wide-area complex scene, equipment and a storage medium, and the method comprises the steps: carrying out the multi-level feature extraction of a wide-area complex scene image through employing a backbone network, and obtaining a low-level feature, a middle-level feature and a high-level feature; encoding the high-level features by adopting an EDFPT module to obtain encoded features; inputting the low-level features, the middle-level features and the coding features into a feature fusion module to obtain fusion features; screening a fixed number of image features from the fusion features by adopting an IoU-perceived query selection strategy to obtain an initial query vector; and processing the initial query vector by adopting a decoder with an auxiliary prediction head to obtain a wide-area complex scene image target detection result. The method has higher feature sensitivity to dense small targets and special-shaped targets in a wide-area complex scene image, the detection precision is effectively improved while the light weight of the model is ensured, and the omission ratio is reduced compared with RT-DETR.
Owner:NANCHANG UNIV +1

Heterogeneous document structured data extraction system and method based on multi-modal fusion

The invention discloses a heterogeneous document structured data extraction system and method based on multi-modal fusion, and the method comprises the following steps: S1, receiving a heterogeneous document, and carrying out the preprocessing; s2, extracting visual and semantic modal information based on the visual backbone network and an OCR module, and fusing spatial coordinates, text content and layout features; s3, receiving a dynamic target mode Schema, executing field-level semantic matching, and calculating semantic similarity; s4, performing consistency verification, verifying and filling the field values, and correcting the associated fields; s5, converting the format into a data file in a specified format, and reserving fields to be mapped with an original document semantic block; s6, automatically switching a cloud mode and a local mode according to a deployment environment; and S7, outputting the data file of the target field. According to the method, high-precision and low-delay structured data extraction of the heterogeneous format document can be realized, and the automatic processing efficiency and the data credibility are improved.
Owner:ANHUI HANGTIAN INFORMATION CO LTD

Density center guide perception enhancement method for detection of dense small targets in unmanned aerial vehicle image

The invention discloses a density center guided perception enhancement method for detection of dense small targets in an unmanned aerial vehicle image, and the method comprises the steps: obtaining an aerial image of an unmanned aerial vehicle, inputting the aerial image of the unmanned aerial vehicle into a preset backbone network and an encoder, and carrying out the extraction to obtain a multi-scale feature map; selecting a specified layer feature map from the multi-scale feature map, and inputting the specified layer feature map into a density guide target center heat map generator; density features are generated through number estimation branches, and a target center thermodynamic diagram is generated through the density features and used for supervising generation of target center features; inputting the density feature and the target center feature into a density-center feature enhancement module, and processing the density feature and the target center feature by the density-center feature enhancement module to obtain an enhanced feature; fusing the enhanced features with the multi-scale feature map to obtain a fused feature set; and inputting the fusion feature set into a preset decoder and a prediction head, and outputting a detection result of the target in the unmanned aerial vehicle image after the fusion feature set is processed by the decoder and the prediction head.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN) +1

Bearing fault diagnosis method based on physical perception KAM network

The invention relates to the field of rotating machinery fault diagnosis, and discloses a bearing fault diagnosis method based on a physical perception KAM network, and the method comprises the steps: carrying out the discretization of a bearing vibration signal through a Gabor filter group based on physical prior initialization, and generating modal feature lexical elements with physical frequency band meanings; and inputting the lexical elements into a PC-KAM backbone network, calculating a hidden state vector by using a state space model branch, and dynamically adjusting the position of a primary function node of a Kolmogov-Arnod network branch to realize collaborative dynamic feature extraction. In the training stage, an orthogonal subspace constraint and physical perception low-rank adaptation fine tuning mechanism is introduced. And finally, searching a historical fault case, performing multi-modal fusion on the historical fault case and the deep feature sequence, mapping a fusion representation into a soft prompt, and inputting the soft prompt into a large language model to generate a diagnosis report. According to the method, the problems of poor physical interpretability of characteristics and few-sample diagnosis under variable working conditions are effectively solved, and the generalization and decision-making ability of a diagnosis system are improved.
Owner:DONGGUAN UNIV OF TECH

Defect detection system and method for glass bottle

The invention discloses a defect detection system and method for a glass bottle, and belongs to the field of industrial visual defect detection and analysis. The system comprises a GBDM module, the GBDM module is a defect detection module, and the GBDM module comprises a front attention enhancement structure, a rear attention enhancement structure and a defect detection structure, a multi-branch feature extraction structure is used for generating four feature branch diagrams; and the feature fusion structure is used for fusing the four feature branch diagrams. The GBDM module is embedded in the backbone network of the YOLOv8 network model, and multi-branch structures such as wavelet transform, Gabor direction filtering and specular reflection suppression are utilized, so that the feature extraction capability of multiple types of complex defects such as cracks, smudginess and damage on the surface of the glass bottle is enhanced, the generalization capability of the model in a small sample scene is remarkably improved, and the accuracy of the model is improved. The real-time requirement of an industrial scene is met, continuous self-optimization is achieved, and the method has high robustness, high accuracy and high application value.
Owner:LUZHOU LAOJIAO CO LTD

Prompt guidance and multi-modal fusion-based class incremental learning method

The invention provides a class incremental learning method based on prompt guidance and multi-modal fusion, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: firstly, performing semantic extension on a category label, and constructing semantic enhanced text representation through a text encoder; then block embedding and hierarchical feature extraction are carried out on the input image by using a pre-trained visual encoder, a cross-modal unified embedding space is constructed, a bimodal prompt gating fusion module is introduced into the unified embedding space, and adaptive weighting is carried out on text prompt and image prompt according to gating weight to generate fusion prompt; through a bimodal prompt collaborative filtering module, screening out a prompt set most relevant to the current task according to the similarity of the semantic features of the image and the text; the pre-training backbone network is frozen in the increment stage, only prompt parameters and fusion layer weights are optimized, a joint loss function is used for parameter updating, finally, image and text data are input in the reasoning stage, cross-modal similarity is calculated, and a classification prediction result is output.
Owner:NORTHEASTERN UNIV CHINA

Metal tube surface defect detection method based on improved RT-DETR

The invention discloses a metal tube surface defect detection method based on an improved RT-DETR (Reverse Transcription-DETR), which is used for carrying out multi-dimensional improvement on an RT-DETR model aiming at the problems that the surface defects of the metal tube are diversified in form, different in scale and complex in background. A backbone network of RT-DETR is modified by using a C3k2 module in a backbone network of yov11, a cylinder position coding module CPE is introduced to model a spatial structure of a cylindrical surface of a metal pipe, and a GLAF module integrated with local, medium and global attention mechanisms is adopted to perform multi-stage feature extraction. Secondly, in a feature fusion layer, a cylinder damage perception attention interaction module AIFIDML is used for replacing a standard Transform encoder, and the context perception ability of the model to defects is enhanced; and finally, in the neck network Neck, multi-level feature integration is optimized by adopting a multi-scale adaptive double-path feature fusion module MSADF, and the detection capability of tiny and large-scale defects is improved. According to the method, high-precision and high-efficiency end-to-end detection on the surface defects of the metal tube is realized.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Super-resolution remote sensing image target detection method based on multi-modal fusion

The invention relates to the technical field of computer vision and remote sensing image processing, in particular to a super-resolution remote sensing image target detection method based on multi-modal fusion, which is realized on the basis of a target detection network obtained by improving a YOLOv5 network. The improvement comprises the following steps: introducing a dynamic cross calibration module and a lightweight super-resolution auxiliary branch behind an input layer; replacing part of common convolution in the backbone network with space-to-depth convolution; the method comprises the following steps: inputting a visible light image and an infrared image into a dynamic cross calibration module, carrying out feature extraction and fusion on the visible light image and the infrared image through the dynamic cross calibration module, and outputting fusion features; inputting the visible light image and the infrared image into a lightweight super-resolution auxiliary branch, and performing super-resolution reconstruction through the branch to generate a high-resolution feature; the problems that an existing remote sensing image target detection method is low in accuracy and prone to loss of detail features when facing small targets, low resolution and complex backgrounds are solved.
Owner:TAIYUAN NORMAL UNIV

Target identification method and system based on sonar image assisted optical image

The invention relates to the technical field of underwater target detection, in particular to a target recognition method and system based on a sonar image assisting an optical image. The method comprises the following steps: acquiring an underwater optical image and a sonar image, inputting the underwater optical image and the sonar image into a double-flow multi-scale feature extraction backbone network, and generating an optical multi-scale feature map and a sonar multi-scale feature map; processing the multi-scale feature map through an optical-sonar adaptive feature fusion module to generate enhanced multi-modal fusion features; inputting the multi-modal fusion features into a detection head, and outputting an underwater target detection result; according to the method, the problems of optical-sonar image modal isomerism and space mismatch are effectively solved, and the precision and robustness of underwater target detection are remarkably improved.
Owner:GUANGDONG OCEAN UNIVERSITY

Dynamic self-adaptive NAND Flash reading method and system based on artificial intelligence, medium and equipment

The invention discloses a dynamic self-adaptive NAND Flash reading method and system based on artificial intelligence, a medium and equipment, and relates to the technical field of solid-state storage, the method comprises the following steps: S1, initiating a reading request, reading page data by using default voltage, decoding the page data through an LDPC algorithm, constructing a feature vector and performing preprocessing, inputting the voltage into a pre-trained neural network model, and carrying out voltage prediction; s2, carrying out fine sampling on the predicted voltage, reconstructing accurate threshold voltage distribution, and calculating a log-likelihood ratio; s3, designing a shared bottom layer backbone network, enabling a neural network model to have independent output branches, predicting the voltage, and evaluating the long-term health state of the block; and S4, designing an online transfer learning mechanism, and enabling the neural network model to quickly adapt to a new environment by using a small amount of new data. According to the method, the delay is greatly reduced, the power consumption is remarkably saved, data loss is prevented through early warning of bad blocks, and the write amplification factor is reduced to 1.26 through dynamic read interference suppression.
Owner:JIANGSU XINSHENG INTELLIGENT TECH CO LTD

Industrial park abnormal behavior detection method based on improved YOLOv11

The invention relates to an industrial park abnormal behavior safety detection method based on improved YOLOv11, and the method comprises the steps: obtaining an industrial park abnormal behavior sample data set, and dividing the sample data set into a training set, a test set and a verification set; modifying a network structure of the YOLOv11, and replacing common convolution of a trunk feature extraction network of an original model with CG Block convolution of a context guided network; replacing a C3K2 module of the backbone network with a multi-scale deep convolution module; and the NECK part is completely replaced by a cross-scale feature fusion module CCFM, so that the adaptability of the model to scale change and the detection capability of the model to small-scale objects are enhanced. According to the method, the YOLOv11 algorithm is improved through methods of improving a network structure, optimizing a model and the like, the calculation complexity is remarkably reduced while the detection precision is ensured, and the method is more suitable for deployment of embedded and mobile devices and has good practicability.
Owner:SHENYANG UNIV

Elevator passenger abnormal behavior recognition method based on time sequence displacement

The invention discloses an elevator passenger abnormal behavior recognition method based on time sequence displacement, which comprises the following steps: acquiring an elevator monitoring video stream and carrying out sparse segmented sampling on the elevator monitoring video stream to obtain a network input tensor; performing feature extraction on the input tensor through a lightweight backbone network to obtain a preliminary feature map; performing spatial attention weighting operation on the preliminary feature map to obtain weighted features; performing gating adaptive time sequence shift fusion operation to obtain a fused time sequence feature map; and obtaining a final feature map from the remaining Block modules of the lightweight convolutional neural network, performing global average pooling on the time dimension and the space dimension to obtain a one-dimensional feature vector, mapping the one-dimensional feature vector to different behaviors through a full connection layer, and obtaining the probability of each type of behavior through a classifier. The space attention module and the gating variable time sequence shifting module are embedded into the lightweight backbone network, and high-efficiency and high-precision recognition of elevator passenger abnormal behaviors is achieved.
Owner:JIANGNAN UNIV

Long-tail data generation method based on feed-forward stylized reconstruction

The invention discloses a long-tail data generation method based on feed-forward stylized reconstruction. The method comprises the following steps: constructing multi-modal geometric perception input and feature codes; reconstructing a semantic guidance scene based on a feature-level linear modulation mechanism; the invention relates to 3D consistency controllable style migration based on a feedforward reconstruction model. According to the invention, a feedforward reconstruction network based on DINOv2 improvement is constructed. Multi-modal adaptation is performed on a backbone network input layer, and decoupled high-level semantic features are explicitly injected by using a feature-level linear modulation mechanism, so that the geometric accuracy of automatic driving scene reconstruction is improved. On the basis of a three-dimensional scene of feedforward reconstruction, a distillation strategy is adopted, and the generation capability of a 2D diffusion model is migrated to a 3D Gaussian field. According to the method, the phenomena of object deformation, disappearance and the like in the generated data are effectively avoided, and long-tail weather data with geometric consistency of rain, snow and the like can be generated through the text instruction.
Owner:DALIAN UNIV OF TECH