Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1621 results about "Encoder decoder" patented technology

Accurate micro-crack segmentation method integrating feature fusion and convolution attention

The invention provides a microcrack precise segmentation method integrating feature fusion and convolution attention, and belongs to the field of image processing. According to the method, a crack segmentation network based on an encoder-decoder architecture is constructed, a convolution block attention module is introduced at an encoder end, background noise is adaptively suppressed and obvious characteristics of cracks are enhanced through a channel and space dual attention mechanism, and the method is suitable for the adaptive segmentation of the cracks on the premise of almost not increasing the calculation overhead. The sensitivity of the model to microcracks is improved; a feature fusion module is introduced at a decoder end, and cooperation of low-layer details and high-layer semantics is realized through cross-layer fusion, so that a semantic gap is effectively bridged, detail loss caused by traditional convolution stacking is avoided, and continuity and a complete topological structure of a long and narrow crack are ensured. According to the method, through collaborative optimization of multi-scale feature extraction and an attention mechanism, accurate capture of the saliency features of the crack and effective suppression of complex background interference are realized, and the detection sensitivity and overall segmentation consistency of the micro-crack are remarkably improved.
Owner:DALIAN UNIV OF TECH

System and method for reconstructing 3D scene data from 2D image data

A method and apparatus for reconstructing a three-dimensional (3D) scene from a two-dimensional (2D) input image of the scene using a fully-differentiable transformer-based encoder-decode. A 2D input image encoded into a set of image features using a pre-trained vision transformer model, wherein the vision transformer model is pre-trained with multi-view RGB image supervision and point cloud supervision. The set of image features is projected onto a 3D triplane representation using a transformer decoder to obtain output triplane tokens. A triplane representation is created from the tokens and queried. 3D point features of color and density for volumetric rendering re predicted using a multi-layer perceptron. The geometry of the generated 3D asset is represented with a surface mesh including vertices and triangular faces. A texture map by is created with a multichannel image in UV space. Multiple views of the 3D scene are simultaneously generated based on the surface mesh.
Owner:FUTUREVERSE IP LTD

Three-dimensional seismic fault identification method based on double-attention multi-scale fusion U-Net

The invention provides a three-dimensional seismic fault identification method based on double-attention multi-scale fusion U-Net. The method comprises the following specific steps: constructing a U-shaped network architecture comprising an encoder, a decoder and jump connection; a constructed multi-scale feature fusion module is embedded in the first level of the encoder, and the extraction capability of fault features of different scales is enhanced through a multi-branch structure; introducing the constructed hole fusion modules into the second and third levels of the encoder, designing and expanding a receptive field by using multiple expansion rates, and capturing fault structures of different scales; a double-attention parallel mechanism is integrated in jump connection, and the sensitivity of channel attention and space attention to fault features is improved; constructing a combined loss function; and finally, performing three-dimensional seismic data training and reasoning based on the optimized model to realize high-precision fault identification. The method has high generalization and accuracy, and especially has good performance in the aspect of seismic image fault identification containing a large fault scale span.
Owner:SOUTHWEST PETROLEUM UNIV

Non-autoregressive transformer-based modeling method for 4-level pulse amplitude modulation high-speed transmitter

Disclosed in the present invention is a non-autoregressive Transformer-based modeling method for a 4-level pulse amplitude modulation high-speed transmitter. The method involves establishing a deep learning model having an encoder-decoder architecture to predict the behavior of a 4-level pulse amplitude modulation transmitter. An encoder processes unordered non-sequential inputs, including an input signal parameter and link parameters, to generate a context vector and then transmit same to a decoder. The decoder uses both the context vector generated by the encoder and a transmitter output signal sequence to generate a categorical probability distribution for each point in the sequence one by one. The model is trained using a random masking strategy, and inference is performed by means of non-autoregressive decoding and filtering, so that the model can perform parallel prediction on an output sequence, and perform a filtering process to predict an output signal. Compared to traditional simulation methods, the present invention achieves a significant acceleration effect, particularly when processing multi-link systems.
Owner:ZHEJIANG UNIV

Medical image segmentation method and system based on residual Mama and multi-scale boundary enhancement

The invention relates to a medical image segmentation method and system based on residual Mama and multi-scale boundary enhancement. The method comprises the following steps: acquiring and preprocessing a medical image; inputting the image into a segmentation model based on an encoder-decoder architecture; the encoder synchronously extracts local texture features and models long-range spatial dependence through residual error convolution blocks and residual error Mama blocks which are alternately connected; fusing and enhancing the jump connection features between the encoder and the decoder through a boundary enhancement module to optimize boundary characterization; integrating a multi-scale gating attention module in a decoding path, and adaptively selecting and fusing multi-scale context features; and finally outputting the high-precision segmentation mask. The method effectively solves the problems that in the prior art, long-range dependence and local details are difficult to consider, the multi-scale feature fusion capability is insufficient, boundary segmentation is fuzzy and the like, and the segmentation accuracy, the boundary continuity and the clinical practicability are remarkably improved.
Owner:NINGBO MEDICAL CENT LIHUILI HOSPITACL

Adaptive mask medical image segmentation method based on self-supervised mask and deep reinforcement learning

The invention discloses an adaptive mask medical image segmentation method based on a self-supervised mask and deep reinforcement learning, and the method comprises the steps: employing a classic encoder-decoder architecture for a self-supervised mask reconstruction network, fusing a Swin Transform encoder, and carrying out the feature fusion of local image blocks through a self-attention mechanism; according to the self-adaptive mask model, a PPO deep reinforcement learning algorithm is adopted, a strategy network and a value network are constructed, mask actions are dynamically regulated and controlled, reconstruction errors are gradually reduced, a mask strategy is continuously optimized in multiple times of strategy updating for self-adaptive optimization, and high-quality reconstruction of a medical image influenced by missing information is achieved; according to the method, high-quality feature representation can be obtained in an unlabeled data environment, and relatively high precision and accuracy are presented on a public data set.
Owner:YUNNAN UNIV

Radar echo extrapolation method and system based on frequency domain enhancement

The invention discloses a radar echo extrapolation method and system based on frequency domain enhancement, and the method mainly comprises the following steps: obtaining and preprocessing a historical radar echo grayscale image sequence, generating a sequence sample through a sliding window, and dividing the sequence sample into a training set, a verification set and a test set; the method comprises the following steps: constructing a frequency domain enhanced U-Net network comprising an encoder-decoder structure, introducing a multi-scale deep convolution structure into an encoder and a decoder, and enhancing frequency domain features by using a frequency domain dynamic attention mechanism in jump connection; inputting the training set into the model for training by adopting a composite loss function comprising intensity weighted loss, frequency domain consistency loss and structural similarity loss; and inputting the test set into the trained model, and outputting a radar echo prediction result at a future moment. The method can be effectively applied to the fields of short-term and temporary weather forecast, severe convection monitoring and the like, and provides more accurate and reliable radar echo prediction support for meteorological disaster early warning.
Owner:HANGZHOU DIANZI UNIV

Unsupervised wind power equipment blade fault detection method based on phase perception parallel attention mechanism

The invention relates to a wind power equipment blade fault detection technology, discloses an unsupervised wind power equipment blade fault detection method based on a phase perception parallel attention mechanism, and solves the problems that an existing wind power equipment blade fault detection method is high in dependence on labeled data, insufficient in generalization ability under strong noise and variable working conditions and high in fault detection efficiency. And a weak transient fault signal and a dynamic change characteristic are difficult to capture robustly. According to the scheme of the invention, the method comprises the steps: collecting a blade operation audio signal, and extracting a dual-channel time-frequency feature containing an amplitude spectrum and a phase spectrum through improved short-time Fourier transform; a deep adversarial auto-encoder is constructed by using an encoder containing a phase perception parallel attention module, a decoder and an auxiliary encoder, and normal working condition feature distribution is learned by reconstructing an error loss, potential representation consistency loss, adversarial loss and phase consistency loss optimization model during off-line training; in the reasoning stage, the fault is judged based on the feature distance score and the reconstruction error score.
Owner:CHINA HYDROELECTRIC ENGINEERING CONSULTING GROUP CHENGDU RESEARCH HYDROELECTRIC INVESTIGATION DESIGN AND INSTITUTE

Evidence obtaining method and system based on image processing

The invention provides an evidence obtaining method and system based on image processing, and the method comprises the following steps: S1, generating a pixel-level depth-of-field distribution diagram of an input image through a multi-scale encoder-decoder network, employing an edge perception optimization layer in a decoding stage, and improving the depth-of-field boundary precision through minimizing a local gradient consistency loss function; s2, performing depth-of-field rationality verification based on an optical imaging physical model, and triggering a first-level tampering alarm by calculating a defocusing fuzzy radius and a gradient direction of a selected region when a difference between the defocusing gradient directions of a target region and a background region exceeds a preset threshold value; and S3, dynamically positioning a key pixel region, identifying a depth-of-field mutation boundary by using an edge detector, calculating by combining local texture complexity, screening a pixel set of which the entropy value is higher than a threshold value and which is located at the mutation boundary, correlating metadata to verify the rationality of the physical size and the spatial position of an object, and eliminating false detection caused by perspective transformation.
Owner:XIAMEN MEIYA ZHONGMIN TECH CO LTD

Human body posture estimation method based on millimeter wave radar point cloud

The invention discloses a human body posture estimation method based on millimeter wave radar point clouds, which comprises the following steps of: processing each frame of input millimeter wave radar sparse point clouds by utilizing a built posture estimation model, and finally predicting and outputting a three-dimensional coordinate sequence of corresponding human body key joint points; wherein the attitude estimation model is composed of a point cloud completion network based on an encoder-decoder architecture, a global-local double-branch network and a feature fusion and prediction mechanism, and the method comprises the following steps: firstly, enhancing the density and integrity of an original sparse point cloud by using the point cloud completion network; the method comprises the following steps: complementing point clouds and original point clouds, respectively processing the complemented point clouds and original point clouds through a global-local double-branch network, extracting global structure features and local detail features, finally fusing the two features through a feature fusion and prediction mechanism, predicting three-dimensional coordinates of human body joint points through a regression layer, and realizing accurate attitude estimation. The method improves the estimation precision and robustness, maintains the privacy protection advantage, and is suitable for complex application scenes.
Owner:SOUTH CHINA UNIV OF TECH

Extreme sea condition parameter identification system based on deep learning

The invention discloses an extreme sea condition parameter identification system based on deep learning, and relates to the technical field of ship navigation auxiliary equipment, in particular to a self-adaptive sea condition identification device which is used for acquiring image data and inertial measurement data of a current sea condition; the wave field visual depth estimation module is used for extracting visible light image features and infrared image features of a wave area from image data of the current sea condition, fusing the extracted visible light image features and infrared image features, using an encoder-decoder architecture and fusing an energy function to obtain a pixel-level wave height field, and outputting the pixel-level wave height field. The three-dimensional reconstruction of the wave surface is realized; the multi-modal data fusion module uses a filter dynamic model and a cost function to eliminate space-time asynchronous errors between inertial measurement data and visual perception data, performs multi-modal data fusion, and outputs wave field real-time parameterization information. According to the invention, the sea condition parameter real-time high-precision identification capability of the autonomous unmanned ship or the offshore carrying platform can be improved.
Owner:WUHAN UNIV OF TECH

2D medical image segmentation method and system based on Mama and UNet

The invention discloses a 2D medical image segmentation method and system based on Mama and UNet, and the method comprises the steps: collecting and preprocessing a medical image segmentation data set, and obtaining a training set; constructing a 2D medical image segmentation model based on Mama and UNet, wherein the 2D medical image segmentation model comprises a block embedding layer, an encoder, a decoder and a prediction generation layer; designing an adaptive hierarchical loss function based on gradient statistics, and training the 2D medical image segmentation model on the training set; and inputting the medical image with segmentation into the trained model to complete image segmentation. According to the invention, the method can achieve the automatic and intelligent segmentation of the medical image through the innovative construction of the 2D medical image segmentation model based on Mamba and UNet, and is higher in segmentation accuracy and efficiency.
Owner:ZHEJIANG UNIV

Medical image segmentation method with adaptive receptive field and feature correction

The invention relates to the technical field of medical image processing, and particularly discloses a medical image segmentation method with adaptive receptive field and feature correction, which comprises the following steps: (1) acquiring an original medical image and a segmentation label thereof, and constructing a training and testing data set; (2) carrying out size normalization and enhancement processing on the image; (3) establishing an improved U-shaped encoder-decoder segmentation network, introducing an adaptive branch mixed shape convolution module in a shallow layer, and improving edge and texture feature modeling capability by adopting a multi-branch banded convolution and channel attention mechanism; (4) a residual directional feature interaction module is introduced into a deep layer, a spatial dependency relationship is modeled through an information interaction structure in the horizontal and vertical directions, and the direction sensing ability of the heterostructure is enhanced; and (5) completing network training and reasoning, and outputting a segmentation result. The method gives consideration to the calculation efficiency and the segmentation precision, and is suitable for the automatic segmentation task of various types of medical images with complex structures.
Owner:SOUTHWEST PETROLEUM UNIV

Deformation online measurement and control method for multi-robot collaborative assembly

The invention relates to a deformation online measurement and control method for multi-robot collaborative assembly. The method comprises the steps that multi-view surface images of workpieces in the assembly process are collected in real time through distributed robots and cameras arranged at the tail ends of the distributed robots; each view angle surface image is input into a deep learning model trained based on a digital image related technology, the optimal pixel displacement corresponding to each view angle surface image is output, and the optimal pixel displacement corresponding to each view angle surface image is converted into a spatial displacement label; the deep learning model takes an encoder-decoder as a trunk network; fusing the spatial displacement labels corresponding to the view angle surface images to obtain a fused displacement field; based on the fusion displacement field and the nominal path planning point, the tail end pose of the distributed robot is determined; the assembly error is calculated based on the reference target positioning and the tail end pose, and the PID controller adjusts the joint space of the distributed robot based on the assembly error. According to the method, the calculation overhead is remarkably reduced, and the measurement precision and robustness are improved.
Owner:HUNAN UNIV

Systems and methods for underwater imagery enhancement

A computer-implemented method for training a generative adversarial network (GAN) for enhancing underwater images. An adversarial loss is computed for updating a discriminator model and a combined loss is calculated for updating a generator model. The combined loss is calculated based on loss components including the adversarial loss and at least one further loss component. Additionally disclosed herein is a generator network for processing underwater images that includes a novel encoder-decoder model architecture. Unlocking insights from Geo-Data, the present invention further relates to improvements in sustainability and environmental developments: together we create a safe and liveable world.
Owner:FNV IP BV

Training method and device for three-dimensional open vocabulary semantic segmentation model

The invention belongs to the technical field of three-dimensional scene understanding, and particularly relates to a training method and device for a three-dimensional open vocabulary semantic segmentation model. The training method comprises the steps of obtaining multi-view RGB-D images of a target area, performing multi-stage reasoning on each image through a visual language model, generating a target vocabulary list, prompting a two-dimensional segmentation model to establish a pixel-level text label, performing depth mapping on the images to generate a first point cloud, and generating a second point cloud; mapping the text tag to the first point cloud to generate a point-by-point text tag; pre-training a neural network model with a sparse encoder-decoder structure by taking the point-by-point text label as a supervision signal, and generating a three-dimensional segmentation model on the first point cloud; and for the second point cloud of the complete scene of the target area, matching point feature embedding and text embedding with the highest similarity in the shared vision-language feature space, generating a credible point-text tag pair, and finely adjusting the three-dimensional segmentation model based on the credible point-text tag pair.
Owner:UNIV OF SCI & TECH OF CHINA

Medical image segmentation method and system based on deep learning

The invention relates to the technical field of medical image processing and computer vision, in particular to a medical image segmentation method and system based on deep learning, the method is based on a U-shaped encoder-decoder architecture, a DSAB module is introduced into an encoder, and context perception of a directional anatomical structure is enhanced through complementary directional space shift and CSA mechanism weighting; an MGCF module is designed in a decoder, and a parallel multi-scale convolution path and an AGCA mechanism are combined, so that multi-level features are efficiently fused to recover boundary details. Meanwhile, links of data preprocessing, Transform structure details, segmentation result post-processing and the like are supplemented, the model performance is improved through a mixed loss function and an optimization training strategy, and the method has remarkable advantages in segmentation precision and boundary definition and provides powerful support for clinical auxiliary diagnosis.
Owner:ANHUI POLYTECHNIC UNIV

Flange sealing element surface defect detection method and system

The invention provides a method and a system for detecting surface defects of a flange sealing element. The method comprises the following steps: expanding an annular surface image into a rectangular image; constructing a defect segmentation network to process the rectangular image to obtain a preliminary defect mask; the defect segmentation network adopts an encoder-decoder structure, a parallel multi-scale feature extraction module is arranged at the tail end of an encoder, and the parallel multi-scale feature extraction module comprises a plurality of parallel convolution branches with different receptive fields; training the defect segmentation network by using a composite loss function, and performing post-processing on the initial defect mask by using a full-connection conditional random field to obtain an optimized defect mask; and inversely transforming the optimized defect mask to a Cartesian coordinate system, and identifying and marking the position and contour of the defect on the original annular surface image.
Owner:山西宝航重工有限公司

Anti-compression coding robust video watermark generation method based on adversarial neural network

The invention is suitable for the field of digital watermarking, and provides an anti-compression coding robust video watermark generation method based on an adversarial neural network, and the method comprises the steps: constructing an MSCA-GAN model; the model comprises an encoder, a decoder, a distortion layer, a discriminator network and an opponent network, and the specific steps are as follows: step S1, the encoder extracts different scale features of a video frame by using a multi-scale convolution attention mechanism, calculates attention weights, efficiently hides and embeds binary watermark information into the video, generates a watermark-containing video, and transmits the watermark-containing video to the decoder; performing confrontation optimization on an embedding strategy with a discriminator in training; according to the method, for H.264 compression layer special training, the multi-scale convolution attention mechanism and the depth separable convolution are combined, and the anti-compression robustness of the watermark under the H.264 standard is effectively improved; through common attack training such as noise layer simulation cutting and zooming, the watermark can still keep high extraction accuracy and robustness in a complex environment.
Owner:ENG UNIV OF THE CHINESE PEOPLES ARMED POLICE FORCE

Medical image segmentation method based on multi-scale convolution bidirectional Mama

The invention relates to the technical field of image processing, in particular to a medical image segmentation method based on multi-scale convolution bidirectional Mama. The method comprises the following steps: constructing a medical image segmentation model comprising an encoder, a decoder and a jump connection module; a medical image segmentation model is trained by using the medical image, local features are extracted through a CNN branch, and a long-range context dependency relationship is captured through a multi-scale bidirectional Mama branch; and updating the parameters of the medical image segmentation model according to the target loss, and performing medical image segmentation through the trained medical image segmentation model, thereby improving the medical image segmentation effect.
Owner:JIANGXI NORMAL UNIV

Distributed photovoltaic power prediction method and system based on high-dimensional gridding numerical weather forecast

The invention relates to the technical field of photovoltaic prediction, in particular to a distributed photovoltaic power prediction method and system based on high-dimensional gridding numerical weather forecast, and the method comprises the steps: carrying out the standardization of the numerical weather forecast data and photovoltaic power historical data of a target region, and achieving the time-space alignment based on a preset grid, generating a gridding data set; utilizing convolution processing to extract local space features, and converting and fusing the local space features into a feature sequence containing space and historical time sequence information at the same time; modeling is carried out through an encoder-decoder architecture, an encoder excavates historical power dependence, and a decoder dynamically couples future meteorological characteristics with historical power through an attention mechanism and outputs a grid-level predicted value; aggregating to obtain a system total power prediction result; by establishing a unified space-time grid, refined alignment of data is realized, cross-space-time dynamic fusion is performed in combination with convolution and an attention mechanism, and prediction precision and stability can be kept in complex weather.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +1

Face image restoration method based on inverse mapping of generative adversarial network

The invention provides a face image restoration method based on generative adversarial network inverse mapping, and belongs to the technical field of image restoration and computer vision. According to the method, an encoder-decoder architecture is adopted, an encoder is a ResNet combined with a channel attention mechanism and can encode a to-be-restored face image (containing block-shaped shielding) to a potential space of a StyleGAN, and a W + vector of 18 * 512 dimensions is obtained; and the decoder is a pre-trained StyleGAN generator, and can convert the potential vector into a complete face image to realize restoration. In the training process, through a weighted combination constraint model of L2 loss, perception loss and face identity loss, it is ensured that a restoration result is consistent with an original image in pixel, feature and identity levels. Meanwhile, by means of the decoupling characteristic of the StyleGAN potential space, the repaired face features (such as smile and age) can be edited. According to the method, complex preprocessing is not needed, large-area missing face images can be efficiently repaired, identity consistency is kept, and the method is suitable for monitoring image enhancement, old photo repair and other scenes.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

RGB-T saliency target detection method based on query guide specific learning network

The invention relates to an RGB-T saliency target detection method based on a query guide specific learning network, and the method comprises the steps: constructing a network model of an encoder-decoder architecture, and setting a visible light image saliency detection module, an infrared image saliency detection module and a bimodal saliency detection module according to the network model, visual features of the visible light RGB image are obtained through an encoder of the visible light image saliency detection module, and infrared features of the infrared thermal imaging image are obtained through the infrared image saliency detection module; in a bimodal information fusion module in the bimodal saliency detection module, generating an RGB modal feature, an infrared modal feature and a cross-modal fusion feature; and the three-feature cross block outputs the enhanced multi-modal fusion features. Compared with the prior art, the method has the advantages of high accuracy, high modal adaptability, high efficiency and the like.
Owner:TONGJI UNIV

Real-time domain adaptive defect detection method based on double alignment and uncertainty filtering

The invention discloses a real-time domain adaptive defect detection method based on double alignment and uncertainty filtering. The method comprises the following steps: firstly, constructing a defect detection data set which comprises source domain data and target domain data; then, constructing a defect detection model which comprises a backbone network, an encoder, a decoder and a detection head; and finally, pre-training the defect detection model by using the source domain data to obtain a teacher model, and storing the multi-scale source domain features extracted by the teacher model backbone network into a feature database in groups according to scales. Initializing the defect detection model by using the teacher model parameters to obtain a student model; and carrying out collaborative optimization on the teacher model and the student model by using a pseudo-bounding box and feature distribution double-alignment strategy to realize real-time domain adaptive detection. According to the method, performance degradation caused by domain offset is effectively relieved through a double-alignment strategy, meanwhile, error tag accumulation is avoided through uncertainty perception and filtering of a pseudo-bounding box, and the stability during domain adaptation is improved.
Owner:HEBEI UNIV OF TECH

Remote sensing image target detection method based on RT-DETR

The invention discloses a remote sensing image target detection method based on RT-DETR, and the method comprises the following steps: 1, obtaining a remote sensing image data set, and completing the preprocessing; 2, inputting a backbone network, and extracting a multi-scale feature map; step 3, inputting a DSF module to realize multi-scale feature adaptive fusion; step 4, inputting an MSFE module, enhancing features and modeling long-range dependence; 5, embedding a G-HCO module in a corresponding stage of the backbone network, and optimizing feature expression; step 6, inputting an RT-DETR encoder-decoder, and outputting a target category and a target position; and 7, filtering the low-confidence prediction frame to obtain a final detection result. According to the remote sensing image target detection method based on the RT-DETR, through multi-scale feature adaptive fusion of the DSF module, frequency domain-spatial domain dual-domain feature enhancement of the MSFE module and global feature optimization of the G-HCO module, the method can also be expanded to multi-task scenes such as video target detection and segmentation-detection combination, and the generalization ability and the application range are better.
Owner:HEFEI UNIV

Infrared small target detection algorithm based on multi-mode guidance

The invention discloses an infrared small target detection algorithm based on multi-modal guidance, and the algorithm comprises the steps: generating three complementary modal features from an input infrared image through a multi-modal extraction module, and enabling the features to reflect the local gray fluctuation, region comparison relation and original brightness information respectively; and then inputting the multi-modal features into a neural network adopting an encoder-decoder structure, introducing a cross-modal guide residual module and explicitly integrating channel and space transconductance information between modals into each layer of an encoder and a decoder, enhancing weak small target features, and after the encoder completes multi-level feature extraction, extracting the multi-modal features by using the neural network. Fine-grained interaction is carried out among different modals through a cross-modal cross fusion module by utilizing multiple groups of multi-head cross attention mechanisms, efficient token mixing is realized by combining depth separable convolution and point-by-point convolution, and dynamic fusion is carried out through adaptive modal weight, so that the detection accuracy and robustness of a weak small target are remarkably improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Remote sensing image reasoning segmentation method based on scene perception guide network

The invention discloses a remote sensing image reasoning segmentation method based on a scene perception guide network, and the method is characterized in that the method comprises the following steps: 1, constructing a multi-scene reasoning segmentation remote sensing data set; step 2, constructing a scene perception guiding network based on an encoder-decoder structure, and processing the data set constructed in the step 1 to obtain a segmentation mask; 3, performing network training by adopting a self-adaptive scene perception loss function; according to the cross-scene generalization performance evaluation method, an efficient remote sensing image reasoning segmentation model is constructed through a scene perception guide mechanism and context-rich adaptive feature transformation, and the adaptation of semantic expressions in different scenes is guided by introducing scene cognition, so that the robustness of the remote sensing image reasoning segmentation model is improved. The problem that feature representation is inconsistent when a traditional method faces a complex geographical environment is effectively avoided.
Owner:安徽省第二测绘院 +1

Medical image segmentation method based on contextual information and multi-scale feature fusion

The invention belongs to the technical field of medical image segmentation, and particularly discloses a medical image segmentation method based on contextual information and multi-scale feature fusion, and the method comprises the steps: obtaining medical image data, and constructing a medical image segmentation model of a double-U-shaped encoder-decoder architecture; inputting the medical image data into a first U-shaped network for feature extraction to obtain multi-scale context features; multiplying the medical image data and the multi-scale context features element by element; inputting an element-by-element multiplication result into a second U-shaped network for global information modeling and local boundary feature enhancement, and outputting to obtain a final fusion feature; and splicing the multi-scale context features and the final fusion features to obtain a medical image segmentation result. According to the method, the problems that the segmentation precision is low, the calculation complexity is high, the small polyp segmentation effect is poor and efficient segmentation cannot be realized when an existing medical image segmentation method is used for carrying out polyp segmentation are solved.
Owner:XIAN UNIV OF POSTS & TELECOMM +1

Complex network disintegration method based on evolution deep reinforcement learning

The invention discloses a complex network disintegration method based on evolution deep reinforcement learning. According to the method, an encoder-decoder model fusing a graph convolutional neural network and a deep Q network is constructed, and is used for efficiently extracting importance features of nodes in a complex network and realizing dynamic decision-making of a node disassembling sequence according to the importance features. In order to optimize model parameters and improve search capability, an evolutionary algorithm is introduced to perform global exploration on the model parameters, and the problem that a directional optimization strategy is easy to fall into local optimum is avoided. Meanwhile, deep mining is carried out on an evolution result in combination with a reinforcement learning strategy, the overall optimization process is accelerated, and advantage complementation of parameter evolution and strategy learning is achieved. Experimental results show that the method significantly improves the efficiency and precision of network disassembly while maintaining the robustness of the model, and has good practical value and wide application prospects.
Owner:NANJING UNIV OF SCI & TECH +2

Large model navigation method guided by historical topological graph based on manifold perception

The invention discloses a manifold perception-based large model navigation method guided by a historical topological graph, and relates to a computer vision technology. The method aims at solving the challenges that in the navigation process, long-distance reasoning experience is insufficient, instruction fragments and dynamic visual observation are difficult to align, and large model reasoning is prone to illusion interference. Firstly, a large model based on an encoder-decoder structure is used for supplementing historical information coding for a visual observation sequence, and therefore global topological information guidance is provided for long-distance reasoning. And secondly, in order to effectively solve the problem that large model reasoning is subjected to illusion interference, significant space-time differences in a visual observation sequence are mined by using a multi-curvature manifold, so that the large model can accurately describe the current environment and make a decision according to a visual reference object in thinking. Besides, in order to strengthen the perception capability of the large model to the space structure and establish a graph self-attention mechanism, the node distance embedded in the constructed historical topological graph is combined with the visual similarity so as to model the space relationship between the nodes.
Owner:WENZHOU TAIYI INTELLIGENT TECHNOLOGY CO LTD +2