Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

248 results about "Feed forward network" patented technology

A fully connected feed forward network is one in which every neuron in a particular layer is connected to every neuron in the subsequent layer, and in which information flows in one direction only, from input to output.

Complex scene-oriented AI large model lightweight deployment method

The invention provides a complex scene-oriented AI large model lightweight deployment method, and relates to the technical field of edge computing, and the method comprises the steps: carrying out the structured pruning of a pre-trained Transform network based on the attention head importance score, carrying out the dynamic sparsification of the activation state of a feedforward network according to the input tensor entropy value, employing the dynamic mixing precision quantization, and carrying out the reconstruction of an AI large model. Obtaining network parameters after pruning quantization; deploying the pruned and quantized network parameters to an edge computing device, distributing a feature extraction operator to a neural network processor through a heterogeneous computing scheduler, and unloading a classification operator to a multi-core central processing unit; and managing an on-chip memory in combination with a virtual memory paging mechanism, realizing zero-copy data transmission by utilizing a direct memory access controller, and outputting a reasoning result tensor. According to the method, efficient and reliable operation of the large model at the resource-constrained edge node is realized.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD

End-side multi-mode large model accelerated reasoning method and system

The invention provides an end-side multi-modal large model accelerated reasoning method and system, and the method comprises the steps: carrying out the two-stage screening and rearrangement of visual tokens based on the CLS attention and text-to-visual attention in a visual encoder and pre-filling stage, and constructing a sparse attention and sparse key value cache; in a decoding stage, an important neuron set is judged according to activation gating or historical statistics, only a corresponding feedforward network weight is pulled and calculated, missed weights are loaded on demand through asynchronous I / O, and hot neurons are maintained in a high-speed memory to utilize model sparsity, so that video memory / memory occupancy and calculation overhead are remarkably reduced on an end side; throughput and time delay performance are improved. According to the method, the internal memory and computing resources required by reasoning of the multi-modal large language model are reduced from two dimensions by utilizing the endogenous sparsity of the end-side large language model in input and the model, so that a higher reasoning speed is achieved by utilizing fewer resources on the premise of keeping the size of the model unchanged, and the performance of the whole system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Medical image classification method and system based on multi-scale spatial state modeling

The invention discloses a medical image classification method and system based on multi-scale spatial state modeling, and the method comprises the steps: firstly dividing an input medical image into a plurality of non-overlapping image blocks, and mapping the non-overlapping image blocks to a feature space through a learnable linear projection layer to obtain an initial feature map; then, multiple layers of stacked MS-SMamba blocks are used for carrying out layer-by-layer feature extraction, each MS-SMamba block comprises a main branch, an auxiliary branch, a dynamic gating fusion network, a residual error connection unit and a feedforward network, and long-range dependency relation capture and multi-scale feature fusion are achieved; and finally, processing the last-layer output feature map through a global feature aggregation and classification module, generating a global feature vector, and outputting a classification result. According to the method, the capturing capability of complex pathological features in the medical image is improved, the calculation efficiency and clinical applicability are improved, and the method is suitable for scenes such as disease screening and auxiliary decision making in medical image diagnosis.
Owner:XIANGJIANG LAB

Physical-data dual-drive retaining wall catastrophe robustness evaluation method

The invention discloses a physical-data dual-drive retaining wall catastrophe robustness evaluation method, and relates to the technical field of structure catastrophe evaluation and computer science crossing. Aiming at the problems that an existing retaining wall catastrophe robustness evaluation method cannot correct physical model errors in real time and is insufficient in accuracy, the invention provides a retaining wall catastrophe process robustness evaluation method driven by fusion of a physical model and a neural network. The method comprises the following steps: 1, carrying out time discretization and parameter vectorization on retaining wall catastrophe monitoring parameters; 2, constructing a performance attenuation function based on the physical model; 3, updating an implicit damage state in combination with a neural network Cell module; 4, correcting an error through a feedforward network to obtain a total performance attenuation function; and 5, finally calculating a robustness index and outputting a corresponding robustness grade, thereby realizing quantitative evaluation of the retaining wall catastrophe robustness. Reliable quantitative support can be provided for retaining wall design optimization, operation monitoring and post-disaster recovery evaluation.
Owner:TONGJI UNIV

Feature hierarchical attention fusion and dynamic optimization method for single-stage target detection

The invention discloses a feature hierarchical attention fusion and dynamic optimization method for single-stage target detection, and relates to the technical field of computer vision. The method comprises the steps of collecting image data, preprocessing an image and dividing the image into a training set, a test set and a verification set; the HAF-DRFPN neck network uses dynamic up-sampling single pixel point sampling to recover the feature resolution; feature extraction is optimized by using a gating residual mechanism and depth separable convolution; a feature refining feed-forward network is introduced through multi-layer perception generation weight, and feature expression is enhanced; respectively capturing and fusing superficial details and deep semantic information by utilizing channel and space attention and cross attention; a multi-branch dynamic sampling stacking framework is adopted, and cross-scale features are adaptively fused and optimized; multi-module cooperative processing is carried out on the whole architecture, and multi-scale features with higher discrimination are provided for the detection head; according to the method, the YOLO12 trunk and the detection head are connected with the HAF-DRFPN to form the model HAF-DRNet suitable for target detection, so that the flexibility and accuracy of single-stage target detection are improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Real haze image defogging method based on haze degradation model

The invention discloses a real haze image defogging method based on a haze degradation model. The method comprises the following steps: constructing a haze degradation model fusing multiple scattering effects and multiple image degradation factors; using the haze degradation model to construct a training data set including the clear image and the corresponding pseudo haze image; constructing a defogging network for a real haze scene; training the defogging network by adopting the training data set until a preset loss function is converged; and inputting a to-be-defogged image into the trained defogging network to obtain a defogging result. According to the haze degradation model constructed by the invention, the difference between a synthetic domain and a real domain is effectively relieved; a designed space-frequency hybrid module improves the adaptability of the model to complex degradation characteristics; the prior-guided feed-forward network fully excavates and fuses dark channel prior information, and the sensing and modeling capability of the model to the haze area is effectively enhanced.
Owner:NAT UNIV OF DEFENSE TECH

Super-resolution remote sensing image reconstruction method, system and equipment based on frequency domain enhancement

The invention belongs to the technical field of image data processing, and particularly relates to a super-resolution remote sensing image reconstruction method, system and device based on frequency domain enhancement, and the method comprises the steps: S1, extracting the shallow features of a low-resolution remote sensing image; s2, inputting the shallow features into a plurality of cascaded frequencies for interactive processing, performing double-branch processing on the input features, performing inverse transformation after radial weighting on different frequency components in a frequency domain to obtain first features, and obtaining second features through depth separable convolution, an activation function, a selection scanning module and layer normalization; fusing the two branch features according to the weight, performing jump connection with the original features, performing enhancement through a feedforward network, performing repeated execution for a set number of times, and performing convolution and residual connection to obtain deep features; and S3, fusing the deep features and the shallow features, and outputting a super-resolution image through convolution and pixel rearrangement up-sampling. According to the method, texture details and edge contours of the remote sensing image can be more accurately reconstructed while the structural consistency is kept.
Owner:YANTAI UNIV

Model illusion detection method and device based on internal state fusion and medium

The invention discloses a model illusion detection method and device based on internal state fusion and a medium, and relates to the technical field of natural language processing. The method comprises the following steps: extracting multi-modal features in a forward propagation process of a target large language model, wherein the multi-modal features comprise a hidden layer embedding feature, an attention feature, a feedforward network activation feature and a text feature; aligning the multi-modal features to the lexical element length of the generated text through an interpolation method, and calculating the weight of the position of the lexical element corresponding to the multi-modal features; performing weighted fusion on the multi-modal features according to the weights to generate a fusion feature sequence; and constructing the fusion feature sequence into a graph structure, reasoning the graph structure by using a multi-layer attention network, and outputting a lexical-level illusion probability through a classification head. According to the method, the target model parameters do not need to be modified, and high-precision detection can be completed only through single-time forward propagation.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

SAR ship image detection method based on improved YOLOv11 model

The invention discloses an SAR ship image detection method based on an improved YOLOv11 model, an improved C2PSA module C2DyMoETAttn is introduced, the core innovation point is that a PSABlock module is replaced by a DyMoETAttnBlock module, the DyMoETAttnBlock module fuses a Dynamic Tanh activation function, a Mona module, a TSSA attention mechanism and a frequency domain enhancement feedforward network (EDFFN), multi-dimensional modeling and robust enhancement of features are realized, and the detection accuracy is improved. The feature expression capability and the noise suppression performance under the background of small targets and complex sea clutters are effectively improved; in the deep feature fusion stage, a C3k2 module of YOLOv11 is optimized, an ScConv structure is introduced, adaptive fusion of space and channel features is realized through a joint reweighting mechanism of SRU and CRU, and the multi-scale target discrimination capability and feature selectivity are enhanced; on the bounding box regression layer, a Focaler-MPDIOU loss function is provided, and a Focaler-IoU sample difficulty adaptive mechanism is combined with MPDIOU positioning matching constraint, so that the learning ability of the model for small targets and shielded targets is enhanced, and the bounding box positioning precision and convergence stability are improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY +1

Signal modulation identification method fused to Transform and hybrid expert mechanism

The invention discloses a signal modulation identification method fused to Transform and a hybrid expert mechanism, and belongs to the technical field of wireless communication and artificial intelligence crossing. The method solves the problem that the model capacity and the adaptability to complex signals are limited by a fixed feed-forward network in the automatic modulation identification of the conventional Transform model. According to the technical scheme, the method comprises the steps of inputting dual-channel time sequence signals of in-phase components and orthogonal components; extracting preliminary features through convolution operation; the features are input into a Transform architecture integrating an attention mechanism and a hybrid expert mechanism; in the hybrid expert mechanism, calculating a routing weight through a gating network with noise routing and sparsely activating an expert network; merging expert outputs according to routing weights to obtain depth features; and finally, outputting a modulation identification result through the classifier. According to the method, the calculation efficiency is maintained while the model parameter scale is expanded, and the recognition accuracy in a complex channel environment is improved.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Big-model multi-client collaborative positioning knowledge editing method based on federal learning

The invention discloses a large-model multi-client collaborative positioning knowledge editing method based on federated learning. The method comprises the following steps: 1) constructing an LLM model; 2) constructing a FedLEKE task; (3) executing a FedLEKE task, optimizing the hidden state of a Transform layer so as to finely adjust the weight of a feed-forward network in the LLM model, and updating the LLM model; 4) each client generates an intermediate knowledge vector and uploads the intermediate knowledge vector to the server; 5) in a predefined time slot ti belongs to T, the server distributes the stored MKVs to each client; and 6) each client dynamically retrieves related intermediate knowledge vectors in the server based on cosine similarity, re-edits the knowledge vectors, and locates and modifies related parameters in the LLM model according to the re-edited knowledge vectors. According to the invention, more than 96% of performance of non-federated LEKE is reserved, and the performance of the non-federated LEKE is obviously twice better than that of a base line based on FedAvg.
Owner:CHONGQING UNIV

User input information processing method and system for model reasoning stage

The invention discloses a user input information processing method and system for a model reasoning stage, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining a plurality of complete prompts inputted by a user; periodically acquiring hardware resource information and current task information; determining hardware resource allocation information corresponding to the complete prompt; pre-executing parallel computing operation on the complete prompt according to the hardware resource allocation information in the pre-filling stage; according to the hardware resource allocation information, calling the first type of computing equipment with the computing power reaching a preset computing power threshold value in the feed-forward network stage to read the intermediate characteristics from the preset resource pool for decoding and storage operation, wherein the computing power of the first type of computing equipment reaches the preset computing power threshold value; and calling the second type of computing equipment with the memory access performance meeting the preset memory access performance condition in the attention stage to read data from the preset resource pool to perform attention computing storage operation, obtaining a model reasoning result and outputting the model reasoning result, thereby solving the technical problem of hardware resource waste, and achieving the technical effects of improving the resource utilization rate and reasoning efficiency.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Conformer-based mixed self-attention and convolution improved speech recognition method and system

The invention discloses a Conformer-based mixed self-attention and convolution improved speech recognition method and system. The recognition method comprises the steps of obtaining speech feature representation; the speech feature representation is input to a Conformer encoder for encoding processing, high-dimensional feature representation is obtained, and the Conformer encoder comprises a convolution module, an attention linear enhancement module and a feedforward network residual scaling module; according to the high-dimensional feature representation, bidirectional decoding is carried out through a bidirectional Transform decoder, and bidirectional decoding output is obtained; and carrying out merging processing on the bidirectional decoding output to obtain a speech recognition result. According to the method, the modeling capacity for long-time dependence and global context is enhanced through a bidirectional decoder architecture, the processing capacity of the model in a long sequence is improved through ALiBi relative position coding, and the model can be flexibly adjusted in a noise environment through a right decoder weighting adjustment mechanism.
Owner:HEBEI UNIV OF ENG

Fault tree Boolean function equivalent mapping method based on untrained neural network

The invention discloses a fault tree Boolean function equivalent mapping method based on an untrained neural network, and relates to the field of fault tree analysis. In order to solve the problems that in the prior art, a Boolean function mapping structure is not beneficial to parallel expansion, the calculation efficiency is limited, and the Boolean function mapping structure is difficult to efficiently realize on high-parallel platforms such as a GPU, the invention provides a method for generating topological structure data by analyzing a fault tree model; the basic events, the intermediate events and the top events are mapped into neurons of an input layer, a hidden layer and an output layer respectively, a feedforward network with fixed weight and bias is constructed, and a logic activation function is defined in nodes to realize Boolean logic propagation. The input layer receives a basic event state vector, outputs a top event result through forward propagation, and can realize large-scale Boolean function mapping on a parallel platform through batch input matrixes. The method is suitable for reliability analysis, minimum cut set simplification, top event probability calculation, parallelization fault tree solving and the like of a large-scale complex system.
Owner:HARBIN ENG UNIV

Tibetan multi-dialect speech recognition system and method

The invention provides a Tibetan multi-dialect speech recognition system, and the system comprises a Tibetan self-supervision model which is used for extracting hidden layer speech representation of an original speech signal to obtain a first input feature vector; the original voice signal is Tibetan dialect voice; the dialect feature extraction module is used for extracting dialect features of multiple pieces of dialect information in a one-hot coding form to obtain a second feature vector; the multiple pieces of dialect information at least comprise Anduo, Kangba and defense and Tibetan dialects; the encoder-decoder model is used for routing the corresponding dialect expert module according to the first input feature vector and the second input feature vector and outputting a Tibetan text; and the dialect expert module is used for routing each input feature to the most relevant feedforward network for joint learning. In a multi-dialect environment, unique characteristics of different dialects can be distinguished and understood more accurately, corresponding dialect experts can be effectively and dynamically called according to characteristics of input signals by introducing the hybrid expert module, and processing efficiency and accuracy are remarkably improved.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Hyperspectral remote sensing image water body extraction method based on CNN-Transform mixed architecture

The invention relates to a hyperspectral remote sensing image water body extraction method based on CNN-Transform mixed architecture, and belongs to the technical field of remote sensing image semantic segmentation. Comprising the following steps: preprocessing a hyperspectral image; constructing a feature extraction network, and obtaining multi-scale spectrum-space features; an enhanced multi-scale context attention module is designed, and local detail features of multi-scale cavity convolution and global context features of cross window global attention are fused; an enhanced residual feedforward network module is designed, and the local feature extraction capability is enhanced by using depth separable convolution; and constructing a dual-path feature fusion module, and fusing deep semantics and shallow detail features through a space attention path and a channel interaction path. According to the method, the problem of missegmentation caused by easy confusion of the water body and the shadow / building in the hyperspectral image and the problem of difficulty in effective fusion of local details and global context information are effectively solved, and the fine water body extraction precision and the boundary definition are remarkably improved.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY

Method and system for detecting corrosion state of overhead ground wire based on deep learning

The invention discloses an overhead ground wire corrosion state detection method and system based on deep learning, and the method comprises the steps: obtaining an overhead ground wire surface image, and constructing a semantic segmentation data set comprising a background and different corrosion grade categories; constructing a deep learning semantic segmentation model based on an encoder-decoder architecture; the encoder comprises a deformable convolution attention module and a multi-scale feed-forward network so as to extract corrosion features; the decoder adopts a full convolution structure, fuses the features extracted by the encoder in each stage, and outputs a semantic segmentation result; training the deep learning semantic segmentation model by using the semantic segmentation data set; and detecting the surface of the on-site overhead ground wire by using the trained deep learning semantic segmentation model, and outputting a corrosion level semantic segmentation result of the surface image of the on-site overhead ground wire. According to the method, a deep learning semantic segmentation model based on an encoder-decoder architecture is combined, and semantic segmentation identification of the corrosion level of the corrosion area of the overhead ground wire is realized.
Owner:ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD

Video generation method and system based on three-dimensional sparse attention

The invention discloses a video generation method and system based on three-dimensional sparse attention, and belongs to the field of video generation. The method comprises the following steps of: performing three-dimensional partitioning on an input video feature according to the size of a time dimension block and the number of space dimension blocks, and performing rearrangement index on each three-dimensional sub-block; adopting a block-level Top-K attention mechanism, and only selecting a key block for each query; carrying out attention calculation on query, key and value features of the key block, and carrying out normalization output on a calculation result; utilizing FlashAttention to execute efficient variable-length attention calculation on the rearranged query, key and value characteristics, and recovering an original sequence; video output features are generated through the residual connection and the feed-forward network. According to the method, the long-sequence video attention calculation complexity and video memory occupation are remarkably reduced, the time-space dependence modeling efficiency is improved, and the method is suitable for large-scale video generation and understanding tasks.
Owner:ZHEJIANG UNIV +1

Edge end large language model reasoning acceleration method and accelerator

The invention relates to the technical field of network acceleration, and discloses an edge-end large language model reasoning acceleration method and accelerator, and the method comprises the following steps: reconstructing a calculation process of a decoding stage, and carrying out the deep fusion of a multi-head attention mechanism and the calculation operation of a feedforward network; the weight and key value data are stored in HBM, and the coefficient and the accumulated attention score are stored in DDR; for linear matrix calculation, a unified matrix calculation unit is used for executing multi-precision matrix operation; for nonlinear function calculation, a mathematical transformation and linear fitting method is adopted, a Softmax function is converted into operation with 2 as the bottom through a bottom conversion formula, and truncation and third-order linear fitting are conducted on a Sigmoid function; and constructing a key value screening algorithm based on the accumulated attention score, dynamically adjusting a key value storage position, maintaining a recent key value cache region and an important key value cache region in a limited cache space, and realizing key value efficient cache in long text reasoning.
Owner:CENT SOUTH UNIV

Intelligent agent action sequence generation method based on visual language action model

The invention discloses an agent action sequence generation method based on a visual language action model, belongs to the technical field of artificial intelligence, and can solve the problems that an existing method is highly dependent on an additional observation mode and an auxiliary module, needs additional training and fine adjustment, and is limited in deployment flexibility and expandability. The method comprises the following steps: S1, determining an output hidden state of a feedforward network of a visual language action model according to observation information and a language instruction of an intelligent agent; s2, determining the probability distribution of each action mark corresponding to the current feed-forward network, and determining the uncertainty of the current feed-forward network; s3, determining observation characteristics to be injected according to the uncertainty of the current feed-forward network, and updating the output hidden state of the next feed-forward network by using the observation characteristics; and S4, taking the next feed-forward network as the current feed-forward network, repeating the steps S2 and S3, and generating an agent action sequence according to the output hidden state of the last feed-forward network. The method is used for generating the agent action sequence.
Owner:ZHONGKE FIFTH CENTURY (HANGZHOU) INTELLIGENT TECHNOLOGY CO LTD +1

Garbage detection method and system based on improved YOLOv11

The invention relates to a garbage detection method and system based on improved YOLOv11, belongs to the field of computer vision and artificial intelligence application, and aims at solving the problems that in existing garbage detection, the small target recognition rate is low, the complex background adaptability is poor, and the reasoning efficiency is insufficient. According to the method, on the basis of an improved YOLOv11 framework, a down-sampling attention enhancement module, a weighted convolution feature extraction module and an enhanced attention feedforward network module are introduced, a multi-scale feature extraction and fusion mechanism is constructed, and accurate classification and positioning of garbage targets are achieved. The system is composed of a backbone network and a detection head which are responsible for feature extraction and fusion decoding respectively, and has the advantages of being high in context modeling capacity, accurate in feature expression and efficient in calculation. According to the method, the detection precision and speed are remarkably improved while the light weight of the model is kept, and the method is suitable for real-time deployment scenes such as intelligent garbage classification and environment monitoring.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Cognitive-driven VLA world model automatic driving system with mixed expert and truncated diffusion

The invention relates to the technical field of artificial intelligence and automatic driving, and discloses a mixed expert and cut-off diffusion cognitive driven VLA world model automatic driving system, which comprises a unified autoregression Transform backbone, a cognitive driven causal attention module and a mixed expert feed-forward network are included in the layer of the unified autoregression Transform backbone, the backbone is used for jointly generating a future world token and an instant action token, an anchor point token is generated, then a truncation diffusion sampling module for executing deep truncation denoising sampling is initialized through the anchor point token, and an instant action token is finally generated. Through a causal reasoning mechanism of cognitive driving, the system is upgraded from perception driving to cognitive driving, and the decision-making problem of automatic driving in a complex interaction scene is solved.
Owner:GUANGZHOU SMART BODY TECH CO LTD

Deep learning CO2 near-real-time inversion method and system fusing multi-scale features

The invention discloses a deep learning CO2 near-real-time inversion method and system fused with multi-scale characteristics, and the method comprises the steps: inputting the observed spectral data into an inversion model, carrying out the in-orbit near-real-time calculation, and outputting a satellite data inversion result; the inversion model comprises a two-stage full connection projection, a Transform encoder, a learnable attention pooling and a regression decoder which are connected in sequence; wherein each layer of the encoder comprises two sequentially connected sub-layers, namely a multi-head self-attention layer and a feedforward network layer, a residual connection layer and a normalization layer are introduced behind each sub-layer, in the multi-head self-attention layer, rotation position coding is applied to Q and K before dot product attention is entered, and relative displacement is implicitly injected into attention weight calculation through phase rotation; according to the inversion method, satellite observation is taken as a characteristic, and ground in-situ observation is taken as a supervision CO2 data driving inversion thought, so that the calculation efficiency is maintained, and meanwhile, the modeling capability of a nonlinear relationship between characteristics is improved.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Music source separation method and system based on hybrid expert self-attention network

The invention discloses a music source separation method and system based on a hybrid expert self-attention network, and belongs to the technical field of audio signal processing and deep learning. The method comprises the following steps: receiving a time domain mixed audio signal, and obtaining a complex frequency spectrum through short-time Fourier transform; the method comprises the following steps: estimating a complex ideal proportion mask through an improved separator network MoEFormer, respectively modeling a time domain dependency relationship and a frequency domain dependency relationship in the separator network by adopting an axial attention mechanism, and replacing a feedforward network in a standard Transformer with a hybrid expert layer; the hybrid expert layer comprises an expert network specially designed for drum sound, Bass and human voice characteristics, and a self-adaptive weight distribution mechanism based on gating routing; and finally, reconstructing the separated time domain signal through inverse short-time Fourier transform. Through expert specialization and feature adaptive fusion, the problem of multi-sound-source feature confusion is effectively solved, and low calculation complexity is kept while the separation precision is improved.
Owner:JINLING INST OF TECH

Three-dimensional point cloud semantic segmentation method and device

The invention discloses a three-dimensional point cloud semantic segmentation method and device. The method comprises the steps of obtaining three-dimensional point cloud data to be segmented; and inputting the semantic tag into a pre-trained segmentation model to obtain the semantic tag. The segmentation model adopts a local geometric enhancement module, a dynamic attention module and a feedforward network which are connected in sequence. The core of the method is that anisotropic graph message passing is carried out through a graph neural network, and local geometric features are explicitly enhanced; and through a grouping vector attention mechanism, a position modulation item based on a relative position between points and a density modulation item based on a local density are introduced in a combined manner, and an attention weight is dynamically adjusted, so that adaptive feature aggregation of a point cloud heterostructure and complex geometry is realized. And the segmentation precision and contour sharpness in sparse and dense mixed regions and high-curvature boundaries are effectively improved.
Owner:THREE GORGES JINSHAJIANG CHUANYUN HYDROPOWER DEV CO LTD

Key value neural network architecture

The present disclosure relates to techniques for improving inference efficiency and memory utilization in transformer-based neural networks. A modified architectural design is introduced that decouples key and value matrix generation from inter-layer dependencies, enabling statically computed or parallelizable projections across layers. The disclosed approach may eliminate the need for layer-wise prefilling, support linear-time inference, and substantially reduce the memory footprint associated with key-value (KV) caching. The disclosed architecture may use shared or layer-specific projections, with a single KV-cache serving all or subsets of layers. In some embodiments, a non-linear transformation (e.g., implemented via a feed-forward network), may preprocess input embeddings prior to query, key, and value generation. A lookup table of transformed embeddings may be precomputed to further accelerate inference. The disclosed system can enhance scalability and may allow deployment of large models on resource-constrained hardware, offering practical benefits for latency-sensitive applications and long-context processing in transformer-based models.
Owner:WRITER INC

A State-Space Model-Based Medical Image Segmentation Method for Abdominal Multi-Organs

PendingCN122089756Aresolve integritySolve the problem of mutual invasion between organsImage analysisBiological modelsComputation complexityFeed forward network
This invention belongs to the field of medical image processing technology, specifically relating to a method for abdominal multi-organ medical image segmentation based on a state-space model. Addressing the shortcomings of existing technologies such as the difficulty of CNNs in modeling long-distance dependencies, the high computational complexity of Transformers, and the large semantic gaps, incomplete segmentation contours, and easy organ encroachment issues in VM-UNet skip connections, this solution makes key improvements: It constructs an improved VM-UNet-Skip architecture, placing skip connections before downsampling to reduce the semantic gap between the small decoder features; it introduces a multi-scale global-local information aggregation module, capturing global anatomical dependencies through multi-head Mamba units and enhancing local detail representations with a convolutional gated feedforward network; and it relies on the VMamba encoder-decoder to achieve efficient feature extraction and reconstruction. The method flow includes: constructing initial feature representations through patch embedding layers, extracting multi-scale hierarchical features through the encoder, enhancing features through an information fusion module, and fusing and mapping the enhanced features to obtain the segmentation result through the decoder. This invention effectively improves the segmentation accuracy of abdominal multi-organs, solves the problems of insufficient contour integrity and organ encroachment, while reducing the number of model parameters and computational complexity, providing reliable technical support for clinical diagnosis and surgical planning.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Hyperspectral image fusion method and system based on variance guidance and heavy tail estimation

The invention relates to the technical field of hyperspectral image processing, and particularly discloses a hyperspectral image fusion method and system based on variance guidance and heavy tail estimation, and the method comprises the steps: obtaining a low-spatial-resolution hyperspectral image and a panchromatic image; performing up-sampling on the low-spatial-resolution hyperspectral image to obtain an up-sampled image; obtaining an absolute difference value weight according to the channel-by-channel variance of the low-spatial-resolution hyperspectral image and the up-sampling image; performing channel-by-channel spatial fusion by means of absolute difference weight, and injecting high-frequency information of the panchromatic image into the up-sampling image to obtain a preliminary fusion image; performing spectral correlation correction learning on the preliminary fusion image through an attention mechanism to obtain spectral mixed output; through a feedforward network, obtaining an up-sampling image subjected to variance guide processing; and obtaining a high-spatial-resolution hyperspectral image through a residual block and convolution. According to the invention, accurate selection and fusion of the spatial-spectral features are realized.
Owner:TIANJIN POLYTECHNIC UNIV

Low-orbit star network target prediction method based on transfer learning

The invention relates to a low-orbit star network target prediction method based on transfer learning, and belongs to the technical field of detection and tracking. A real target simulation flight path is generated, and spatial position information of the target at each moment is accurately obtained; an observation error conforming to normal distribution is introduced into the simulation data so as to simulate an observation value of a satellite when the target flies, namely, the spatial position information of the observed target at each moment is obtained; a prediction model based on a self-attention mechanism and a feedforward network is constructed, training is carried out, and spatial position information of the target at each moment is predicted; comparing the predicted value with a real track to obtain a compensation value; through iterative correction, a predicted value is gradually close to a real value; outputting target high-precision indication data; the trajectory prediction process is repeatedly executed, and low-orbit star network target prediction is completed; the problems that real satellite trajectory data is difficult to obtain and a corresponding verification method is lacked for trajectory prediction capability simulation verification of a low-orbit star network detection target are solved.
Owner:CHINA ACAD OF LAUNCH VEHICLE TECH

Infrared and visible light unified basic model pre-training method and device based on LoRA

The invention provides an infrared and visible light unified basic model pre-training method based on LoRA, and the method comprises the steps: obtaining a large-scale visible light pre-trained VIT model, and constructing a bimodal aligned image pair data set; freezing a VIT weight, and interpolating a trainable low-rank adaptation matrix in an encoder attention layer, a feedforward network layer and an output projection layer; inputting a bimodal image, and extracting visible light features from the frozen VIT network to generate a binary block-level pseudo tag; optimizing the cosine similarity matrix of the infrared and frozen visible light features and the consistency of a pseudo label through a first loss function to realize cross-modal alignment training, and optimizing the similarity matrix of the training branch and the frozen VIT visible light features and the consistency of the pseudo label through a second loss function to realize visible light knowledge maintenance training; and combining the low-rank matrix and the VIT weight, and carrying out segmentation or detection head fine tuning to obtain a target model. Therefore, the calculation and storage overhead of pre-training can be obviously reduced, and the model detection precision is obviously improved.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI