Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

301 results about "Feed forward network" patented technology

A fully connected feed forward network is one in which every neuron in a particular layer is connected to every neuron in the subsequent layer, and in which information flows in one direction only, from input to output.

Complex scene-oriented AI large model lightweight deployment method

The invention provides a complex scene-oriented AI large model lightweight deployment method, and relates to the technical field of edge computing, and the method comprises the steps: carrying out the structured pruning of a pre-trained Transform network based on the attention head importance score, carrying out the dynamic sparsification of the activation state of a feedforward network according to the input tensor entropy value, employing the dynamic mixing precision quantization, and carrying out the reconstruction of an AI large model. Obtaining network parameters after pruning quantization; deploying the pruned and quantized network parameters to an edge computing device, distributing a feature extraction operator to a neural network processor through a heterogeneous computing scheduler, and unloading a classification operator to a multi-core central processing unit; and managing an on-chip memory in combination with a virtual memory paging mechanism, realizing zero-copy data transmission by utilizing a direct memory access controller, and outputting a reasoning result tensor. According to the method, efficient and reliable operation of the large model at the resource-constrained edge node is realized.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD

End-side multi-mode large model accelerated reasoning method and system

The invention provides an end-side multi-modal large model accelerated reasoning method and system, and the method comprises the steps: carrying out the two-stage screening and rearrangement of visual tokens based on the CLS attention and text-to-visual attention in a visual encoder and pre-filling stage, and constructing a sparse attention and sparse key value cache; in a decoding stage, an important neuron set is judged according to activation gating or historical statistics, only a corresponding feedforward network weight is pulled and calculated, missed weights are loaded on demand through asynchronous I / O, and hot neurons are maintained in a high-speed memory to utilize model sparsity, so that video memory / memory occupancy and calculation overhead are remarkably reduced on an end side; throughput and time delay performance are improved. According to the method, the internal memory and computing resources required by reasoning of the multi-modal large language model are reduced from two dimensions by utilizing the endogenous sparsity of the end-side large language model in input and the model, so that a higher reasoning speed is achieved by utilizing fewer resources on the premise of keeping the size of the model unchanged, and the performance of the whole system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Medical image classification method and system based on multi-scale spatial state modeling

The invention discloses a medical image classification method and system based on multi-scale spatial state modeling, and the method comprises the steps: firstly dividing an input medical image into a plurality of non-overlapping image blocks, and mapping the non-overlapping image blocks to a feature space through a learnable linear projection layer to obtain an initial feature map; then, multiple layers of stacked MS-SMamba blocks are used for carrying out layer-by-layer feature extraction, each MS-SMamba block comprises a main branch, an auxiliary branch, a dynamic gating fusion network, a residual error connection unit and a feedforward network, and long-range dependency relation capture and multi-scale feature fusion are achieved; and finally, processing the last-layer output feature map through a global feature aggregation and classification module, generating a global feature vector, and outputting a classification result. According to the method, the capturing capability of complex pathological features in the medical image is improved, the calculation efficiency and clinical applicability are improved, and the method is suitable for scenes such as disease screening and auxiliary decision making in medical image diagnosis.
Owner:XIANGJIANG LAB

Magnetic particle imaging resolution improving method and system based on frequency domain information filtering

The invention provides a magnetic particle imaging resolution improving method and system based on frequency domain information filtering, and relates to the technical field of magnetic particle imaging. Comprising the following steps: inputting original magnetic particle imaging data into a Transform model, and selectively filtering frequency domain information in the magnetic particle imaging data through a fusion frequency domain discrimination feedforward network module in an encoder to obtain the output of the encoder; inputting the output of the encoder into a bottleneck layer, aggregating global and local information through a multi-scale attention mechanism module, and then selectively filtering frequency domain information in the magnetic particle imaging data again to obtain the output of the bottleneck layer; inputting the output of the bottleneck layer into a decoder with the same structure as the bottleneck layer to obtain deep features; inputting the deep features into a convolutional layer to obtain a residual image; and calculating the sum of the original magnetic particle imaging data and the residual image to obtain a reconstructed image. According to the invention, the overall imaging resolution of the MPI image under the low-gradient field acquisition condition is improved.
Owner:SHANDONG UNIV

Power transmission line insulator multi-mode multi-type defect detection method based on MEWL-YOLO model

The invention discloses a power transmission line insulator multi-modal multi-type defect detection method based on an MEWL-YOLO model. The method comprises the following steps: acquiring a multi-modal power transmission line data set; constructing an insulator multi-type defect detection model of the MEWL-YOLO model; the method specifically comprises the steps that on the basis of a YOLOv10n network framework, a multi-head attention mechanism and a gated convolution feedforward network are combined, a multi-head gated convolution feedforward network MGDB module is designed, and a C2K layer is optimized; a multi-scale efficient attention mechanism (EMA) is inserted into the Neck part of the YOLOv10n network; an Inner-Wise-MPDIOU loss function is adopted to replace a traditional loss function; a detection head in a YOLOv10n network is replaced by an LSCD lightweight detection head, and defect detection is carried out on the power transmission line insulator based on an insulator multi-type defect detection model of a constructed MEWL-YOLO model. According to the method, the MEWL-YOLO model is provided to realize defect detection of the multi-modal power transmission line insulator image, the detection effect of multi-modal and multi-type defects is improved, and the method has good practical application value.
Owner:XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER

Physical-data dual-drive retaining wall catastrophe robustness evaluation method

The invention discloses a physical-data dual-drive retaining wall catastrophe robustness evaluation method, and relates to the technical field of structure catastrophe evaluation and computer science crossing. Aiming at the problems that an existing retaining wall catastrophe robustness evaluation method cannot correct physical model errors in real time and is insufficient in accuracy, the invention provides a retaining wall catastrophe process robustness evaluation method driven by fusion of a physical model and a neural network. The method comprises the following steps: 1, carrying out time discretization and parameter vectorization on retaining wall catastrophe monitoring parameters; 2, constructing a performance attenuation function based on the physical model; 3, updating an implicit damage state in combination with a neural network Cell module; 4, correcting an error through a feedforward network to obtain a total performance attenuation function; and 5, finally calculating a robustness index and outputting a corresponding robustness grade, thereby realizing quantitative evaluation of the retaining wall catastrophe robustness. Reliable quantitative support can be provided for retaining wall design optimization, operation monitoring and post-disaster recovery evaluation.
Owner:TONGJI UNIV

Double-branch pyramid attention time sequence prediction method

The invention discloses a double-branch pyramid attention time sequence prediction method, which belongs to the field of time sequence prediction, and comprises the following steps of: supplementing and standardizing input sequence deletion, and embedding and aligning time covariables; a horizontal branch coding time domain and a vertical branch coding variable are dynamically fused to obtain a unified feature; multi-scale representation is formed through multi-layer convolution down-sampling, and information interaction is completed through local-cross-scale mask attention. The result is subjected to pyramid residual normalization and feedforward network updating, a global context is generated through multilayer stacking, and a linear head outputs future multi-step prediction. The complexity of the method is O (L log L), long-range dependence and a fine-grained mode are captured at the same time, the method is remarkably superior to thirteen existing models on five data sets including ECL, Exchange and Trafic, and the method has the advantages of being high in precision, high in speed and easy to expand.
Owner:JIANGSU OCEAN UNIV

Method for converting trained language model into language model having architecture of mixture of experts and computing device using same

A processor-implemented method for converting a trained language model into a language model in an architecture of mixture of experts (MoE), and a computing device using the same is provided. The method for converting a trained language model into a language model in an architecture of mixture of experts using a computing device according to an embodiment of the disclosure may include dividing a plurality of layers included in a target language model and extracting a feed-forward network (FFN) included in each of the plurality of layers, generating an MoE block of the MoE language model, which corresponds to the feed-forward network, generating an input tensor, comparing output tensors between the feed-forward network and the MoE block for the input tensor to obtain a first loss, and updating a weight of the MoE block, based on the first loss.
Owner:SAMSUNG SDS CO LTD

Feature hierarchical attention fusion and dynamic optimization method for single-stage target detection

The invention discloses a feature hierarchical attention fusion and dynamic optimization method for single-stage target detection, and relates to the technical field of computer vision. The method comprises the steps of collecting image data, preprocessing an image and dividing the image into a training set, a test set and a verification set; the HAF-DRFPN neck network uses dynamic up-sampling single pixel point sampling to recover the feature resolution; feature extraction is optimized by using a gating residual mechanism and depth separable convolution; a feature refining feed-forward network is introduced through multi-layer perception generation weight, and feature expression is enhanced; respectively capturing and fusing superficial details and deep semantic information by utilizing channel and space attention and cross attention; a multi-branch dynamic sampling stacking framework is adopted, and cross-scale features are adaptively fused and optimized; multi-module cooperative processing is carried out on the whole architecture, and multi-scale features with higher discrimination are provided for the detection head; according to the method, the YOLO12 trunk and the detection head are connected with the HAF-DRFPN to form the model HAF-DRNet suitable for target detection, so that the flexibility and accuracy of single-stage target detection are improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Real haze image defogging method based on haze degradation model

The invention discloses a real haze image defogging method based on a haze degradation model. The method comprises the following steps: constructing a haze degradation model fusing multiple scattering effects and multiple image degradation factors; using the haze degradation model to construct a training data set including the clear image and the corresponding pseudo haze image; constructing a defogging network for a real haze scene; training the defogging network by adopting the training data set until a preset loss function is converged; and inputting a to-be-defogged image into the trained defogging network to obtain a defogging result. According to the haze degradation model constructed by the invention, the difference between a synthetic domain and a real domain is effectively relieved; a designed space-frequency hybrid module improves the adaptability of the model to complex degradation characteristics; the prior-guided feed-forward network fully excavates and fuses dark channel prior information, and the sensing and modeling capability of the model to the haze area is effectively enhanced.
Owner:NAT UNIV OF DEFENSE TECH

Super-resolution remote sensing image reconstruction method, system and equipment based on frequency domain enhancement

The invention belongs to the technical field of image data processing, and particularly relates to a super-resolution remote sensing image reconstruction method, system and device based on frequency domain enhancement, and the method comprises the steps: S1, extracting the shallow features of a low-resolution remote sensing image; s2, inputting the shallow features into a plurality of cascaded frequencies for interactive processing, performing double-branch processing on the input features, performing inverse transformation after radial weighting on different frequency components in a frequency domain to obtain first features, and obtaining second features through depth separable convolution, an activation function, a selection scanning module and layer normalization; fusing the two branch features according to the weight, performing jump connection with the original features, performing enhancement through a feedforward network, performing repeated execution for a set number of times, and performing convolution and residual connection to obtain deep features; and S3, fusing the deep features and the shallow features, and outputting a super-resolution image through convolution and pixel rearrangement up-sampling. According to the method, texture details and edge contours of the remote sensing image can be more accurately reconstructed while the structural consistency is kept.
Owner:YANTAI UNIV

Cerebral stroke early diagnosis model construction method and device, electronic equipment and storage medium

The invention provides a construction method and device of a cerebral apoplexy early diagnosis model, electronic equipment and a storage medium, which are applied to the field of model construction, and in a transformer training process, channel reduction or channel expansion is performed on a feedforward network based on an activation matrix and a current width of the feedforward network of transformer to adjust the width of the feedforward network, so that the accuracy of the cerebral apoplexy early diagnosis model is improved. On the premise of not sacrificing the model performance, the parameter quantity and video memory occupation are reduced, the computing resource overhead of a cerebral apoplexy early diagnosis system is reduced, and the adaptability is enhanced.
Owner:ATHENAEYES CO LTD

Model illusion detection method and device based on internal state fusion and medium

The invention discloses a model illusion detection method and device based on internal state fusion and a medium, and relates to the technical field of natural language processing. The method comprises the following steps: extracting multi-modal features in a forward propagation process of a target large language model, wherein the multi-modal features comprise a hidden layer embedding feature, an attention feature, a feedforward network activation feature and a text feature; aligning the multi-modal features to the lexical element length of the generated text through an interpolation method, and calculating the weight of the position of the lexical element corresponding to the multi-modal features; performing weighted fusion on the multi-modal features according to the weights to generate a fusion feature sequence; and constructing the fusion feature sequence into a graph structure, reasoning the graph structure by using a multi-layer attention network, and outputting a lexical-level illusion probability through a classification head. According to the method, the target model parameters do not need to be modified, and high-precision detection can be completed only through single-time forward propagation.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Multi-stage anomaly detection system based on multivariable time series data

The invention relates to the field of data anomaly detection, and particularly discloses a multi-stage anomaly detection system based on multivariable time sequence data, which comprises an input data acquisition module used for acquiring time sequence multivariable data; the MTS-Mixer spatio-temporal feature extraction module is used for extracting multi-level time sequence correlation features through a feature mixing technology and generating low-rank compression prediction features; the time sequence anomaly detection module is used for realizing anomaly positioning and classification by adopting alternate stacking of an Angle-Attention layer and a feed-forward network, expanding anomaly differences in combination with prior constraints and adversarial learning and fusing reconstruction errors and association differences to generate anomaly scores; and the attention mechanism soft thresholding module is used for suppressing noise, enhancing key signals and outputting high signal-to-noise ratio representation. According to the method, multiple technologies are integrated, and real-time anomaly detection and interpretable positioning of the high-dimensional time series data are realized through multi-stage collaborative optimization.
Owner:GUANGZHOU UNIVERSITY

SAR ship image detection method based on improved YOLOv11 model

The invention discloses an SAR ship image detection method based on an improved YOLOv11 model, an improved C2PSA module C2DyMoETAttn is introduced, the core innovation point is that a PSABlock module is replaced by a DyMoETAttnBlock module, the DyMoETAttnBlock module fuses a Dynamic Tanh activation function, a Mona module, a TSSA attention mechanism and a frequency domain enhancement feedforward network (EDFFN), multi-dimensional modeling and robust enhancement of features are realized, and the detection accuracy is improved. The feature expression capability and the noise suppression performance under the background of small targets and complex sea clutters are effectively improved; in the deep feature fusion stage, a C3k2 module of YOLOv11 is optimized, an ScConv structure is introduced, adaptive fusion of space and channel features is realized through a joint reweighting mechanism of SRU and CRU, and the multi-scale target discrimination capability and feature selectivity are enhanced; on the bounding box regression layer, a Focaler-MPDIOU loss function is provided, and a Focaler-IoU sample difficulty adaptive mechanism is combined with MPDIOU positioning matching constraint, so that the learning ability of the model for small targets and shielded targets is enhanced, and the bounding box positioning precision and convergence stability are improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY +1

Signal modulation identification method fused to Transform and hybrid expert mechanism

The invention discloses a signal modulation identification method fused to Transform and a hybrid expert mechanism, and belongs to the technical field of wireless communication and artificial intelligence crossing. The method solves the problem that the model capacity and the adaptability to complex signals are limited by a fixed feed-forward network in the automatic modulation identification of the conventional Transform model. According to the technical scheme, the method comprises the steps of inputting dual-channel time sequence signals of in-phase components and orthogonal components; extracting preliminary features through convolution operation; the features are input into a Transform architecture integrating an attention mechanism and a hybrid expert mechanism; in the hybrid expert mechanism, calculating a routing weight through a gating network with noise routing and sparsely activating an expert network; merging expert outputs according to routing weights to obtain depth features; and finally, outputting a modulation identification result through the classifier. According to the method, the calculation efficiency is maintained while the model parameter scale is expanded, and the recognition accuracy in a complex channel environment is improved.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Big-model multi-client collaborative positioning knowledge editing method based on federal learning

The invention discloses a large-model multi-client collaborative positioning knowledge editing method based on federated learning. The method comprises the following steps: 1) constructing an LLM model; 2) constructing a FedLEKE task; (3) executing a FedLEKE task, optimizing the hidden state of a Transform layer so as to finely adjust the weight of a feed-forward network in the LLM model, and updating the LLM model; 4) each client generates an intermediate knowledge vector and uploads the intermediate knowledge vector to the server; 5) in a predefined time slot ti belongs to T, the server distributes the stored MKVs to each client; and 6) each client dynamically retrieves related intermediate knowledge vectors in the server based on cosine similarity, re-edits the knowledge vectors, and locates and modifies related parameters in the LLM model according to the re-edited knowledge vectors. According to the invention, more than 96% of performance of non-federated LEKE is reserved, and the performance of the non-federated LEKE is obviously twice better than that of a base line based on FedAvg.
Owner:CHONGQING UNIV

Logging curve reconstruction method, device and equipment, storage medium and product

The embodiment of the invention relates to the technical field of exploration and development, and provides a well logging curve reconstruction method, device and equipment, a storage medium and a product, the method comprises the steps that an original well logging curve in a specified time period is obtained, the original well logging curve comprises a to-be-reconstructed well logging curve and a plurality of basic well logging curves, carrying out representation learning on the original logging curve to obtain an initial feature vector matrix; the initial feature vector matrix is input into a pre-trained curve reconstruction model, a reconstruction curve of the to-be-reconstructed logging curve is obtained, the curve reconstruction model comprises at least one decoding layer, and the decoding layer comprises an asymmetric causal attention layer, a first normalization layer, a feedforward network layer and a second normalization layer; the asymmetric causal attention layer is used for extracting the space-time correlation weight of the initial feature vector matrix. According to the embodiment of the invention, the reconstruction precision of the logging curve can be improved, so that accurate and reliable data support is provided for exploration and development decisions of oil and gas fields.
Owner:RICHFIT INFORMATION TECH +1

Accurate prediction of gas hydrate formation conditions with artificial neural networks (ANN) and multilayer perceptrons (MLPS)

The determination of the probability of gas hydrate formation using artificial neural network (ANN) and multilayer perceptrons (MLPs) models. ANN generally refers to a network of interconnected neurons (also referred to as “nodes”) that model the neurons in a human brain. An MLP refers to a feed-forward network having a specific arrangement of neurons and includes an input layer, one or more hidden layers, and an output layer. Input data such as temperature, pressure, gas mixture composition, and indicators of gas hydrate formation may be obtained and preprocessed for use in training and testing. The ANN and MLPs may be trained using a training set of the input to output a probability of gas hydrate formation. The trained ANN and MLP models may then be used to determine a gas hydrate formation probability for new data associated with a pipeline transporting a gas mixture.
Owner:SAUDI ARABIAN OIL CO

User input information processing method and system for model reasoning stage

The invention discloses a user input information processing method and system for a model reasoning stage, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining a plurality of complete prompts inputted by a user; periodically acquiring hardware resource information and current task information; determining hardware resource allocation information corresponding to the complete prompt; pre-executing parallel computing operation on the complete prompt according to the hardware resource allocation information in the pre-filling stage; according to the hardware resource allocation information, calling the first type of computing equipment with the computing power reaching a preset computing power threshold value in the feed-forward network stage to read the intermediate characteristics from the preset resource pool for decoding and storage operation, wherein the computing power of the first type of computing equipment reaches the preset computing power threshold value; and calling the second type of computing equipment with the memory access performance meeting the preset memory access performance condition in the attention stage to read data from the preset resource pool to perform attention computing storage operation, obtaining a model reasoning result and outputting the model reasoning result, thereby solving the technical problem of hardware resource waste, and achieving the technical effects of improving the resource utilization rate and reasoning efficiency.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Conformer-based mixed self-attention and convolution improved speech recognition method and system

The invention discloses a Conformer-based mixed self-attention and convolution improved speech recognition method and system. The recognition method comprises the steps of obtaining speech feature representation; the speech feature representation is input to a Conformer encoder for encoding processing, high-dimensional feature representation is obtained, and the Conformer encoder comprises a convolution module, an attention linear enhancement module and a feedforward network residual scaling module; according to the high-dimensional feature representation, bidirectional decoding is carried out through a bidirectional Transform decoder, and bidirectional decoding output is obtained; and carrying out merging processing on the bidirectional decoding output to obtain a speech recognition result. According to the method, the modeling capacity for long-time dependence and global context is enhanced through a bidirectional decoder architecture, the processing capacity of the model in a long sequence is improved through ALiBi relative position coding, and the model can be flexibly adjusted in a noise environment through a right decoder weighting adjustment mechanism.
Owner:HEBEI UNIV OF ENG

Optical-SAR (Synthetic Aperture Radar) fusion target detection method under cloud and mist conditions

The invention relates to an optical-SAR (Synthetic Aperture Radar) fusion target detection method under a cloud and mist condition. The method comprises the following steps: constructing an optical-SAR fusion target detection data set under the cloud and mist condition; an optical-SAR fusion target detection model is constructed; the model is composed of a double-flow backbone network based on YOLOv5 expansion and a continuous dual-polymerization Transform module embedded in the middle of the double-flow backbone network. Training the detection model based on the data set; and performing fusion target detection by using the trained detection model. According to the method, information of two dimensions of a channel and a space in the Transform can be fused, feature aggregation of cross-space and channel dimensions is realized through dual modes of a dual aggregation Transform block between modules and in the modules, and deep fusion of features in the modules is performed by means of an adaptive interaction module and a space door feedforward network. The multi-modal feature expression can be integrated more efficiently, and the performance of optical-SAR fusion target detection is improved comprehensively.
Owner:CHINA ACADEMY OF SPACE TECHNOLOGY

Fault tree Boolean function equivalent mapping method based on untrained neural network

The invention discloses a fault tree Boolean function equivalent mapping method based on an untrained neural network, and relates to the field of fault tree analysis. In order to solve the problems that in the prior art, a Boolean function mapping structure is not beneficial to parallel expansion, the calculation efficiency is limited, and the Boolean function mapping structure is difficult to efficiently realize on high-parallel platforms such as a GPU, the invention provides a method for generating topological structure data by analyzing a fault tree model; the basic events, the intermediate events and the top events are mapped into neurons of an input layer, a hidden layer and an output layer respectively, a feedforward network with fixed weight and bias is constructed, and a logic activation function is defined in nodes to realize Boolean logic propagation. The input layer receives a basic event state vector, outputs a top event result through forward propagation, and can realize large-scale Boolean function mapping on a parallel platform through batch input matrixes. The method is suitable for reliability analysis, minimum cut set simplification, top event probability calculation, parallelization fault tree solving and the like of a large-scale complex system.
Owner:HARBIN ENG UNIV

Tibetan multi-dialect speech recognition system and method

The invention provides a Tibetan multi-dialect speech recognition system, and the system comprises a Tibetan self-supervision model which is used for extracting hidden layer speech representation of an original speech signal to obtain a first input feature vector; the original voice signal is Tibetan dialect voice; the dialect feature extraction module is used for extracting dialect features of multiple pieces of dialect information in a one-hot coding form to obtain a second feature vector; the multiple pieces of dialect information at least comprise Anduo, Kangba and defense and Tibetan dialects; the encoder-decoder model is used for routing the corresponding dialect expert module according to the first input feature vector and the second input feature vector and outputting a Tibetan text; and the dialect expert module is used for routing each input feature to the most relevant feedforward network for joint learning. In a multi-dialect environment, unique characteristics of different dialects can be distinguished and understood more accurately, corresponding dialect experts can be effectively and dynamically called according to characteristics of input signals by introducing the hybrid expert module, and processing efficiency and accuracy are remarkably improved.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Efficient simultaneous interpretation method based on expert routing threshold

The invention discloses an efficient simultaneous interpretation method based on an expert routing threshold, and relates to the field of voice processing, the method is based on a classical Transform architecture model, an expert routing strategy model based on the expert routing threshold is constructed, and multi-language streaming translation is realized, the model comprises a streaming voice encoder, and the streaming voice encoder adopts a hybrid design and is connected with the expert routing threshold. The block-by-block autoregression block is composed of an autoregression block and a non-autoregression block; the text decoder is used for simultaneously processing the complete offline voice and the randomly truncated voice prefix to generate a hidden state; the routing threshold module is realized by a feedforward network and projects the final hidden state into a scalar value to determine an expert weight; and the hybrid expert post-processing module shares a language model head with the text decoder, and predicts a target translation sequence in combination with prefix information and global information. According to the method, a mixed expert threshold scheme is adopted to learn the strategy, the self-learning ability of the neural network is fully played, good effects are achieved in streaming translation and streaming TTS, and the method can be used for generating more streaming sequences.
Owner:SHANGHAI JIAOTONG UNIV

Hyperspectral remote sensing image water body extraction method based on CNN-Transform mixed architecture

The invention relates to a hyperspectral remote sensing image water body extraction method based on CNN-Transform mixed architecture, and belongs to the technical field of remote sensing image semantic segmentation. Comprising the following steps: preprocessing a hyperspectral image; constructing a feature extraction network, and obtaining multi-scale spectrum-space features; an enhanced multi-scale context attention module is designed, and local detail features of multi-scale cavity convolution and global context features of cross window global attention are fused; an enhanced residual feedforward network module is designed, and the local feature extraction capability is enhanced by using depth separable convolution; and constructing a dual-path feature fusion module, and fusing deep semantics and shallow detail features through a space attention path and a channel interaction path. According to the method, the problem of missegmentation caused by easy confusion of the water body and the shadow / building in the hyperspectral image and the problem of difficulty in effective fusion of local details and global context information are effectively solved, and the fine water body extraction precision and the boundary definition are remarkably improved.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY

Method and system for detecting corrosion state of overhead ground wire based on deep learning

The invention discloses an overhead ground wire corrosion state detection method and system based on deep learning, and the method comprises the steps: obtaining an overhead ground wire surface image, and constructing a semantic segmentation data set comprising a background and different corrosion grade categories; constructing a deep learning semantic segmentation model based on an encoder-decoder architecture; the encoder comprises a deformable convolution attention module and a multi-scale feed-forward network so as to extract corrosion features; the decoder adopts a full convolution structure, fuses the features extracted by the encoder in each stage, and outputs a semantic segmentation result; training the deep learning semantic segmentation model by using the semantic segmentation data set; and detecting the surface of the on-site overhead ground wire by using the trained deep learning semantic segmentation model, and outputting a corrosion level semantic segmentation result of the surface image of the on-site overhead ground wire. According to the method, a deep learning semantic segmentation model based on an encoder-decoder architecture is combined, and semantic segmentation identification of the corrosion level of the corrosion area of the overhead ground wire is realized.
Owner:ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD

Distributed task processing method, controller, controller cluster and electronic equipment

The invention discloses a distributed task processing method, a controller, a controller cluster and electronic equipment, and relates to the technical field of computers.The method comprises the steps that the controller cluster comprises a plurality of baseboard management controllers, and when a to-be-processed task is received, the baseboard management controllers send the to-be-processed task to the controller cluster; a to-be-processed task is split based on the number of attention heads of a multi-head attention layer of a large language model in a baseboard management controller and a weight matrix of a feedforward network layer, each to-be-processed sub-task is allocated to a corresponding idle controller, the idle controller processes the to-be-processed sub-tasks to obtain processed sub-tasks, and the processed sub-tasks are integrated. And obtaining a final processed task. According to the method, idle controller computing resources in a cluster are utilized, and when one baseboard management controller receives user input and needs to call model computing, a computing task is allocated to other idle controllers according to a division strategy, so that the operation pressure of the baseboard management controller directly interacted by a user is reduced.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Video generation method and system based on three-dimensional sparse attention

The invention discloses a video generation method and system based on three-dimensional sparse attention, and belongs to the field of video generation. The method comprises the following steps of: performing three-dimensional partitioning on an input video feature according to the size of a time dimension block and the number of space dimension blocks, and performing rearrangement index on each three-dimensional sub-block; adopting a block-level Top-K attention mechanism, and only selecting a key block for each query; carrying out attention calculation on query, key and value features of the key block, and carrying out normalization output on a calculation result; utilizing FlashAttention to execute efficient variable-length attention calculation on the rearranged query, key and value characteristics, and recovering an original sequence; video output features are generated through the residual connection and the feed-forward network. According to the method, the long-sequence video attention calculation complexity and video memory occupation are remarkably reduced, the time-space dependence modeling efficiency is improved, and the method is suitable for large-scale video generation and understanding tasks.
Owner:ZHEJIANG UNIV +1

Edge end large language model reasoning acceleration method and accelerator

The invention relates to the technical field of network acceleration, and discloses an edge-end large language model reasoning acceleration method and accelerator, and the method comprises the following steps: reconstructing a calculation process of a decoding stage, and carrying out the deep fusion of a multi-head attention mechanism and the calculation operation of a feedforward network; the weight and key value data are stored in HBM, and the coefficient and the accumulated attention score are stored in DDR; for linear matrix calculation, a unified matrix calculation unit is used for executing multi-precision matrix operation; for nonlinear function calculation, a mathematical transformation and linear fitting method is adopted, a Softmax function is converted into operation with 2 as the bottom through a bottom conversion formula, and truncation and third-order linear fitting are conducted on a Sigmoid function; and constructing a key value screening algorithm based on the accumulated attention score, dynamically adjusting a key value storage position, maintaining a recent key value cache region and an important key value cache region in a limited cache space, and realizing key value efficient cache in long text reasoning.
Owner:CENT SOUTH UNIV