Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

43 results about "Attentional network" patented technology

The Attentional Network theory proposes three independent cognitive concepts: physiological state, and prepares the organism for fast reactions. Orienting involves selective allocation of attention to a source of signals in space.

Multi-modal data drawing logical relationship analysis method, electronic equipment and medium

The invention discloses a multi-modal data drawing logical relationship analysis method, electronic equipment and a medium, and the method comprises the steps: generating a node set based on drawing image data and text data; generating a cross-modal hyperedge set based on the spatial proximity relationship, the visual feature similarity and the semantic correlation between the node sets; generating a hypergraph embedding input representation based on the node set and the cross-modal hyperedge set; the hypergraph is embedded into the input representation input improved hypergraph self-attention network model, and a hyperedge logic relation type and a corresponding hyperedge confidence coefficient are generated; generating a graph structure result based on the hyperedge logic relationship type and the node set, wherein the graph structure result meets the structure legality requirement; and performing hyper-parameter automatic adjustment and convergence control on the atlas structure result based on hyper-edge confidence, and generating an optimal atlas analysis model and a structured output result. According to the method, the reliability and the quality of analysis of component nodes, logic edge relationships and semantic structures in the drawing are improved.
Owner:NANJING ELECTRIC POWER ENG DESIGN +1

RIS-assisted MIMO implicit channel estimation method based on graph attention network

The invention discloses a reconfigurable intelligent surface (RIS)-assisted multiple-input-multiple-output (MIMO) implicit channel estimation method based on a graph attention network, which is used for efficient downlink transmission in a multi-user scene. Firstly, a graph attention network is designed, user nodes and RIS nodes are modeled in a unified mode, received pilot signals serve as initial features, spatial feature expression is enhanced in combination with user three-dimensional position information, and therefore interference between users and a spatial correlation structure are accurately represented; secondly, end-to-end feature aggregation is achieved based on a message passing mechanism, a base station beam forming matrix and an RIS reflection coefficient are directly predicted under the condition that explicit channel estimation is not needed, and the total transmitting power constraint and the unit mode constraint are met through normalization processing so as to complete joint optimization; according to the method, the users and the speed of the system can be remarkably improved under limited pilot frequency overhead, and the method has excellent generalization performance and robustness in a multi-user complex propagation environment.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Multi-modal sentiment analysis method based on graph-attention collaborative optimization cross-modal recombination

PendingCN121350723ABiological modelsSequence learningData mining
The invention provides a multi-modal sentiment analysis method based on graph-attention collaborative optimization cross-modal recombination, and relates to the technical field of multi-modal sentiment analysis. The method comprises the following steps: firstly, designing a modal self-adaptive multi-modal graph construction module, constructing a local sparse graph based on KNN-RBF for a language modal, and adopting a low-rank representation method combined with nuclear norm regularization for an audio and video modal; secondly, the processed modal features are transmitted into a graph attention network to realize high-order feature aggregation; then, a language-guided hierarchical cross-modal interaction mechanism is constructed, and multi-granularity semantics are accumulated in combination with an advanced multi-modal feature container module; and finally, designing an advanced feature recombination strategy based on dynamic matching, and realizing feature alignment by taking a language feature container as an anchor point. According to the method, graph learning and sequence learning are unified in a collaborative framework through a graph-attention collaborative optimization cross-modal recombination model, so that the problems of cross-modal attention noise interference, modal imbalance, insufficient cross-modal feature alignment efficiency and the like can be effectively solved.
Owner:ZHONGYUAN ENGINEERING COLLEGE

Multi-modal image fusion method based on deep coding and decoding axis interactive attention network

The invention relates to the technical field of image processing, in particular to a multi-modal image fusion method based on a deep coding and decoding axis interactive attention network, which comprises the following steps of: S1, acquiring data of an infrared image and a visible light image, normalizing the infrared image and the visible light image, and then inputting a model for feature extraction; the feature extraction method comprises an encoder, a fusion strategy and a decoder. S2, in a feature extraction stage of an encoder, multiple times of extraction is performed on two paths of infrared and visible light images through a convolutional neural network and transformer, so that local information of two modal images is captured, and encoding representation is obtained; s3, performing multi-time layered fusion on the features coded by the encoder by a fusion strategy; s4, reconstructing and decoding the fused image by a decoder; the feature representation capability is higher, the utilization rate of original information is higher, image detail mining is more sufficient, and the fusion effect is better.
Owner:SHAANXI SILK ROAD DIGITAL INTELLIGENT NAVIGATION TECHNOLOGY CO LTD

Three-dimensional tooth model segmentation method of double-branch geometric attention network based on centroid guidance

A three-dimensional tooth model segmentation method of a double-branch geometric attention network based on centroid guidance is oriented to oral cavity three-dimensional scanning point cloud data and comprises the steps that firstly, normalization and normal vector estimation are conducted on the oral cavity three-dimensional point cloud data, a double-branch encoder for coordinate and normal decoupling is constructed, and a global topological structure and local geometric features are extracted respectively; secondly, a separable attention mechanism guided by the mass center is introduced into a coordinate branch so as to improve the distinguishing ability of adjacent teeth, and a graph convolution attention mechanism is introduced into a normal branch so as to strengthen the boundary expression of the teeth and gingiva; further, in the fusion stage, a mass center thermodynamic diagram and multi-scale feature aggregation are combined, and joint modeling of global and local structures is achieved; and finally, collaborative optimization of instance segmentation and centroid localization tasks is carried out through a dynamically weighted joint loss function. According to the method, the segmentation precision and robustness under complex cases are remarkably improved while the light weight is kept, and the automation and clinical practicability of oral diagnosis and treatment are enhanced.
Owner:ZHEJIANG UNIV OF TECH

A trajectory generation and simulation method for sparse data completion-oriented attention mechanism

The application discloses a kind of attention mechanism trajectory generation and simulation methods for sparse data completion, belong to intelligent transportation and trajectory prediction field.The application fills in trajectory missing value using high-precision sensor and map information, combines graph attention network (GAT) and multi-modal fusion technology;Adopt the attention module based on distance (D-GAT) and based on view (V-GAT), capture the interaction between vehicles, improve the understanding of complex traffic scene;Through prediction supervision generator and multi-modal trajectory generator, combine LSTM and Gaussian mixture model (GMM) to generate multiple possible trajectories, and use Kalman filter for online adjustment, ensure the accuracy and real-time of trajectory.The application realizes the intelligent completion of sparse traffic data and the accurate generation of trajectory, provides reliable data support and decision basis for intelligent transportation system, helps the efficient operation and sustainable development of urban traffic planning and management.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Power topological graph node classification method, system and device based on federal asynchronous graph attention network and storage medium

The invention discloses an electric power topological graph node classification method, system and device based on a federal asynchronous graph attention network, and a storage medium, and relates to the field of electric power system automation, and the method comprises the steps: collecting local electric power topological graph data of each regional power grid client, constructing a data set, and generating weighted embedding through label semantic embedding learning; based on this, a weighted tag semantic graph and an auto-encoder are constructed to obtain coding representation and loss, the two are combined to obtain a classification result through backbone network processing, and client model training is carried out. And finally, generating a global model at a server side through operations such as spectrum similarity and the like, wherein the global model is used for power topological graph node classification. According to the invention, the privacy of power grid data is guaranteed, and the generalization ability and robustness of the model are improved. The node type and state of the power topological graph are accurately identified, a decision basis is provided for troubleshooting and load distribution, the operation and maintenance efficiency of a power grid is effectively improved, and safe and stable operation of the power grid is guaranteed.
Owner:HAINAN POWER GRID CO LTD

A method for semantic segmentation of oceanic internal wave ripples in SAR imagery

This invention discloses a semantic segmentation method for ocean internal wave stripes in SAR images, relating to the field of semantic segmentation of remote sensing images. The method model consists of an encoder and a decoder. The encoder comprises four Transformer modules, each containing a self-attention layer, a feedforward neural network, and an overlap patch merging module. Within each module, the input image is processed N times through a multi-head self-attention mechanism, and then the merging module generates feature maps at four scales. The decoder consists of three modules: a serpentine convolution, an EVC module, and an expectation-maximization attention network. The advantages of this invention are: the model fully utilizes the multi-scale fusion module, improving performance and robustness; the use of serpentine convolution can better extract features of linear shapes; and the use of the expectation-maximization attention network improves model accuracy while reducing computational complexity.
Owner:HOHAI UNIV

Methods, apparatus, and devices for child reading and attention deficit risk screening

PendingCN122320544Aefficient extractionEfficient characterizationFunctional connectivityNetwork connection
This application relates to a method, apparatus, and device for screening the risk of reading and attention deficit disorder in children. The method includes acquiring multi-channel raw brain blood oxygenation signals under task-induced conditions using a specific layout fNIRS array integrated into a wearable headband, based on a rapid naming cognitive paradigm. Based on the raw brain blood oxygenation signals, a fusion feature vector representing the reading and attention networks is generated by calculating temporal waveform features and frontotemporal functional connectivity strength. The multi-dimensional fusion feature vector is then processed and analyzed using a Transformer classification model to generate classification results indicating the risk level of reading disorders and comorbid ADHD. This application achieves portable and rapid brain function signal acquisition by integrating a targeted fNIRS array with a standardized cognitive paradigm. By fusing temporal dynamics and brain network connectivity features, a multi-dimensional neural representation is constructed. Finally, a lightweight Transformer model is used to output the risk level of reading disorders and comorbid ADHD end-to-end, achieving high-precision automated assisted screening.
Owner:INSTITUTE OF MENTAL HEALTH OF PEKING UNIVERSITY (SIXTH HOSPITAL OF PEKING UNIVERSITY)

A visual question answering method and system based on a multi-level visual feature enhancement network

The application provides a visual question answering method and system based on a multi-level visual feature enhancement network, which can enhance the relationship between local objects and local objects and the relationship between regional objects and global concepts, thereby jointly learning the visual semantic relationship of multiple spatial contexts. A separation visual feature module based on a graph attention network is used to capture pixel-level visual features and object-level regional features; a joint visual feature representation based on a graph attention network is used to jointly represent the pixel-level features and the object-level features, simultaneously learn the semantic relationship between different levels, better associate with the question text, and thus provide more rich visual feature representation. The application solves the problem that the traditional visual feature representation loses the context relationship between the regional features and the global features, so that the global semantic cannot be fully utilized, and the visual features are lost.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Texture feature guided texture preserving low dose ct image denoising

The application discloses a texture feature guided texture preserving low dose CT image denoising method, and belongs to the field of medical image processing.The application specifically discloses a multi-scale deep residual attention network model with texture feature guidance, which is applied to low dose CT imaging.The main network model comprises four sub-models, one is a multi-scale initial denoising network model for denoising low dose CT, and the other network is used for extracting texture details after the initial denoising network, and the two network parts work cooperatively.The extracted texture details and the initial low dose CT are fused through a multi-scale image and texture feature fusion network model, and then enter a multi-scale main denoising network for further denoising of the low dose CT, which is beneficial to the main denoising network to learn more unobvious details.The low dose CT image denoising method disclosed by the application efficiently removes the noise and stripe artifacts in the low dose CT image, and meanwhile, the structural information and texture feature detail information in the image are preserved.
Owner:QUFU NORMAL UNIV

Spatial domain identification method based on single cell large model and graph attention network

PendingCN121768479ABiostatisticsBiological modelsCellular modelData pre-processing
The invention belongs to the technical field of bioinformatics, and more specifically relates to a spatial domain identification method based on a single cell large model and a graph attention network. The method comprises the following steps: S1, data preprocessing; s2, extracting hidden features by adopting an scGPT full-human model; and S3, screening out high-variation genes with top 3000 ranks from the preprocessed data, splicing the high-variation genes with hidden features, and finally inputting the spliced high-variation genes into a graph attention network to obtain a clustering result. According to the method, the missing part of ST data is acquired by using scGPT, the problem that high-quality single cell data is difficult to acquire is solved, and meanwhile, with the development of a single cell large model, GFGAT is also very convenient to use different single cell models. A finer graph attention network is used, and a part of the network is trained by adopting a cut adjacent matrix, so that more site information is reserved, and a more accurate spatial domain is identified.
Owner:NANKAI UNIV

Underground pipe network facility service life prediction system and method based on multi-parameter time sequence analysis

The invention relates to the technical field of digital operation and maintenance, in particular to an underground pipe network facility life prediction system and method based on multi-parameter time sequence analysis, and the method comprises the steps: generating a holographic state set through a multi-scale window and a tensor filling algorithm, building a pipe network topology through three-layer logic verification, and paying attention to network quantification risk conduction characteristics based on a space-time diagram. And performing state recursion and continuous integration by using a particle filter algorithm to predict the remaining life. According to the method, heterogeneous monitoring data is converted into a low-rank tensor model, and sparse graph matrix operation is combined, so that the calculation complexity and memory overhead in a high-dimensional spatial-temporal feature extraction process are effectively reduced; meanwhile, based on a numerical integration strategy of a Bayesian sequence Monte Carlo method, the particle degradation problem in nonlinear system state estimation is solved, and convergence and numerical stability of a residual life prediction result in a computer simulation environment are guaranteed.
Owner:NANJING TOWNGAS CO LTD

Alzheimer's disease classification method and system based on multi-modal hypergraph attention network

The application provides an Alzheimer's disease classification method and system based on a multi-modal supergraph attention network, and the method comprises the following steps: acquiring sMRI image data of the brain of a plurality of Alzheimer's disease patients and performing preprocessing; performing feature extraction on the preprocessed sMRI image data, and constructing a plurality of cross-modal supergraphs according to the image features and morphological features of the brain regions of the patients; establishing a supergraph attention neural network model, training the cross-modal supergraphs, finally acquiring sMRI image data of the brain of a patient to be diagnosed, constructing a corresponding supergraph, inputting the trained supergraph attention neural network model for classification, and obtaining an Alzheimer's disease classification result and the attention weight corresponding to each supergraph; the application can effectively improve the accuracy of the Alzheimer's disease classification task, and can also find out which brain regions and morphological supergraphs have a significant contribution degree in the model, which is helpful for accurate diagnosis by doctors.
Owner:GUANGDONG UNIV OF TECH

Building damage change detection method and device based on change guidance and interactive attention

The invention discloses a building damage change detection method and device based on change guidance and interactive attention. The method comprises the following steps: acquiring a double-time-phase image of a building to be detected; inputting the double-time-phase image into a trained interactive attention network based on change guidance to obtain a change detection graph of the to-be-detected building, the change detection graph being used for indicating a damage change area of the to-be-detected building; wherein the interactive attention network is used for extracting multi-scale dual-time-phase features of the dual-time-phase image, generating a prior change diagram based on the deepest-scale dual-time-phase features, and under the guidance of the prior change diagram, executing interactive attention operation on the multi-scale dual-time-phase features to obtain a change detection diagram. The method is good in detection effect and high in detection precision in a complex scene.
Owner:XIDIAN UNIV

Entity relationship identification method and device, computer equipment and medium

The invention relates to an entity relationship recognition method and device, computer equipment and a medium, the method applies an entity relationship extraction model to recognize entity relationship data for a target text, and the method comprises the following steps: a feature representation network performs feature representation on a word segmentation sequence of the target text of a to-be-recognized entity relationship to generate a full-text feature vector; a conversion network performs entity classification based on the full-text feature vectors to obtain entity boundary information, and word element vector segments of all entities in the full-text feature vectors are constructed into entity feature vectors; performing convolution enhancement processing on the entity feature vector by a convolution network to obtain a structure enhancement vector; the attention network performs context enhancement processing on the structure enhancement vector by using the full-text feature vector to generate an entity enhancement vector; and determining entity relationship data between every two entities contained in the target text by the classification network according to the entity enhancement vectors. According to the method, the entity feature representation precision and breadth can be improved, and the accuracy and efficiency of relation extraction are remarkably improved.
Owner:CHENGDU HARIT MEDICAL TECH CO LTD

Drug recommendation methods and related equipment based on drug representation and user dynamic modeling

This application relates to the field of healthcare informatics technology, providing a drug recommendation method and related equipment based on drug representation and dynamic user modeling. User features are generated based on the acquired user's historical health records and current health status. Diagnostic features and procedural features are sequentially input into a GRU network and a Transformer network, respectively, to generate user representations through dynamic modeling. Drug features are input into a pre-constructed graph attention network to construct a heterogeneous graph between drug attributes and molecular motifs. Drug representations are generated by message propagation and stacking on this heterogeneous graph. User and drug representations are input into a pre-constructed feedforward neural network, outputting fused features between the drug and the user. The fused features are used to generate probabilities through an activation function, and recommendation information is generated based on the target drugs corresponding to these probabilities. This method can accurately match user health needs and provide personalized and effective drug recommendations.
Owner:XIAN HOSPITAL OF TRADITIONAL CHINESE MEDICINE +1

Spoken-to-written conversion method, device and equipment based on graph attention network

This invention provides a method, apparatus, and device for spoken-to-written language conversion based on graph attention networks. The method includes: semantically encoding a spoken document to obtain a semantic representation of the spoken document; determining the initial representation of each node in the document structure graph of the spoken document based on the semantic representation, wherein the document structure graph includes document nodes, sentence nodes, and word segmentation nodes; performing message propagation on the initial representation of each node in the document structure graph based on an attention mechanism to obtain a structure graph representation of the document structure graph; and performing semantic decoding based on the structure graph representation to obtain the written document corresponding to the spoken document. The method, apparatus, and device provided by this invention, by constructing a document graph structure diagram, can obtain a more concise and readable written document, avoiding the omission of spoken terms crossing sentence boundaries during text conversion, and ensuring the effectiveness of document-level spoken text conversion to written text.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

A session recommendation method, device and equipment and computer storage medium

The application provides a conversation recommendation method, device and equipment and a computer storage medium, comprising: obtaining an item set and a conversation set; obtaining a first feature matrix of the item set, and performing hypergraph convolution processing on the first feature matrix of the item set by using a hypergraph neural network of the item set to obtain a second feature matrix of the item set; inputting the second feature matrix of the item set into a multi-layer self-attention network for learning to obtain a first feature matrix of the conversation set; constructing a graph attention network by using the conversation set, and inputting the first feature matrix of the conversation set into the graph attention network for learning to obtain a second feature matrix of the conversation set; and calculating a recommendation score of each item by using the second feature matrix of the item set and the second feature matrix of the conversation set, and calculating a loss function by using the recommendation score of each item and a recommendation true value, which improves the recommendation accuracy through the combination of the hypergraph neural network, the multi-layer self-attention network and the graph attention network.
Owner:SOUTH CHINA NORMAL UNIV

A knowledge graph-based service recommendation method

The application provides a service recommendation method based on a knowledge graph, comprising the following steps: S1, converting the interactive matrix data of a user into a two-part graph, and then matching the non-user entities in the two-part graph with the entities in a knowledge graph to form a joint graph in combination with the knowledge graph; S2, using a knowledge graph embedding method to parameterize the entities and relationship parameters of the joint graph into vector representations; S3, inputting the representations of the entities into a multi-layer graph attention network, using an attention mechanism to calculate the neighbor entity weight of each entity respectively, and performing weighting; S4, aggregating the representation of the node and the weighted result obtained in step 3; S5, repeating steps 3-4, so that each entity recursively aggregates its neighbor entities to obtain the final representation of the user and the entity; S6, predicting the probability of the user's service preference according to the final representation of the user and the entity. The method introduces auxiliary information of the knowledge graph, and improves the recommendation effect of the recommendation system.
Owner:TONGJI UNIV

Power load inflection point prediction method based on multi-persistence graph attention network

The invention discloses a power load inflection point prediction method based on a multi-persistence graph attention network, and the method comprises the steps: carrying out the collection and preprocessing of the historical active load and corresponding meteorological data of a power grid region, and adaptively determining the number of key load points through a perception PIP algorithm in combination with a fitness function; extracting a key load point set representing the morphological change of the load curve; on this basis, a key load point visible graph is constructed, node degrees and node eccentricity rates are used as two-parameter filtering functions, a series of filtering sub-graphs are generated according to multiple sets of threshold values, Euler-Poincare features are calculated, and a multi-parameter Euler-Poincare feature matrix is formed; carrying out normalization, vectorization and row-based first expansion on the feature vector to obtain a one-dimensional multi-persistence topological feature vector of a fixed dimension; and then, introducing a graph attention network formed by stacking multiple layers of graph attention layers to output inflection point category probabilities of prediction days step by step, and realizing refined prediction of the power load inflection points.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A crane anti-collision early warning method based on a space-time diagram attention network and a transformer

PendingCN122286713AImprove collision risk perception capabilitiesRealize early warningFeature vectorWorking environment
This invention provides a crane collision avoidance early warning method based on a spatiotemporal graph attention network and a Transformer architecture, comprising the following steps: S1, real-time collection of raw monitoring data within the work area using sensing sensors to construct a set of historical trajectories of the target; S2, construction of a dynamic heterogeneous spatiotemporal topology graph and feature initialization of nodes; S3, modeling the spatial interaction relationships between different targets using a graph attention network, extracting spatial interaction features and assigning corresponding risk weights; S4, inputting the spatial enhanced feature vectors into a Transformer-based network in chronological order, performing temporal encoding, and outputting predicted trajectories for multiple future time steps; S5, combining the predicted trajectories with the crane's planned motion path to assess collision risk, and outputting corresponding early warning signals or control commands when the risk exceeds a preset threshold. This method achieves early perception and dynamic prevention of potential collision risks during crane operation, improving operational safety and intelligence in complex working environments.
Owner:YICHANG WTAU ELECTRONICS EQUIP

Multi-scale three-attention network construction method and device for pixel-level crack segmentation

The embodiment of the specification provides a multi-scale three-attention network construction method and device for pixel-level crack segmentation, wherein the method comprises the following steps: adopting a ResNet18 network as a basic network to construct a backbone network; embedding a multi-scale input strategy in the backbone network, detecting feature information through the multi-scale input strategy to obtain a multi-scale feature map; fusing multi-scale feature maps of different resolutions to obtain a fused multi-scale feature map; using feature learning for guiding bottom feature mapping through an AAF block based on the fused multi-scale feature map to obtain an enhanced multi-scale feature map; based on the enhanced multi-scale feature map, detecting attention features through a TA block to obtain an aggregated attention feature map, and guiding training of an end-to-end pixel crack detection network, i.e. a multi-scale three-attention network, based on a crack defect data set according to the aggregated attention feature map through a loss function of the MSTA-Net.
Owner:GUANGZHOU UNIVERSITY

Image enhancement network method based on compressed self-attention Transform and standardized flow

The invention provides an image enhancement network method based on a compressed self-attention Transform and a standardized flow. The image enhancement network method comprises the following steps: step 1, data preprocessing; step 2, constructing a zero element joint height and channel compression self-attention Transform-U network; 3, triple condition features are led out from the zero-element joint height and channel compression self-attention Transform-U network constructed in the step 2, standard normal distribution is sampled, the standard normal distribution and the triple condition features are input into a standardized flow network together, an attention network is used for learning affine / linear transformation parameters of a condition feature driving layer, and the zero-element joint height and channel compression self-attention Transform-U network is constructed; and reversible transformation is guided to fit standard normal distribution, and the image enhancement network based on the compressed self-attention Transform and the standardized flow is generated by the standardized flow network. According to the image enhancement network method based on the compressed self-attention Transform and the standardized flow provided by the invention, the designed low-illumination image enhancement network can effectively realize restoration of missing information of the image, and the produced enhanced image has an excellent visual effect and more excellent peak signal-to-noise ratio and structural similarity.
Owner:NANJING UNIV OF POSTS & TELECOMM

Coarse-to-fine attention network for optical signal detection and recognition

A vehicle light signal detection and recognition method, system, and computer program product includes defining one or more regions of an image of a car using a coarse attention module to generate one or more defined regions, the image including at least one of a brake light and a signal light generated by the car, the one or more regions including an illuminated portion, removing noise from the one or more defined regions using a fine attention module to generate one or more noise-free defined regions, and identifying the at least one of the brake light and the signal light from the one or more noise-free defined regions.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Training method, audiovisual segmentation method, electronic device and storage medium

The invention relates to a training method, an audio-visual segmentation method, an electronic device and a storage medium, and the method comprises the steps: obtaining a training sample which comprises an audio signal, an image, a target semantic tag and a segmentation tag; extracting audio features and image features, and performing self-attention enhancement on the audio features based on the target semantic tag to apply target semantic consistency constraint to obtain enhanced audio features; performing cross attention reweighting on the unenhanced audio features based on the guidance of the enhanced audio features to obtain target pointing audio features; bidirectional interactive fusion is executed based on the image features and the target pointing audio features, and a sparse self-attention network is adopted in the image feature fusion process to suppress unmatched image region response; generating a segmentation prediction result based on the fused image features and audio features, and training to obtain an audiovisual segmentation model; and inputting a to-be-processed video into the trained audio-visual segmentation model, and outputting a segmentation result. The image segmentation accuracy of the sounding target can be improved.
Owner:SUZHOU UNION INTELLIGENT TECH CO LTD

Single-channel electroencephalogram driving fatigue recognition method based on time-frequency attention network

The application discloses a single-channel electroencephalogram driving fatigue recognition method based on a time-frequency attention network. It belongs to the field of electroencephalogram signal analysis, and the operation steps are: driving experiment paradigm design and single-channel electroencephalogram signal acquisition; the collected electroencephalogram signals are pretreated; the pretreated electroencephalogram data are subjected to continuous wavelet transformation to obtain the spectrum-time representation corresponding to each data; the spectrum-time representation of each sample is input into a time-frequency attention network model, so that the model automatically extracts valuable feature information and completes the recognition of the fatigue state. The spectrum-time representation of the single-channel electroencephalogram signal is combined, a time-frequency attention mechanism and an adaptive feature fusion module are used to fully mine and capture key features related to driving fatigue, and the recognition of the driving fatigue state is realized; the method is reasonable in design, convenient to realize, good in detection effect and high in practical value.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Deformable face spoofing detection network and spatiotemporal consistent face spoofing model construction method

This invention relates to the field of face authentication technology, and discloses a deformable face authentication network and a spatiotemporally consistent face authentication model construction method. The deformable face authentication network includes a backbone network and a deformable temporal self-attention network. The backbone network is used to extract the spatial features of the face in the input image, and the deformable temporal self-attention network is used to process the spatial features of the face output by the backbone network to extract the temporal features of the face. The spatiotemporally consistent face authentication model construction method, based on the deformable face authentication network and the temporal self-attention feature mechanism, can alleviate the interference of semantic feature offset between face video frames on the extraction of temporal forgery features, thereby effectively improving the generalization ability of the authentication model. By constructing a spatiotemporally consistent face authentication model, accurate separation and representation of the spatiotemporal features of real faces and forged faces are achieved, thereby effectively improving the robustness of the authentication model against unknown forged faces.
Owner:HAOHAN DATA +1

Speech authentication method and system based on residual attention network

The present disclosure provides a voice authentication method and system based on a residual attention network, which comprises: obtaining audio data to be detected and performing corresponding preprocessing; performing feature extraction on the preprocessed audio data, and performing frame processing on the extracted voice feature data to obtain voice signal feature data with a fixed frame length; based on the voice signal feature data, using a pre-trained residual attention network model to obtain enhanced feature data; wherein the residual attention network model comprises sequentially connected convolution modules, multi-scale residual modules, shrinkage excitation units, attention pooling modules and a fully connected layer; and inputting the enhanced feature data into a pre-trained classifier to obtain a voice authentication result.
Owner:SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1

A sar image ship instance segmentation method based on global semantic boundary attention network

ActiveCN115272842Bprecise positioningOvercoming the problem of limited positioning capabilitiesCharacter and pattern recognitionNeural learning methodsPattern recognitionEngineering
The application discloses a SAR ship instance segmentation method based on a global semantic boundary attention network, and aims at solving the problem of limited target frame positioning capability in the prior art. The application is based on the deep learning theory and mainly comprises a global context information modeling module and a boundary attention prediction module. The global context information modeling module establishes a long-distance dependency relationship by enhancing the semantic information of features for multiple times, thereby effectively reducing background interference. The boundary attention prediction module improves the positioning capability of the target frame by predicting the boundary information of the target twice. The method provided by the application is superior to other SAR ship instance segmentation methods based on the deep learning in average precision (AP). The application can overcome the problem of limited target frame positioning capability in the prior art, and improve the instance segmentation precision of ships in SAR images.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA