Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

267 results about "Feature aggregation" patented technology

Face spoofing detection method based on query-driven forensic adapter

This invention presents a face forgery detection method based on a query-driven forensic adapter, belonging to the fields of artificial intelligence and machine learning. It aims to address issues such as insufficient modeling of local forgery regions, limited interaction between semantic and visual features, and poor cross-distribution generalization performance. First, this invention proposes a dynamic semantic query module. This semantic query is injected as a dynamic attention signal into the multi-head attention mechanism of the CLIP encoder, effectively guiding CLIP to focus on potential forgery regions, improving its response to local anomalies such as boundary discontinuities and texture inconsistencies, and compensating for its insufficient local modeling capabilities without altering the CLIP's core structure. Second, this invention proposes a multi-scale feature aggregation mechanism. It proposes a semantically guided attention fusion module and a feature enhancement strategy based on linear interpolation. This method can effectively improve sample diversity and enhance the model's robustness and generalization ability across forgery methods and datasets.
Owner:BEIJING UNIV OF TECH

Neurofibromatosis prediction method based on clinical medical information

The present application relates to the technical field of medical information processing and intelligent disease prediction, in particular to a neurofibromatosis prediction method based on clinical medical information. The method comprises the following steps: acquiring multi-modal clinical medical information with time labels; performing anatomical system classification processing, extracting clinical feature nodes and calculating correlation, and establishing an initial multi-system prediction network graph; determining the asynchronous graph node feature aggregation rate according to the feature update rate difference of adjacent nodes under different time labels, performing local feature diffusion processing, and determining a global multi-system prediction network graph; further determining the cross-domain evolution incubation period weight, the phenotype cascade transfer probability matrix and the disease topology state information entropy, and determining the neurofibromatosis prediction result accordingly. The present application solves the contradiction of long-term and ineffective consumption of a large amount of computing resources to cope with low-frequency but high-impact clinical events, and optimizes the trade-off relationship between risk identification accuracy and system response timeliness.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

Pipe network node defect visual detection method and system based on a detection robot

This application relates to the field of pipeline network detection technology, and discloses a visual detection method and system for pipeline network node defects based on a detection robot. The method includes: acquiring node images of a sewer network using a detection robot, and performing circle detection and sector segmentation on the node images to obtain multiple sector region blocks; extracting the block feature vector of each sector region block; calculating angular attenuation weights and radial attenuation weights, and constructing a node connection relationship matrix based on the angular attenuation weights and radial attenuation weights; combining the block feature vectors and the node connection relationship matrix to perform feature aggregation and illumination compensation to obtain a global feature vector; calculating angular deviation index, radial mutation rate index, and leakage index, and generating defect visual detection results, thereby realizing comprehensive diagnosis and spatial positioning of pipeline network node connection defects, and improving the defect detection accuracy of sewer detection robots in complex node environments.
Owner:GUANGDONG XINTUO NETWORK TECHNOLOGY CO LTD

A rotating target detection method based on frequency domain hypergraph fusion and geometric optimization

This invention discloses a rotating target detection method based on frequency domain hypergraph fusion and geometric optimization, mainly comprising three core components: a dynamic frequency domain hypergraph feature aggregation mechanism, a multi-scale reparameterization module, and a vertex geometric consistency optimization strategy. By combining the frequency domain relationship modeling mechanism, the multi-scale structure reparameterization module, and the geometric consistency constraint strategy, a rotating target detection model containing frequency domain structure information and geometric consistency constraints is formed, suitable for remote sensing image analysis, aerial target detection, and other visual application scenarios involving rotating target localization.
Owner:NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER

An object grasping method based on a denoising diffusion model

The application discloses an object grabbing method based on a denoising diffusion model, and relates to the technical field of robots, and comprises the following steps: S100, scene point cloud data processing; S200, forward process noise adding network training; S300, point cloud data completion; S400, object grabbing simulation; S500, feature aggregation; S600, aggregated feature noise adding; S700, feature heat map grabbing; and S800, final grabbing posture matrix calculation. The application realizes generation of a six-degree-of-freedom diversity grabbing posture for parallel clamping jaws, and improves the success rate, accuracy and diversity of grabbing.
Owner:NINGBO ARTIFICIAL INTELLIGENCE RES INST OF SHANGHAI JIAOTONG UNIV

3D Target Detection Method and Apparatus for Autonomous Driving

PendingCN122090417AImprove recallSuppress oversamplingBiological modelsScene recognitionPoint cloudAlgorithm
This invention provides a method and apparatus for 3D target detection in autonomous driving, relating to the field of autonomous driving technology. The method includes: acquiring point cloud data collected by LiDAR; for each point in the point cloud data, defining a neighborhood of a set radius centered on the current point, counting the number of points within the neighborhood, and normalizing the number of points within the neighborhood to obtain a normalized value; determining the density feature of the current point based on the negative of the normalized value; extracting semantic features of each point in the point cloud data through multiple SDSA layers, predicting the foreground confidence of each point based on the semantic features, and performing point sampling by combining the density features and the foreground confidence to generate sampling points; aggregating the features of each sampling point to generate aggregated features for each sampling point, and predicting 3D target bounding boxes based on the sampling points and their aggregated features. This improves the sampling bias problem caused by the unevenness of the point cloud data.
Owner:UNIV OF SCI & TECH BEIJING

A tobacco professional knowledge document processing and analysis method based on multi-level feature aggregation

The application discloses a tobacco professional knowledge document processing and analysis method based on multi-level feature aggregation, and relates to the technical field of natural language processing. The method can adapt to tobacco professional knowledge documents with different degrees of structuring through intelligent document classification and differentiated preprocessing; can capture semantic and knowledge features of different granularities of the documents in layers through three-level feature progressive extraction and cross-level feature fusion; can realize accurate blocking in line with the knowledge semantic boundary based on the fused multi-level features through semantic segmentation matching the corresponding processing pipeline, so as to guarantee the semantic independence and integrity of each document block; can realize the uniformity of professional expression and the enrichment of information dimension of the document block content through term standardization, entity link processing and context enhanced information supplement; and can convert the standardized knowledge block into vector data which can be efficiently searched and accurately called through vectorization processing and writing of the vector library of the supporting hierarchical feature metadata.
Owner:WUHAN DEFA INFORMATION TECHNOLOGY CO LTD

A method and system for diagnosing faults of a high-frequency transformer

ActiveCN122174127BData setTimestamp
The application relates to the technical field of fault diagnosis, and provides a high-frequency transformer fault diagnosis method and system, which comprises the following steps: collecting a target signal with a time stamp and extracting corresponding features, simultaneously relying on a transformer structure, material parameters and physical rules to build a digital twin model, simulating insulation and structure degradation equivalent working conditions, solving multi-physical field data and generating multi-physical field mechanism samples; then, the mechanism samples and field measured data are fused through a generative adversarial network to expand the fault sample data set and solve the sample scarcity problem; in the running stage, the digital twin model is updated in real time through parameter online inversion, a feature dynamic graph representing multi-physical coupling is constructed by combining the parameter deviation of internal mechanism degradation and the multi-source features of external working conditions, finally, the feature aggregation and time sequence reasoning are completed through the graph neural network and the time sequence neural network trained offline, the fault probability is output, and the optimal diagnosis result is determined; thereby, the fault recognition accuracy and the robustness in the running stage are improved.
Owner:SOUTHWEST JIAOTONG UNIV

Lithium ore microscopic image segmentation method and system based on improved Unet model

This invention provides a method and system for lithium ore micro-image segmentation based on an improved Unet model, belonging to the field of machine vision and image segmentation. The method constructs a lithium ore micro-image dataset and acquires micro-images of lithium ore samples to be tested, performing image preprocessing; pixel-level annotation of the micro-image dataset is performed to generate corresponding mineral category mask images, which are then segmented and data augmented; a micro-image segmentation model is constructed and improved based on the Unet algorithm, constructing a cascaded structure of depthwise convolution and pointwise convolution in the downsampling path, embedding an anti-adhesion dynamic serpentine convolution module in the skip connections, introducing a channel-space dual-path adaptive attention mechanism module, and constructing a multi-scale feature aggregation module in the upsampling path; a mature model is obtained after training; the input micro-image to be tested yields a preliminary segmented image, which is then subjected to adaptive area filtering. This invention improves the segmentation accuracy and precision of lithium ore micro-images.
Owner:YICHUN JIANGLI LITHIUM BATTERY NEW ENERGY IND RES INST +1

A source-load joint probability prediction method and system of a physically constrained graph attention network

The application discloses a source-load joint probability prediction method and system of a physically constrained graph attention network. The method collects multi-dimensional feature data of source-load nodes in a prediction area to construct an initial node feature matrix. A similarity matrix is generated through differentiable graph structure learning. A dynamic adjacency matrix is generated through normalization and introduction of a sparse mask. Spatial feature aggregation is performed through a multi-head graph attention network to obtain node spatial encoding. The node spatial-temporal hidden state is output through an encoder. The node spatial-temporal hidden state is input into a probability prediction head to output Gaussian distribution parameters of the source-load node power. A joint loss function is constructed. The joint loss function is used for soft constraint training to output a probability prediction result. Posterior projection hard constraint correction is performed in an inference stage to obtain a corrected prediction result. The application solves the problems of lack of physical consistency and inability to quantify uncertainty in the prior art.
Owner:BEIJING NORTH STAR DIGITAL REMOTE SENSING TECH CO LTD +1

A transformer partial discharge identification method and device

PendingCN122449298AGraph spectraTransformer
The application relates to the technical field of transformer partial discharge identification, and discloses a transformer partial discharge identification method and device. The method comprises the following steps: acquiring multi-terminal high-frequency pulse current signals through a multi-terminal physical sensing array arranged on a transformer electric circuit; extracting time-domain features, frequency-domain features and polarity features from the multi-terminal high-frequency pulse current signals, and constructing a time-frequency-polarity three-dimensional tensor atlas; constructing a topological graph model of the transformer based on the physical structure of the transformer and a measured injection test; inputting the time-frequency-polarity feature tensor atlas and the topological graph model into a graph neural network, performing feature aggregation on the time-frequency-polarity feature tensor atlas based on the topological graph model; performing global graph feature pooling and probability classification on the aggregated features output by the graph neural network, and obtaining an identification result of transformer partial discharge. The application can solve the problem of high misjudgment rate of transformer partial discharge detection in a complex electromagnetic environment.
Owner:STATE GRID BEIJING ELECTRIC POWER CO

A traffic flow prediction method based on a transformer

PendingCN122369257AData graphEngineering
This invention discloses a traffic flow prediction method and system based on Transformer, belonging to the fields of intelligent transportation and deep learning technology. Addressing the technical problems of existing traffic flow prediction models, such as difficulty in simultaneously considering long-term and short-term dependencies, inability of static road network topology to characterize dynamic spatial heterogeneity, and poor modeling performance of spatiotemporal feature coupling, this invention proposes a multi-timescale adaptive graph attention Transformer model. This method first reconstructs the original traffic data at low, medium, and high time scales, and then aggregates spatiotemporal features through a temporal convolutional network and a compressed excitation network. Next, an adaptive data graph generation module learns node embedding vectors to generate an adaptive adjacency matrix that integrates static topology and dynamic associations. Finally, an encoder incorporating temporal one-dimensional convolutional multi-head attention and spatial graph attention, and a decoder integrating causal convolution and temporally gated convolution, are constructed to achieve high-precision multi-step prediction of traffic flow. This invention effectively captures the spatiotemporal dependencies of traffic flow, with prediction accuracy and generalization superior to mainstream models, and can be widely applied to urban intelligent traffic management, dynamic path planning, and traffic congestion mitigation scenarios.
Owner:CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY +1

Heterogeneous unmanned cluster game decision method and device

The application provides a heterogeneous unmanned cluster game decision method and device, belonging to the technical field of data processing, comprising: obtaining local observation information of each first-camp agent in a heterogeneous unmanned cluster in a multi-agent interaction scene, and a historical observation sequence of each agent in the multi-agent interaction scene; based on the historical observation sequence, performing behavior intention prediction on each agent through an intention prediction network to obtain an intention prediction vector of each agent; splicing the local observation information and the intention prediction vector to construct a node feature of each agent, and performing hierarchical attention feature aggregation through a hierarchical attention network to obtain a decision feature vector; wherein the hierarchical attention feature aggregation comprises same-camp feature aggregation, cross-camp feature aggregation and global feature aggregation performed in sequence; based on the decision feature vector, generating an action strategy of each first-camp agent through a strategy network to control each first-camp agent to perform a corresponding action.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Sparse point cloud classification method based on mse-mamba

The method for classifying sparse point cloud based on MSE-Mamba network relates to the technical field of three-dimensional data processing, and solves the technical problem that the existing sparse point cloud classification method is affected by the sparsity of the point cloud, resulting in insufficient capture of local features, low efficiency of global correlation modeling, and difficulty in coordinating local feature extraction and global modeling, thereby causing the classification accuracy to decrease. The method uses a MSE-Mamba multi-scale local feature coding module to complete the extraction and enhancement of the local geometric features of the sparse point cloud, then captures the long-range semantic correlation between the core points through a Transformer module based on a global attention mechanism, realizes the fusion of the local geometric features and the global semantic information, and finally completes the class probability calculation through a multi-feature aggregation strategy to realize the high-precision and high-efficiency classification of the sparse point cloud. The method realizes the coordinated improvement in classification accuracy, calculation efficiency and anti-sparsity robustness.
Owner:XIAN TECH UNIV

A two-stage highway lane line detection method and related device

This application provides a two-stage highway lane detection method and related equipment, belonging to the fields of computer vision and intelligent transportation technology. The method includes: a first stage, using a lane detection network to perform preliminary detection on the input image to obtain an initial set of lane lines; and a second stage, constructing a lane line topology refinement module. Using the initial lane lines as graph nodes, edge connections are constructed based on geometric attributes and visual features to form a lane line topology graph. Iterative message passing and feature aggregation are performed through a graph neural network to refine the initial detection results, enhancing the confidence and positioning accuracy of detected lane lines, and recovering missed lane lines based on topological relationship inference. This application can effectively improve the accuracy, completeness, and real-time performance of lane line detection in highway monitoring scenarios, and is suitable for lane-level fine-grained perception tasks in intelligent transportation systems.
Owner:SOUTH CHINA UNIV OF TECH

Local feature matching system based on keypoint detection

Disclosed in the present invention is a local feature matching system based on keypoint detection. The system comprises: an encoder, a deep Transformer model and a matching module. An image pair to be matched (IA, IB) is input into the encoder, and fine features (formula I), coarse features (formula II), keypoints PA and PB of the image pair (IA, IB) are extracted; the coarse features (formula II) and the image pair (IA, IB) are input into the deep Transformer model, and the deep Transformer model performs feature aggregation on an image IA and an image IB to obtain keypoint features (formula III); and the matching module converts the keypoint features (formula III) into a confidence matrix C, performs matching between the keypoint PA and the keypoint PB on the basis of the confidence matrix C, and then performs matching enhancement on the basis of the fine features (formula I), so as to complete the matching of the image pair (IA, IB). A weight (parameter) reuse technique is used to share task parameters between consecutive Transformer layers on the basis of task requirements, such that a model can maintain a feature expression capability, improving the model performance, and can also effectively reduce the model size. In addition, the use of a multi-scale keypoint detector reduces the propagation of redundant information and enhances feature specificity, thereby improving the model efficiency.
Owner:YANGTZE RIVER DELTA HIT ROBOT TECH RES INST

An abnormal driving behavior detection method based on an improved residual double-layer graph attention network

PendingCN122454542AEncoder decoderSimulation
The application discloses an abnormal driving behavior detection method based on an improved residual double-layer graph attention network (Res-DBiGATv2), and belongs to the technical field of Internet of Vehicles. Firstly, the vehicle trajectory data is constructed into a space-time dynamic graph sequence according to time steps. Then, a ResBiGATv2 module is designed, and spatial feature aggregation is realized through double-layer graph attention, a multi-head mechanism and residual connection. Then, a ContraNorm contrast normalization layer is used to enhance feature uniformity and inhibit dimension collapse and oversmoothing. A graph external attention enhancement module GEA is introduced, and a learnable external memory unit is used to inject global information. Finally, a GRU is used to model time sequence features, and an abnormal node is identified through reconstruction error in an encoder-decoder framework. The application can accurately and timely detect various abnormal driving behaviors such as slow driving, overspeeding, following, and stagnation in complex traffic scenes, and significantly improves detection accuracy and robustness.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method and device for assessing refractive development of neural circuit based on visual development

This invention provides a method and device for refractive development assessment based on visual development prediction neural circuits, relating to the field of myopia management technology. The method includes: acquiring color retinal optical imaging signals and refractive ground state quantification values ​​of the subject; performing retinal stabilization preprocessing; extracting multi-level morphological features through a reentrant layered retinal choroid encoder; performing spatial feature aggregation gating on the extracted feature maps to obtain feature vectors; concatenating and integrating the feature vectors with the refractive ground state quantification values ​​to obtain fused feature vectors; constructing a fused feature vector sequence and inputting it into a refractive development dynamics internal model inferrer to generate a future multi-time-point refractive state prediction sequence; converting it into an interpretable refractive state numerical sequence; and providing risk warnings based on preset risk thresholds. This invention explicitly embeds physiological mechanisms prior into the model structure and parameter adaptive strategies, thereby maintaining the stability and interpretability of predictions even in a single medical visit scenario.
Owner:BEIJING TONGREN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Medical image segmentation method based on lightweight wavelet enhancement fusion

PendingCN122368082AArnold transformationData set
The application relates to a light wavelet enhancement fusion medical image segmentation method, and belongs to the technical field of medical image processing and computer-aided diagnosis. The core of the method is to construct a light WEF-Net network, which comprises an encoder, a bottleneck layer and a decoder. The encoder adopts a dual-domain perception module, extracts complementary features from the frequency domain and the time domain through wavelet transformation and convolution, and simultaneously extracts complementary features from the frequency domain and the time domain. The bottleneck layer is designed with a feature aggregation Kolmogorov-Arnold transformation module, which is used for efficiently fusing multi-scale semantic information and enhancing the nonlinear modeling capability. The network level also integrates a sawtooth rolling feature fusion module, which enhances the global continuity of the features through the channel rolling and spatial scanning mechanism to improve the integrity of the boundary segmentation. Experiments show that the method realizes excellent performance on multiple public medical image datasets, while maintaining low parameter quantity and low computational quantity, and significantly improves the segmentation precision.
Owner:FUJIAN PROVINCIAL HOSPITAL

Feedforward event camera three-dimensional reconstruction method and system based on spatiotemporal feature aggregation

The application discloses a kind of feedforward event camera three-dimensional reconstruction method and system based on space-time feature aggregation, comprising: obtaining at least two asynchronous event streams, convert each event stream into space-time voxel tensor, space-time voxel tensor includes multiple time boxes, each time box corresponds to the event accumulation in a short time fixed time slice;Space-time voxel tensor is input into time attention encoder, the feature of each spatial position is aggregated in time sequence on different time boxes by self-attention mechanism, obtain space-time feature map rich in time context information;Space-time feature map is input into the space encoder-decoder based on feedforward architecture, and the globally aligned three-dimensional point graph is generated by regression.The method and system can effectively extract the motion information and geometric clues in event data, can accurately predict three-dimensional point graph, and are significantly better than existing event camera methods in depth estimation, camera pose estimation and three-dimensional reconstruction tasks.
Owner:ZHEJIANG UNIV +1

A multi-target connected unmanned relay control method and device based on reinforcement learning

PendingCN122119740APower managementNetwork topologiesSimulationBelief state
The application provides a multi-target connected unmanned relay control method and device based on reinforcement learning, which is aimed at an emergency communication scene containing a command center, a relay unmanned aerial vehicle and multiple mobile task nodes, utilizes an attention mechanism to aggregate features of variable-length multi-node observation data, processes a historical state sequence by introducing a memory-enhanced deep reinforcement learning model, extracts a belief state and outputs a flight acceleration instruction of the relay unmanned aerial vehicle, simultaneously constructs a multi-target reward function containing communication quality and flight energy consumption, dynamically adjusts weights by using an adaptive balance factor based on real-time signal strength, and performs model training in combination with signal offset exploration and a priority sequence experience playback mechanism. The application solves the technical problems of poor generalization ability, easy falling into local optimization and difficulty in dynamically balancing each optimization target of existing methods when facing dynamic changes in the number of task nodes, partial observability of the environment and multi-target conflicts of communication and energy consumption.
Owner:BEIJING UNIV OF POSTS & TELECOMM

An image super-resolution method based on omnidirectional spatial feature learning, a terminal and a storage medium

The application discloses an image super-resolution method based on omnidirectional space feature learning, a terminal and a storage medium, and relates to the technical field of image processing. The method comprises the following steps: using an encoder network to extract an initial shallow feature map of an input low-resolution image; inputting the initial shallow feature map into a deep feature extraction network composed of a plurality of cascaded omnidirectional feature extraction modules to perform spatial relationship perception feature extraction, channel relationship perception feature extraction and multi-scale detail perception feature extraction on the initial shallow feature map, and obtaining an omnidirectional feature map; inputting the omnidirectional feature map into a high-frequency texture enhancement module to perform frequency modulation feature enhancement and sparse non-local feature extraction on the omnidirectional feature map, and generating a texture enhancement feature map; performing context-aware feature aggregation on the texture enhancement feature map to obtain an aggregated feature vector; and performing implicit decoding and image reconstruction on the aggregated feature vector to obtain a high-resolution image. The application can comprehensively capture image features, effectively restore high-frequency textures and intelligently aggregate features.
Owner:SHENZHEN MSU-BIT UNIVERSITY

A resource scheduling method for smart property based on deep learning

PendingCN122311717AData streamService experience
This application relates to the field of smart property management and discloses a resource scheduling method for smart property management based on deep learning. The method includes: constructing a spatial topology influence map of the target property management area; spatially aligning the multi-source heterogeneous data streams according to the nodes of the spatial topology influence map and generating a sequence of node feature vectors corresponding to each node; inputting the node feature vector sequence into a spatiotemporal influence propagation prediction model; performing spatial feature aggregation on the node feature vector sequence based on the graph attention layer and the static topology constraint parameters in the spatiotemporal influence propagation prediction model to generate a spatiotemporal influence representation characterizing the degree of impact of sudden events on each functional area; predicting the resource demand heat value of each functional area within a preset time period based on the spatiotemporal influence representation to form a resource demand heat distribution; and generating a resource scheduling instruction set based on the resource demand heat distribution. This technical solution improves the user's property service experience.
Owner:HUBEI LIANTOU CITY OPERATION CO LTD

A driving intention recognition method based on functional connection and graph neural network

ActiveCN117076980BSensorsDiagnostic recording/measuringFunctional connectivityElectroencephalogram feature
The application provides a driving intention recognition method based on functional connectivity and a graph neural network, comprising the following steps: S1, collecting electroencephalogram signals of a driver during driving, and preprocessing original electroencephalogram data; S2, calculating power spectral densities of each frequency band of the preprocessed electroencephalogram signals as electroencephalogram signal frequency domain features; S3, constructing an adjacency matrix as an initial graph structure based on electrode spatial proximity and functional connectivity; and S4, inputting the frequency domain features in S2 and the adjacency matrix obtained in S3 into a graph attention network for feature aggregation, inputting the extracted feature expression into a classifier to classify driving intentions and outputting results. The method can solve the problems of single electroencephalogram feature extraction and poor expression ability of the classification model in the existing driving intention prediction method, improve the accuracy of driving intention classification, improve the interpretability of the classification results, and better apply the method to driving state perception and auxiliary decision-making of a human-machine co-driving system.
Owner:BEIJING JIAOTONG UNIV

Sparse anomaly behavior detection method and system based on privacy protection semantic evidence

This invention relates to the field of cyberspace security and threat intelligence analysis technology, and discloses a sparse anomaly behavior detection method and system based on privacy-preserving semantic evidence. First, irreversible semantic abstraction is performed on collected multi-source security data to extract semantic evidence elements related to anomalies, which are then semantically encoded to obtain semantic evidence units. Next, based on the semantic evidence elements and their semantic feature representations, a semantic evidence hypergraph containing intra-event and cross-event hyperedges is constructed. Finally, based on the semantic evidence hypergraph, feature aggregation is performed on semantic evidence nodes and their relationships, anomaly scores for each behavioral event or semantic evidence are calculated, and sparsity is determined by combining the frequency of behavior occurrence to identify anomalies. This invention can construct cross-event semantic behavior representations without exposing the original data and model multi-entity collaborative relationships based on the hypergraph structure, achieving effective detection of low-frequency, cross-stage recurring anomalies.
Owner:HAINAN NORMAL UNIV

A speech deepfake attribution method and system under few-shot data

PendingCN122347960AReference sampleMedicine
The application provides a voice deep forgery attribution method and system under few sample data, and relates to the technical field of voice forgery attribution. The method comprises: obtaining a to-be-tested audio and a reference sample set; performing feature processing on the to-be-tested audio and the reference audio to obtain an embedding vector of the to-be-tested audio and an embedding vector of the reference audio; performing feature aggregation on the embedding vector of the reference audio to obtain a multi-prototype representation of the reference audio; and generating an attribution result of the to-be-tested audio according to the embedding vector of the to-be-tested audio and the multi-prototype representation of all reference audios, wherein the attribution result is a forgery category of the to-be-tested audio. The application overcomes the defect that a traditional single-prototype representation is difficult to cover a complex distribution, can effectively adapt to a few sample data scene, improves the attribution robustness in an unknown forgery scene, and gets rid of excessive dependence on a known category statistical hypothesis.
Owner:HEFEI UNIV OF TECH +1

An article classification method and device based on graph attention diffusion

The application provides an article classification method and device based on graph attention diffusion, the article classification method comprising: obtaining article information of an article to be classified, and constructing graph structure data of the article to be classified; determining an attention coefficient between two adjacent article nodes in the graph structure data based on a latent representation vector of each article node and using a graph attention mechanism; performing iterative attention diffusion on an attention matrix based on an adjacency matrix of the graph structure data and a jump back probability coefficient, to obtain a deep-level attention matrix; and performing feature aggregation on the latent representation vector based on the deep-level attention matrix, to obtain an attention aggregation feature of each article node, so as to determine a classification result corresponding to each article to be classified. Through the above method, the information from a high-order neighborhood is quantitatively aggregated, and an over-smoothing phenomenon is avoided, thereby improving the accuracy and stability of processing an article classification task.
Owner:CHINA ELECTRONICS CORP 6TH RES INST

Elliptic multiscale latent space method for intelligent fast prediction of turbulent flows

PendingCN122263739AImprove learning effectstrong interactionDesign optimisation/simulationNeural learning methodsState spaceComputational physics
The application discloses an elliptical multi-scale hidden state space method for intelligent rapid prediction of turbulent flow, and belongs to the cross field of machine learning and computational fluid dynamics. The method first encodes the velocity field, external force field, boundary information and Reynolds number in a multi-scale manner to construct a feature pyramid; then elliptical local attention and four types of physical bias are introduced, and physical consistent feature aggregation is realized through anti-symmetric weight construction; adjacent scale low-frequency bidirectional exchange is used to depict positive and inverse order series of turbulent flow energy; long time series evolution is completed with adaptive time step relying on physical perception hidden state space and system identification; the flow field is output through coarse-to-fine decoding, and autoregressive training is carried out by using joint loss. The application can accurately capture the characteristics of turbulent flow rotation, multi-scale coupling and long-time dynamics, significantly improve the prediction accuracy and stability of high Reynolds number instantaneous turbulent flow, suppress error accumulation, and is suitable for rapid intelligent solution of complex turbulent flow scenes, and has strong generalization ability and engineering practical value.
Owner:HARBIN INST OF TECH

Foreign object detection method for power transmission line based on fusion of edge perception and multi-scale feature fusion

This invention discloses a method for detecting foreign objects on power transmission lines that integrates edge perception and multi-scale feature fusion. The method includes the following steps: scaling and normalizing images to obtain batch tensor data; the backbone network performs spatial compression and channel expansion of features through hierarchical convolutional downsampling, extracts multi-directional edge gradients in parallel using multi-directional Sobel convolutional kernels, filters effective edge information through directional attention mechanisms, and achieves dynamic balance between edge features and semantic features through dual-branch gating fusion to generate a multi-scale basic feature map; the neck network transforms the multi-scale basic feature map into a multi-scale fused feature map through dynamic alignment, background noise suppression, and multi-scale feature aggregation; and the head network obtains the detection result by performing target detection and redundant box filtering on the multi-scale fused feature map. This invention improves the detection accuracy and environmental adaptability for various types of foreign objects on power transmission lines while maintaining real-time performance.
Owner:NORTH CHINA ELECTRIC POWER UNIV

A method and apparatus for atmospheric visibility estimation based on the fusion of binocular stereo vision and deep learning

PendingCN122090261AHigh-precision non-contact surface telemetryEfficient captureCharacter and pattern recognitionBiological modelsBinocular stereoFeature fusion
This invention discloses a visibility estimation method and apparatus based on binocular vision. The method includes: simultaneously acquiring left and right views using a calibrated binocular camera; inputting the image pairs into a stereo matching deep neural network to obtain a disparity map and converting it into a depth map; inputting the left view and depth map into a dual-branch deep convolutional neural network to extract multi-scale features; fusing RGB features and depth features through a cross-modal feature fusion module; and finally outputting a visibility estimate through a feature aggregation and regression module. The apparatus includes a binocular image acquisition unit, a data processing and visibility estimation unit, and a result output unit. This invention solves the depth ambiguity problem of monocular vision by fusing binocular depth information and image appearance information, achieving high-precision, non-contact visibility surface measurement. It has the advantages of low cost and flexible deployment, and is suitable for visibility monitoring in traffic scenarios such as highways.
Owner:NANJING MEIJISEN INFORMATION TECH CO LTD