Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

68 results about "Global modeling" patented technology

The term Global Delivery Model is typically associated with companies engaged in IT consulting and services delivery business and using a model of executing a technology project using a team that is distributed globally.

Unified framework for solving automatic driving track prediction and planning consistency based on world model

The invention discloses a unified framework for solving automatic driving track prediction and planning consistency based on a world model. According to the method, through cooperative work of the automatic driving domain controller and the vehicle-mounted sensing system, end-to-end joint optimization of track prediction and planning in a complex traffic scene is realized, time sequence dependence and interaction dynamics among intelligent agents are accurately captured, and the prediction capability and robustness of a model are remarkably improved. The method comprises the following specific steps: firstly, constructing a generative world model, and generating potential future state representation by utilizing a behavior conditional and backtracking expansion technology; secondly, in combination with global modeling and a local convolutional network, multi-scale features are extracted, adaptive fusion is carried out, and a multi-modal prediction trajectory is generated; then, a multi-target planning model is adopted to integrate various driving indexes, and a track with the minimum loss function is generated; finally, path planning parameters are dynamically optimized through real-time environment perception and decision feedback, and the problems of prediction uncertainty and planning consistency of the automatic driving track are effectively solved.
Owner:EAST CHINA UNIV OF SCI & TECH

Short-term power load prediction method and system based on CNN-Transform hybrid model

The invention discloses a short-term power load prediction method and system based on a CNN-Transform hybrid model, and the method comprises the steps: carrying out the data collection of historical load data and meteorological data of a power system, carrying out the data preprocessing of the collected data, and carrying out the coding of a periodic time feature, and obtaining periodic time coding information; local space-time features of the load data and the meteorological data are extracted by using a convolutional neural network, and hierarchical expression of the features is realized through a multi-layer convolutional structure during extraction; the local spatiotemporal features and the periodic time coding information are fused to obtain fusion features containing a load sequence, the long-period dependency relationship of the load sequence is modeled through a Transform network, global modeling of the fusion features is achieved through a multi-head self-attention mechanism, and a CNN-Transform hybrid model is obtained; a CNN-Transform hybrid model is used for prediction, and a load prediction result is output; according to the invention, the precision of load prediction and the generalization ability of the model are significantly improved.
Owner:STATE GRID ELECTRIC POWER RES INST +2

Deep learning method for realizing mechanical fault diagnosis

The invention discloses a deep learning method for realizing mechanical fault diagnosis, and belongs to the technical field of intelligent manufacturing fault prediction and diagnosis. Aiming at the problem that the diagnosis precision is sharply reduced along with the improvement of noise due to insufficient front-end feature extraction and mutual superposition of time domain limitation of a self-attention mechanism, a fault diagnosis classification model composed of three levels of feature processing layers is constructed; each stage comprises a wavelet-guided adaptive multi-scale convolution module and a frequency domain enhanced self-attention module; the wavelet-guided adaptive multi-scale convolution module can extract abundant multi-scale features under high noise; the frequency domain enhanced self-attention module carries out global modeling in the frequency domain, and the influence of noise on the overall recognition precision is reduced. The method has strong multi-scale feature extraction capability and anti-noise interference capability, effectively solves the problem of inaccurate diagnosis precision caused by a high-noise environment under an actual industrial condition, and is suitable for fault diagnosis of rotating mechanical equipment such as bearings and gears.
Owner:TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Multi-scale time series prediction method based on adaptive sparse expert selection strategy and closed continuous time neural network

The invention discloses a multi-scale time sequence prediction method based on an adaptive sparse expert selection strategy and a closed continuous time neural network. The method comprises the following steps: carrying out normalization and low-dimensional feature mapping based on RevIN; performing trend-seasonal structure enhancement processing on the feature sequence after linear mapping; constructing a multi-scale expert model based on the feature sequence after trend-season enhancement; self-adaptive sparse expert selection and load balancing loss calculation are carried out; carrying out weighted aggregation and residual fusion on multi-scale expert output; global modeling of a closed continuous time neural network based on channel weighting is carried out; and finally performing prediction generation and reverse normalization. The multi-scale time series prediction method has the multi-time-scale adaptive modeling capability, the sparse expert efficient selection mechanism and the global continuous time modeling capability, and can be applied to various multivariable time series prediction scenes such as power load prediction, weather prediction, industrial production monitoring, traffic flow prediction and financial price prediction.
Owner:HUNAN UNIV

Handwritten text recognition method based on multi-stage enhancement

The invention discloses a handwritten text recognition method based on multi-stage enhancement. The handwritten text recognition method comprises the following steps: acquiring a handwritten text image; constructing a hierarchical dynamic multi-scale CNN backbone network to obtain a visual feature sequence; inputting the visual feature sequence into a time sequence multi-scale module to obtain a local enhanced feature sequence; performing global modeling on the local enhanced feature sequence by using a Transform encoder to obtain a global visual feature sequence; enhancing the global visual features to obtain a time sequence context feature sequence; dynamic weighted fusion is carried out on the global visual features and the time sequence context features through a gating fusion module; and sending the fused features into a linear classifier and a CTC decoder to obtain an identification result. According to the method, intelligent arbitration of global and time sequence features is realized through new technology application of a hierarchical multi-scale CNN trunk, time sequence context enhancement and a gating fusion mechanism, and the recognition accuracy and robustness are remarkably improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

CNN-Mmba double-branch fusion network for outdoor fire detection

The invention relates to a CNN-Mama double-branch fusion network for outdoor fire detection, and belongs to the technical field of image detection, the CNN-Mama double-branch fusion network comprises a VMama branch and a CNN branch which are processed in four stages, the VMama branch provides a global receptive field and a global modeling capability for outdoor fire, the CNN branch is responsible for capturing local details, and the VMama branch is responsible for providing a global receptive field and a global modeling capability for outdoor fire. And the VMama branch and the CNN branch carry out feature fusion interaction and iteration of down-sampling processing in a parallel mode. The invention provides a novel CNN-Mmba double-branch fusion network, and the network combines the skilled local detail capturing capability of the CNN and the strong long-range context modeling capability of the Mmba, so that the robustness of OF detection is remarkably enhanced in a complex real scene. Different from ViT, Mamba realizes linear calculation complexity while maintaining strong global modeling capability, so that the method is more efficient and suitable for actual scenes.
Owner:CHENGDU ZVAN TECH

Wind power blade image super-resolution reconstruction method and system based on graph neural network

The invention provides a wind power blade image super-resolution reconstruction method and system based on a graph neural network. The method comprises the following steps: S1, dynamically establishing edges between nodes according to the similarity between pixel features; s3, dynamically distributing an aggregation degree for each node based on a detail perception index; s4, performing multi-scale feature aggregation on the graph structure to obtain an aggregated graph structure, the multi-scale feature aggregation including local feature aggregation and global feature aggregation; and S5, performing up-sampling reconstruction on the aggregated graph structure to generate a high-resolution wind power blade image, and optimizing the reconstruction process by using a weighted loss function. The method solves the problems of weak global modeling capability, poor geometric adaptability and uneven calculation resource distribution when a traditional method is used for processing the wind power blade image, improves the reconstruction definition and detail accuracy of the blade defect area, and is suitable for supporting subsequent high-precision automatic detection and state monitoring of the wind power blade.
Owner:SHANGHAI JIAO TONG UNIVERSITY INNER MONGOLIA RESEARCH INSTITUTE

Artificial intelligence-based speech processing method and apparatus, computer device, and medium

PendingCN122658332AEngineeringVoice data
The application belongs to the technical field of artificial intelligence, and relates to a voice processing method based on artificial intelligence, which comprises the following steps: receiving an input voice signal; pre-processing the voice signal to obtain voice data; calling a preset voice processing model; wherein the voice processing model comprises a modeling module, an attention module and a convolutional neural network module; performing feature extraction on the voice data based on the modeling module to obtain feature data; performing global modeling processing on the feature data based on the attention module to obtain global features; processing the global features based on a skip connection mechanism in the convolutional neural network module to generate a target complex spectrum; and performing inverse transformation processing on the target complex spectrum to obtain a target voice signal. The application also provides a voice processing device based on artificial intelligence, a computer device and a storage medium. The application can be applied to the voice enhancement processing scene in the fields of financial technology and digital medical treatment, and effectively improves the generation quality of the target voice signal.
Owner:PING AN TECH (SHENZHEN) CO LTD

Sparse point cloud classification method based on mse-mamba

The method for classifying sparse point cloud based on MSE-Mamba network relates to the technical field of three-dimensional data processing, and solves the technical problem that the existing sparse point cloud classification method is affected by the sparsity of the point cloud, resulting in insufficient capture of local features, low efficiency of global correlation modeling, and difficulty in coordinating local feature extraction and global modeling, thereby causing the classification accuracy to decrease. The method uses a MSE-Mamba multi-scale local feature coding module to complete the extraction and enhancement of the local geometric features of the sparse point cloud, then captures the long-range semantic correlation between the core points through a Transformer module based on a global attention mechanism, realizes the fusion of the local geometric features and the global semantic information, and finally completes the class probability calculation through a multi-feature aggregation strategy to realize the high-precision and high-efficiency classification of the sparse point cloud. The method realizes the coordinated improvement in classification accuracy, calculation efficiency and anti-sparsity robustness.
Owner:XIAN TECH UNIV

Visual representation global modeling method and system based on feature extraction

The invention discloses a visual representation global modeling method and system based on feature extraction, and the method comprises the steps: processing input visual data through employing a visual vocabulary bag feature coding algorithm, generating a feature vector, and obtaining space-time continuous feature representation through the dimension reduction reconstruction of a space-time continuous visual autoencoder; pixel and semantic category mapping is established by using a pixel-level semantic feature mapping algorithm, a pixel-level semantic feature map is generated, and the step comprises the sub-steps of dimension analysis, model construction and the like; inputting the feature map into a high-dimensional visual representation intelligent analysis platform, and screening through sub-steps of segmentation, standardization, correlation analysis and the like to obtain high-dimensional screening features; and on the basis of the screening features, constructing and optimizing a global feature incidence matrix through a visual representation global modeling technology, and outputting a global model. The system comprises corresponding function units, the visual data global association rule is accurately captured, and adaptability and analysis accuracy in a complex scene are enhanced.
Owner:ZHENJIANG ZHIGU HIGH END EQUIP RES INST CO LTD

A small target detection method for improving YOLOX network structure

The present application relates to the technical field of target detection, and particularly relates to a small target detection method for improving YOLOX network structure, by introducing and improving CSPDarkNet network, integrating multi-scale spatial pyramid pooling layer, global self-attention and multi-scale feature fusion modules into the network model, small target features of images can be extracted from complex data sets, and the positioning and effective detection of small targets can be accurately detected. Three technical problems are mainly solved, one is that limited use of maximum pooling convolution makes the top convolution too sparse, resulting in incomplete features extracted; two is that CNN lacks the ability of global modeling and long-distance modeling; three is that single-level extracted features will cause the final prediction result to be far from the true situation.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A micro target detection method based on implicit transformer

The application discloses a kind of micro target detection methods based on implicit Transformer, belong to computer vision and target detection technical field, comprising: S1, input image is extracted high-level semantic feature by backbone network, obtains single scale feature representation and low resolution feature map;S2, single scale feature is enhanced semantic expression by global modeling of encoder;S3, initial query set containing feature representation and position information is generated by initializing query through center point guide dense query initialization module;S4, decoder realizes low resolution feature map continuous position query using implicit deformable attention module;S5, decoder directly outputs target class and position by multilayer optimization, completes end-to-end detection.The micro target detection method based on implicit Transformer provided in the application considers detection precision and efficiency, and is suitable for remote sensing, automatic driving and other scenes.
Owner:SHENZHEN TECH UNIV

Medical image automatic segmentation method based on simplified meta learning

The invention provides a medical image automatic segmentation method based on simplified meta learning, and belongs to the technical field of artificial intelligence and medical image processing. Comprising the following steps: step 1, preprocessing a medical image; 2, constructing a double-branch structure composed of a visual Transform branch and a convolutional network branch, respectively extracting global semantic and local detail features, and realizing information complementation through a weighted fusion module; 3, performing model training under a simplified meta-learning framework; and step 4, inputting a to-be-segmented image into the trained model to generate a segmentation result, and outputting a binary mask through a set threshold. According to the method, the advantages of global modeling and local optimization are fused, and the adaptability, generalization and segmentation precision of the model are remarkably improved through a lightweight meta-learning mechanism.
Owner:NANTONG UNIV

Federal learning distribution external generalization detection method based on local attention enhancement and singular vector global modeling

The invention relates to a federated learning out-of-distribution generalization detection method based on local attention enhancement and singular vector global modeling, and belongs to the technical field of federated learning and out-of-distribution detection. The method comprises the following steps: pre-training a local encoder of a client by using constructed positive and negative sample pairs to obtain initial parameters of the local encoder; a PCAA module is introduced into the local model, and fine tuning is performed on the improved local model based on initial parameters of a local encoder; performing singular value decomposition on the local training data by using a fine tuning model, extracting a first singular vector and adding Gaussian noise to obtain a disturbed first singular vector, and performing weighted aggregation to obtain a global first singular vector; and calculating an included angle between the semantic feature vector of the to-be-detected data and the global first singular vector, and determining whether the to-be-detected data is an out-of-distribution sample based on an included angle value to realize out-of-distribution generalization detection. The technical problem of OOD detection and generalization ability degradation caused by data heterogeneity in a federated learning environment is solved.
Owner:KUNMING UNIV OF SCI & TECH

Remote sensing image change detection method and system based on Fourier frequency domain feature enhancement

The invention discloses a remote sensing image change detection method and system based on Fourier frequency domain feature enhancement. The method comprises the following steps: acquiring double-time-phase images: a first image and a second image; performing shallow convolution flow on the first image and the second image through a shared weight to obtain a first initial feature map and a second initial feature map; respectively inputting the first initial feature map and the second initial feature map into a corresponding multi-cascade encoder to obtain a first deep feature and a second deep feature; inputting the first deep feature and the second deep feature into a self-adaptive dual-time-phase interaction module, and performing cross-time-phase interaction and guide fusion to generate a multi-scale difference feature; and inputting the multi-scale difference characteristics into a decoder to obtain a change probability graph. According to the method, the contradiction between global modeling efficiency and local detail reservation of the existing architecture can be overcome, explicit and efficient fusion of the local spatial features and the global frequency domain features is realized, and the change detection precision and efficiency are improved.
Owner:HAIYANG AEROSPACE IND TECH RES INST +1

Computer network security situation intelligent prediction method based on big data analysis

PendingCN122513160AComprehensive reflection of interaction statusreflect the status comprehensivelyData setFeature set
The application discloses a computer network security situation intelligent prediction method based on big data analysis, and comprises the following steps: obtaining a standardized security data set; performing feature extraction to obtain a multi-dimensional security feature set; constructing a network asset graph and determining a target priority sequence; constructing an attack target set, performing probability modeling, and obtaining an attack intention probability distribution model; generating an attack path set; constructing an attack path evolution sequence, screening and sorting the attack path evolution sequence by using MCDM decision, and obtaining a target attack path set; performing risk assessment processing to obtain a risk value corresponding to each attack path, and performing aggregation processing to generate a network security situation assessment result; generating a network security situation prediction result within a preset time range, and outputting risk change trend information, so that global modeling of the network security situation and early prediction of the future risk change trend are realized.
Owner:SHANGHAI MAITONG INFORMATION TECHNOLOGY CO LTD

Ablation area image segmentation method and system in lung cancer thermal ablation operation, medium and equipment

The invention discloses an ablation area image segmentation method and system in lung cancer thermal ablation, a medium and equipment. The segmentation method comprises the following steps: constructing a lightweight converter based on in-window self-attention and inter-window self-attention cascading; inputting the image into a lightweight converter, and encoding to obtain a multi-scale feature; performing linear mapping on the multi-scale features to obtain mapping features with the same embedding dimension, enabling the spatial dimensions of the mapping features to be the same by using an up-sampling operation, fusing the mapping features by using a channel superposition operation, and generating a coarse segmentation result by using a convolution module; and after superposing the coarse segmentation result and the image channel, inputting the result into an encoder structure of a convolutional network to obtain a convolutional coding feature, fusing the convolutional coding feature with a multi-scale feature, and performing channel superposition with a decoder feature to generate a fine segmentation result. According to the method, complex image features can be extracted, the robustness is high, and the problem that the global modeling capability of the convolutional network is weak is solved.
Owner:QUZHOU PEOPLES HOSPITAL (QUZHOU CENT HOSPITAL)

Atmospheric turbulence image restoration method based on multi-modal condition fusion diffusion transformer

PendingCN122453665AImaging processingAlgorithm
The present application relates to the technical field of image processing, and more particularly to an atmospheric turbulence image restoration method based on a multi-modal condition fusion diffusion transformer. It comprises the following steps: obtaining a clear true value image, generating a turbulence degradation image in combination with an atmospheric turbulence degradation model, constructing a training sample pair and configuring a multi-modal physical prior condition to form a training sample set; building a restoration architecture composed of a multi-modal condition fusion module, a physical perception modulation module and a backbone network, fusing turbulence intensity parameters, semantic labels and diffusion time steps to generate a global condition vector and modulation parameters; injecting the backbone network, completing network iteration training and saving model weights, and combining a non-classifier guide strategy to complete condition encoding fusion during the training process; pre-processing the image to be measured, closing the empty embedding logic and carrying out reverse diffusion iteration, and obtaining the restored image through encoding, feature updating and decoding. The advantage is that the image details are preserved, global modeling and regional adaptive correction are realized, and the restoration effect is improved.
Owner:CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Method and device for improving engineering problem modeling quality and storage medium

The invention discloses a method and equipment for improving engineering problem modeling quality and a storage medium, and belongs to the field of engineering optimization problem modeling by a large language model. The method comprises the following steps: step 1, organizing a series of engineering optimization problems according to hierarchical classification and complexity, and constructing a tree structure knowledge base; step 2, receiving natural language description of a target engineering problem to be modeled, performing recursive search on the tree structure knowledge base in the step 1, and identifying modeling sub-problems corresponding to nodes related to the target engineering problem until the most matched and most specific modeling sub-problem is found; 3, retrieving an advanced modeling thought corresponding to the modeling sub-problem in the step 2, and combining the advanced modeling thought with the description of the target engineering problem to generate a global modeling thought; and 4, automatically generating an optimization model and solver codes of the target engineering problem by using the large language model according to the global modeling thought in the step 3. According to the method, the automatic modeling quality of solving a complex engineering optimization problem by using a large language model can be improved.
Owner:UNIV OF SCI & TECH OF CHINA

An interpretable intelligent analysis method for three-dimensional aircraft aerodynamic parameter prediction

The application discloses an interpretable intelligent analysis method for three-dimensional aircraft aerodynamic parameter prediction, and belongs to the technical field of aerodynamics, deep learning and computer vision, which first converts a three-dimensional model into a two-dimensional image sequence through standardized orthogonal projection, constructs a deep neural network containing five parallel branches to extract geometric features, combines with flow field working condition coding, carries out training by introducing a fine-grained physical constraint step, generates an aerodynamic sensitivity heat map by using an attention mechanism, and verifies the interpretation result by using occlusion sensitivity analysis; the scheme reduces the dimension of a complex three-dimensional problem through standardized orthogonal projection, combines with the global modeling capability of ViT, maintains the second-level prediction speed, and realizes higher prediction accuracy than traditional point cloud networks.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Video target tracking method based on multi-source heterogeneous data collaboration

The invention discloses a video target tracking method based on multi-source heterogeneous data collaboration, and belongs to the technical field of computer vision. The method aims at solving the problems that existing single-mode tracking robustness is insufficient and multi-mode fusion efficiency is low. The method comprises the following core steps: firstly, carrying out adaptive cutting on a red-green-blue (RGB) image and a depth map which are synchronously acquired, and constructing six-channel input based on pseudo-color mapping and channel splicing of a JET color table; secondly, respectively extracting independent feature tokens of RGB and depth modes by adopting a double-branch embedded structure; furthermore, a visual prompt module is embedded in each layer of the visual Transform backbone network, so that hierarchical progressive fusion of two modal features is realized; meanwhile, deploying candidate elimination modules in the third layer, the sixth layer and the ninth layer, and dynamically screening high-confidence-coefficient feature tokens based on the attention weight of a template-search area so as to reduce the calculation overhead; then, a center heat map, a size map and an offset map of the target are predicted in parallel through a center positioning head structure, and an accurate bounding box is generated; and finally, performing end-to-end training in combination with a multi-task loss function, executing a complete process including initialization and tracking in a reasoning stage, and assisting with a post-processing strategy to improve stability. According to the method, the geometric advantages of depth information and the global modeling capability of Transform are fully utilized, and the tracking precision, robustness and reasoning efficiency in a complex scene are remarkably improved.
Owner:BEIHANG UNIV

A chiller fault data generation method, device, equipment, medium and product

ActiveCN121935615BGlobal modelingChiller
This application discloses a method, apparatus, equipment, medium, and product for generating chiller fault data, relating to the field of chiller faults. The method includes acquiring noise data and corresponding category labels for the chiller; using a trained fault data generation module to generate synthetic fault data from the noise data and the corresponding category labels; using the synthetic fault data to train a fault diagnosis module for fault diagnosis; the self-adversarial encoder of the fault data generation module extracts a local feature subspace through its local modeling branch and a global feature subspace through its global modeling branch; and fusing and linearly mapping the local and global feature subspaces to obtain latent variables. This application can improve the stability and quality of fault sample generation.
Owner:HANGZHOU YIQI FUTURE ENERGY TECHNOLOGY CO LTD

CT image motion artifact classification model construction method and system based on feature prototype contrast learning

The invention belongs to the technical field of image recognition, and particularly relates to a CT image motion artifact classification model construction method and system based on feature prototype contrast learning. According to the classification method based on the prototype, the class center is used as feature storage, the inter-class separability and the intra-class consistency in the artifact classification task are enhanced, and the robustness and the classification precision of the model are remarkably improved. The method specifically adopts a Vision Transform (ViT) as a basic model for feature extraction, combines a strong global modeling capability, effectively captures long-range dependence and fine-grained features in artifact detection, and improves the capability of identifying the complexity of artifact types. And by introducing a prototype contrast learning strategy, the feature representation of the artifact image is optimized, and the overfitting problem of the model when training data is insufficient is relieved, so that the generalization ability and robustness of the model in practical application are improved, and an excellent classification result is obtained on a clinical data set.
Owner:CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI +1

A method for image segmentation using a semantic segmentation network

The application discloses a kind of high-efficiency semantic segmentation networks, suitable for medical image analysis, automatic driving etc. The network adopts encoder-decoder architecture, combines CNN with Transformer (converter model), balances global modeling and computational efficiency. Encoder extracts multi-scale features through lightweight convolution, and introduces spatial selection module: its gate convolution splits channel into gate signal and reserved information, and key spatial features are activated and strengthened by Sigmoid activation function; Grouped pooling module extracts details using multi-scale pooling, and restores channels after upsampling and splicing. The decoder fuses multi-scale features and restores resolution through upsampling, enhances efficient channel attention, fuses global max pooling and average pooling, generates channel weight using one-dimensional convolution, and optimizes feature dependence. The design improves small target segmentation accuracy through gate mechanism and multi-scale pooling, is lightweight and easy to expand, and has high precision and practicality.
Owner:TIANJIN PUXIN TECH CO LTD

Nuclear fuel assembly appearance micro-defect detection method based on multi-modal fusion and Transformer

The invention belongs to the technical field of automatic detection and artificial intelligence visual identification, and particularly relates to a nuclear fuel assembly appearance micro-defect detection method based on multi-modal fusion and Transform. Comprising the following steps of 1, data preparation and model training; 2, model reasoning and detection output; the method has the beneficial effects that multi-modal image input is adopted, so that the recognition capability on low-contrast and weak edge defects is improved, and the perception robustness of the model is enhanced; an attention guiding fusion mechanism and a defect guiding attention module are introduced, feature enhancement of a key area is achieved, and background interference is avoided; the trunk network based on Transform has global modeling capability, improves context understanding capability of a defect area, and is suitable for complex texture and multi-scale defect scenes; the multi-loss function collaborative optimization mode can improve the edge segmentation quality while maintaining the pixel-level accuracy, and significantly reduces the false detection and omission ratio.
Owner:CNNC JIANZHONG NUCLEAR FUEL +1

Speech synthesis quality prediction method and device, computer equipment and storage medium

The invention relates to the technical field of voice processing of financial scenes, and discloses a voice synthesis quality prediction method and device, computer equipment and a storage medium, and the method comprises the steps: in the application of voice synthesis quality prediction of a financial scene, through a coherent technical path of multi-scale partitioning-feature extraction-multi-scale fusion-frame-level prediction, the voice synthesis quality of the financial scene is predicted; the problem of local distortion submerging caused by global modeling in the prior art is solved; wherein local modeling is structurally realized through the block processing; and multi-scale fusion gives consideration to distortion characteristics of voice units with different time lengths, so that fine-grained and interpretable analysis is realized. Therefore, the frame-level interpretable prediction of the voice quality is realized under the condition of weak supervision, not only can the overall quality score be given, but also the specific distortion segment causing the quality reduction can be accurately positioned, and a direct and reliable basis is provided for the optimization of a voice system.
Owner:PING AN TECH (SHENZHEN) CO LTD

Power battery internal short circuit early warning method, computer equipment and computer readable medium

The invention discloses a power battery internal short circuit early warning method, computer equipment and a computer readable medium, and the method comprises the steps: obtaining the voltage and temperature of each single cell, the system-level total current and the four-dimensional data of SOC, and constructing a four-dimensional input vector feature space; inputting the four-dimensional input vector feature space into a short-circuit early-warning model of a CNN-Transform hybrid architecture, extracting microscopic abnormal local features of cell-level voltage and / or temperature in each time window through a CNN network, performing global modeling through a Transform model, analyzing a long-time dependency relationship of cell states in a battery pack in the whole time window by using a multi-head attention mechanism, and performing short-circuit early-warning on the cell states of the battery pack in the whole time window according to the long-time dependency relationship. Capturing a lithium precipitation accumulation equal-length time sequence degradation rule to obtain a global dependency feature; and carrying out adaptive weighted fusion processing on the global dependency features and the microscopic abnormal local features to obtain the abnormal probability of each battery cell, thereby realizing short-circuit early warning. According to the invention, the prediction accuracy and timeliness are improved.
Owner:LISHEN (QINGDAO) NEW ENERGY CO LTD

Asymmetric dual-path gating and VLM arbitration method for RGB-infrared target detection

This invention discloses an RGB-infrared target detection method based on asymmetric dual-path gating and VLM arbitration, involving multimodal target detection and cross-modal representation fusion technology in embodied intelligent scenarios. The method involves acquiring and preprocessing multimodal input data from the embodied scenario to generate cross-modal input pairs; constructing an asymmetric topology based on these cross-modal input pairs and establishing a constraint relationship of "semantic mainstream - structural cue flow - semantic arbitration flow"; converting the infrared modalities in the multimodal input data into injectable structural cues; introducing StarFusion star-shaped fusion into the semantic mainstream to achieve efficient global modeling; performing semantic arbitration using a frozen VLM to generate a semantic consistency graph as an interpretable adjudication signal for gating fusion; and performing gating fusion through semantically guided denoising fusion to achieve a unified closed loop for nighttime recovery and thermal noise suppression. This invention has advantages in cross-modal manifold consistency, noise robustness, and multi-scenario generalization ability.
Owner:ZHEJIANG NORMAL UNIV

Lightweight fruit and vegetable identification method based on superpixels

The invention discloses a fruit and vegetable image recognition method based on super-pixel lightweight improvement, and belongs to the technical field of computer vision. According to the method, a focusing token module (SFA) and a channel-space attention mechanism module (CSAM) are constructed, and an adaptive Patch Merging lightweight module is combined, so that the balance between the global modeling capability and the calculation complexity is realized. Wherein a focusing token mechanism (SFA) performs semantic compression on an input image, reduces the number of redundant tokens, improves the collaborative perception ability of a model for local and global features of the image by introducing sparse correlation mapping and low-dimensional attention interaction, and then adopts an adaptive patch merging strategy to reduce spatial dimensions while maintaining high semantic density, so as to improve the robustness of the model. And the information utilization efficiency is improved. An introduced channel-space attention mechanism module (CSAM) emphasizes an area where an object is located and suppresses background interference in a spatial dimension; and in the channel dimension, the attention on the high-distinction-degree features is improved. The method is small in parameter quantity and low in calculation overhead, can be conveniently deployed on resource-limited platforms such as edge equipment, can be widely applied to scenes such as automatic fruit and vegetable sorting, agricultural robots, food processing and the like needing efficient image recognition, and has important practical significance and economic value.
Owner:LUDONG UNIVERSITY

Multi-modal video target tracking method and system based on comparative learning modal alignment

The invention discloses a multi-modal video target tracking method and system based on comparative learning modal alignment, and belongs to the field of target tracking. The method comprises the following steps: acquiring an RGB image and a depth image of an input video sequence, and performing synchronous registration and preprocessing; dividing the preprocessed image into a plurality of patches, respectively inputting the patches into a feature extraction network, and extracting semantic texture features of RGB and geometric structure features of depth; taking the extracted features as input, constructing a cross-modal comparison learning modal alignment module, performing positive and negative sample comparison on the RGB features and the depth features, realizing modal alignment, and obtaining cross-modal consistent fusion feature representation; and inputting the fusion features into a Transform backbone network, performing global modeling by using a multi-head self-attention mechanism, and outputting a position prediction and tracking result of the target in the video sequence. According to the method, the tracking robustness can be remarkably enhanced in complex scenes such as illumination variation, target shielding and similar backgrounds, and high-precision video target tracking under a multi-mode condition is realized.
Owner:LANZHOU CITY UNIV