Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

158286results about "Character and pattern recognition" patented technology

System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance

A system and method are disclosed for generating hyperspectral images from multi-modal sensor data including RGB, LiDAR, thermal, and near-infrared inputs. Training data includes hyperspectral images and corresponding multi-modal measurements. Spectral band grouping is performed based on correlation coefficients. A multi-modal decomposition network with cross-modal attention mechanisms generate reconstructed hyperspectral images by fusing complementary sensor information. A fine-tuning network creates reconstructed RGB images. A comprehensive quality assurance system analyzes spectral consistency, cross-modal coherence, and fusion artifacts to generate quality metrics. Missing data compensation strategies handle corrupted sensor inputs using information from other modalities. The system includes temporal integration for video sequences and multi-resolution processing for different sensor resolutions. Quality metrics guide network weight adjustments to improve reconstruction accuracy while maintaining robustness to sensor failures and environmental variations.
Owner:ATOMBEAM TECH INC

Salient contour matching-based method for target measurement in severe imaging environment

Disclosed in the present invention is a salient contour matching-based method for target measurement in a severe imaging environment. The method specifically comprises: (1) acquiring a binocular image of a target; (2) establishing a global-local joint constraint-based background light estimation model, and removing a scattering effect of a medium in an imaging environment to obtain a restored left eye image and a restored right eye image; (3) learning an original image, and on the basis of a residual between a network reconstructed image and the original image, obtaining target localization prediction maps of the left eye image and the right eye image; and (4) respectively extracting contour lines of the target in the left eye image and the right eye image, constructing feature matching descriptors of contour points, performing stereo matching on the two sets of contour lines by minimizing matching cost, and performing three-dimensional reconstruction on the contour lines in light of calibrated intrinsic and extrinsic parameters to complete the measurement of a key size. According to the present invention, the key sizes of different targets in a severe environment can be accurately measured, thereby providing an effective solution for the problem of measuring the sizes of targets in a severe environment.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LTD YANCHENG POWER SUPPLY BRANCH

Deep learning-based facial recognition system with privacy-preserving features

The present invention provides a facial recognition system using deep learning methodologies while integrating privacy-preserving capabilities. This system employs convolutional neural networks (CNNs) to extract and classify facial features, ensuring high accuracy in recognition tasks. Moreover, the system addresses privacy concerns by incorporating techniques such as facial feature encryption and anonymization, thereby enhancing user privacy and data security. This invention is applicable across various domains, including security, surveillance, access control, and personalized services, where facial recognition is utilized while preserving individual privacy.
Owner:TRIPATHI BHASKAR +11

Power plant intelligent maintenance method and system based on multi-modal dynamic graph learning

The invention discloses a power plant intelligent maintenance method and system based on multi-modal dynamic graph learning. The method comprises the following steps: acquiring structured sensor data, unstructured data and equipment physical topology data of equipment operation in real time through a multi-source sensor cluster and an industrial terminal; the method comprises the following steps: preprocessing multi-modal data, and fusing multi-modal features by using a double-flow Transform architecture and a gated attention mechanism to generate a joint embedded representation; constructing a dynamic causal graph based on equipment physical topology data and sensor time sequence characteristics, updating an edge weight through a GraphSAGE algorithm, fusing domain rule constraints, and outputting equipment state information; generating a maintenance strategy through an improved near-end strategy optimization algorithm according to the state and the equipment health index; and finally, the maintenance strategy triggers third-level early warning of the DCS through an OPC UA protocol, and a maintenance instruction is accurately issued. According to the method, the defects of a traditional method in the aspects of data fusion, fault modeling and decision making are overcome, and the safety, the economical efficiency and the operation and maintenance intelligent level of power plant equipment are remarkably improved.
Owner:SEVENTH SENSE IOT (SHANGHAI) CO LTD

Transform-based cross-modal fusion multi-modal emotion recognition method

The invention discloses a Transform-based cross-modal fusion multi-modal emotion recognition method and device, which are used for solving the problems of modal isomerism, difficulty in time alignment and insufficient dynamic emotion modeling in a multi-modal emotion recognition task, and the method takes the accuracy and robustness of emotion recognition as performance evaluation indexes. Firstly, feature information of three modes of vision, voice and text is obtained, feature extraction is performed on each mode through a deep learning model, then features of different modes are fused by using a cross-mode Transform module, and a complex dependency relationship between the modes is dynamically modeled through a multi-head self-attention mechanism, so that more accurate emotion recognition is realized, and the emotion recognition efficiency is improved. And finally, performing emotion prediction on the fused features based on time sequence modeling and an emotion classification module. According to the method, the problems of modal isomerism, difficulty in time alignment and insufficient dynamic emotion modeling in multi-modal emotion recognition can be effectively solved.
Owner:SOUTHEAST UNIV

Defect detection method for high-voltage equipment based on deep learning and multispectral image fusion

The invention relates to a high-voltage equipment defect detection method based on deep learning and multispectral image fusion, and relates to the technical field of electric power high-voltage equipment state detection. The method comprises the following steps: acquiring an ultraviolet image, an infrared image and a visible light image of the surface of the high-voltage equipment; carrying out image pixel feature-based fusion processing on the ultraviolet image, the infrared image and the visible light image through an image fusion method; establishing a high-voltage equipment defect detection model, and training the high-voltage equipment defect detection model by using the fused image data to obtain a high-voltage equipment defect identification model based on the YOLO-STrans multispectral fusion network; and inputting the ultraviolet image, the infrared image and the visible light image of the outer surface of the power high-voltage equipment into a high-voltage equipment defect identification model to obtain a fault identification result of the to-be-detected power high-voltage equipment. The method can improve the recognition precision of the extremely early insulation degradation and temperature anomaly defects of the surface of the high-voltage power equipment.
Owner:ANHUI NANRUI JIYUAN POWER GRID TECH CO LTD

Explanatory model architecture for image scoring reasoning

A method includes obtaining an image, the image associated with a mask corresponding to a portion of the image, generating a plurality of images based on the image and the mask, each image of the plurality of images depicting a different color in the portion of the image corresponding to the mask, executing a machine learning model to generate an image performance score for each of the plurality of images, ranking the plurality of images according to the image performance scores for the plurality of images, and generating a record comprising one or more images of the plurality of images based on the rankings of the plurality of images.
Owner:VIZIT LABS INC

Adaptive sensing-based lightweight monitoring method for fine crack in complex background region

The present invention relates to an adaptive sensing-based lightweight monitoring method for a fine crack in a complex background region. The method comprises the following steps: step S1, on the basis of region division, performing automatic acquisition of crack information, wherein PTZ camera sensors are used to automatically perform block-wise acquisition on crack regions; step S2, performing an adaptive complex scale calibration process, using a multi-scale template matching algorithm to adaptively correct distortion information of all regions, and performing real-scale conversion from pixel precision; step S3, constructing a lightweight crack segmentation network to process data processed in step S2; and step S4, by means of a quantitative crack-tracking algorithm based on Euclidean distance similarity classification, performing real-time monitoring on each piece of crack dynamic information. Compared with the prior art, the present invention has advantages such as achieving efficient, accurate, and online monitoring and analysis of cracks.
Owner:SOUTHEAST UNIV

Defect detection method for semiconductor packaging material based on deep learning

The invention relates to the field of semiconductor packaging material defect detection, in particular to a semiconductor packaging material defect detection method based on deep learning, which comprises the following steps: acquiring a surface image, and extracting a two-dimensional contour and a feature point set; preprocessing the image, and separating a packaging material main body area; constructing a two-dimensional defect identification model based on Transform, and outputting a two-dimensional detection result; scanning suspected and unknown defect areas to obtain three-dimensional point cloud data, and extracting geometric and texture features; fusing two-dimensional and three-dimensional data through a space-time alignment model; utilizing the multi-modal fusion model to output defect positions and types; and evaluating the defect importance based on the material node connectivity and the stress distribution, and generating a visual detection report. According to the invention, high-precision detection of semiconductor packaging material defects is realized, the defect identification rate, the positioning precision and the detection efficiency are improved through multi-modal data fusion and a deep learning model, and a visual report can be generated based on material structure quantification defect importance.
Owner:XIAN UNIV OF POSTS & TELECOMM

LED display defect prediction and process adjustment method and system based on multi-modal fusion

The invention relates to the technical field of LED display, solves the problem that the existing LED display defect detection and parameter adjustment technology is lack of multi-modal information fusion and intelligent process control capability and is difficult to meet the quality control requirement of a high-precision display product, and provides an LED display defect prediction and process adjustment method and system based on multi-modal fusion. The method comprises the following steps: performing multi-modal data fusion processing on optical image data, electrical test data and thermal infrared imaging data corresponding to a to-be-tested LED display screen to obtain fused data; inputting the fused data into a pre-trained defect recognition model to obtain a defect recognition result; according to a process parameter adjustment strategy corresponding to the defect identification result, adjusting the original process parameter to obtain a target process parameter; and according to the target process parameters, process flow correction processing is carried out, and a qualified LED display screen is produced. According to the method, the defect identification precision is improved, and the quality control requirement of high-precision LED display screen production is met.
Owner:XIAMEN PROD QUALITY SUPERVISION & INSPECTION INST +1

Auxiliary dental implant generation method based on diffusion model

The present invention relates to the technical field of stomatology. Provided is an auxiliary dental implant generation method based on a diffusion model. The method in the present invention comprises: acquiring oral CBCT image data of historical patients, preprocessing the oral CBCT image data of the historical patients to obtain a CBCT image dataset, using the CBCT image dataset to train a multi-task segmentation network, and using the segmentation network to obtain an intraoral tissue segmentation result; using the intraoral tissue segmentation result to train detection networks from the three dimensions of a cross-sectional plane, a coronal plane and a sagittal plane, respectively; using the detection networks to obtain detection results in the three directions of the cross-sectional plane, the coronal plane and the sagittal plane; fusing the detection results in the three directions of the cross-sectional plane, the coronal plane and the sagittal plane, and using a majority voting algorithm to construct a three-dimensional bounding box, so as to acquire an edentulous area; and using the intraoral segmentation result and the edentulous area as prompt information to guide, by means of an iterative process, a network to generate a post-implantation effect. The implantation effect obtained by the present invention is highly accurate, thereby providing a more precise auxiliary tool for stomatology.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Mama-based spectrum dynamic fusion and double attention enhancement medical image segmentation method

The invention discloses a Mama-based spectrum dynamic fusion and double-attention enhancement medical image segmentation method, which comprises the following steps of: firstly, constructing a Mama integrated spectrum domain and attention pyramid module, fusing spectrum dynamic characteristics and a self-attention pooling mechanism, and performing frequency domain information compensation and local characteristic enhancement to obtain a spectrum dynamic fusion image; the spatial correlation loss caused by image blocking processing is relieved; secondly, designing a layered enhanced U-shaped architecture, deploying an MISAP module in a shallow layer of an encoder to capture multi-scale global context features, introducing a bipolar routing attention mechanism in a deep layer, and dynamically allocating sparse attention weights to focus a key pathological region; according to the method, the segmentation precision of complex edge textures and tiny lesions in medical images can be remarkably improved, and the Dice coefficient in breast tumor, polyp and abdominal organ segmentation tasks is averagely improved by 6.5%.
Owner:SHAANXI UNIV OF SCI & TECH

Large-scene three-dimensional reconstruction method based on three-dimensional Gaussian sputtering

The invention discloses a large-scene three-dimensional reconstruction method based on three-dimensional Gaussian sputtering, and relates to computer graphics. The method comprises the following steps: collecting a multi-view image set of a large scene; obtaining a scene sparse point cloud according to the multi-view image set; performing monocular depth estimation on the multi-view image by using a pre-trained depth prediction network to obtain monocular depth estimation priori; the method comprises the following steps of: performing global training on a scene by utilizing scene sparse point cloud and monocular depth estimation prior to obtain an initial three-dimensional Gaussian model, and performing space grid division on the initial three-dimensional Gaussian model to obtain a plurality of scene blocks with axis alignment bounding boxes; setting image view angle data of each scene block; performing deep supervised training on the Gaussian ellipsoids in the plurality of scene blocks by using a parallel GPU (Graphics Processing Unit); combining the trained scene blocks to obtain a final three-dimensional Gaussian model; in view of low geometric structure reconstruction precision caused by only depending on color information of a multi-view image in large-scene three-dimensional rendering, the method improves the reconstruction precision of large-scene rendering.
Owner:JSTI GRP CO LTD +2

Model deployment method, end-side device, and storage medium

The present disclosure relates to the technical field of target detection, and particularly relates to a model deployment method, an end-side device and a storage medium, which are used for solving the problem in the related art of the accuracy of a deployed model being low. The method comprises: performing target detection on a video frame image input into a first model, and acquiring a first target detection result and a first confidence; if the first confidence is greater than or equal to a first confidence threshold value, recording the video frame image and the first target detection result as samples in a training set; if the first confidence is less than the first confidence threshold value, performing target detection on the video frame image on the basis of a second model, and recording the video frame image and an acquired second target detection result as samples in the training set; and training the first model on the basis of the training set, and replacing the current first model with a trained first model for subsequent target detection. In this way, the accuracy and model generalization capability of a first model are improved.
Owner:HISENSE GRP HLDG CO LTD

PCB (Printed Circuit Board) defect detection method and system based on image recognition

The invention relates to the field of defect detection, in particular to a PCB defect detection method and system based on image recognition. The method comprises the following steps: collecting a multidirectional PCB detection image, carrying out pixel-level registration correction and adaptive pixel stability compensation, and constructing a space-time stability compensation image sequence; performing reverse pyramid structure division on the space-time stability compensation image sequence, performing normalized similarity probability calculation, and constructing an initial region classification result; based on an initial region classification result, depth image visual analysis is carried out, pseudo defect comprehensive elimination optimization is carried out, and a pseudo defect purification high-confidence image is constructed; pCB connection defect identification is carried out on the pseudo defect purification high-confidence image, global defect point distribution marking is carried out, and a defect point space distribution diagram is constructed. According to the invention, high-credibility, high-precision and high-closed-loop PCB defect detection is realized.
Owner:SHENZHEN HTWY TECH CO LTD

Image enhancement method and system in complex coal mine environment

The invention discloses an image enhancement method and system in a complex coal mine environment, and relates to the technical field of image processing, and the method comprises the steps: carrying out the preprocessing of a collected coal mine image of a target region, and dividing the coal mine image into different semantic regions, including a bright region, a dark region and a dust shielding region, through a deep learning semantic segmentation model; according to semantic region characteristics, a differentiation enhancement strategy is made; a traditional Retinex model is improved, non-local mean filtering is introduced, and an illumination component and a reflection component are decomposed through pixel similarity matching. According to the method, the image is divided into the bright area, the dark area and the dust shielding area through the deep learning semantic segmentation model, differential enhancement strategies are formulated according to different area characteristics, detail distortion caused by global adjustment is avoided, local contrast suppression is adopted in the bright area, illumination compensation is enhanced in the dark area, and the image quality is improved. Noise diffusion of the dust shielding area is inhibited through edge preservation smoothing, the image quality of each area is remarkably improved, and it is ensured that image details in a complex coal mine environment are clear and visible.
Owner:CHINA COAL TECH GRP INFORMATION TECH CO LTD

Construction site safety risk intelligent early warning system and method based on BIM and big data analysis

The invention discloses a construction site safety risk intelligent early warning system and method based on BIM and big data analysis, relates to the technical field of building engineering construction safety, and solves the problem that it is difficult to transmit construction site multi-source data which is collected and preprocessed in real time in real time and carry out space mapping with a BIM model. A rule engine is difficult to carry out initial early warning; a machine learning model is difficult to analyze time series data, predict collapse risks and identify dangerous behaviors; a risk prediction model is difficult to construct and is difficult to integrate into a BIM model; and pushing and closed-loop management are difficult to carry out on the risk early warning information. According to the method, the multi-source data is collected at the construction site, the digital twinborn scene is constructed by mapping the multi-source data to the BIM model by means of space-time alignment, the multi-source data is analyzed and processed by applying technologies such as a rule engine and a machine learning algorithm, and the result is integrated to the BIM model, so that visual risk monitoring and early warning are realized.
Owner:BEIJING ZHENDONG LIANKE TECH CO LTD

PCB (Printed Circuit Board) defect detection method and system

The invention relates to the technical field of PCB detection, and discloses a PCB defect detection method and system, and the method comprises the steps: collecting multispectral imaging data through an image collection module, and generating an original image data set; the defect analysis server receives the synchronous imaging data to construct a three-dimensional surface topology matrix; in combination with the original image data set and the real-time imaging data, performing multi-scale decomposition on the three-dimensional surface topological matrix, extracting texture features, positioning a defect region, outputting defect type space distribution features through a layered recognition model, and updating the original image data set; and dynamically calibrating the detection parameters according to the feature categories. The system comprises an image acquisition module group, a data transmission module, a three-dimensional modeling module, a defect identification module and a parameter calibration module. According to the scheme, the accuracy, comprehensiveness and efficiency of defect detection are improved, and the detection requirements of modern PCB production are met.
Owner:SHENZHEN UNITED MULTILAYER CIRCUIT BOARD CO LTD

Enterprise smart legal affair platform system based on generative language large model

The invention discloses a hybrid enhanced enterprise smart law platform system based on a generative language large model. Four modules including a hybrid enhanced legal knowledge engine, a multi-modal legal document analysis module, a risk quantitative evaluation module and a compliance verification workflow work cooperatively. The hybrid enhanced legal knowledge engine integrates multi-source data, realizes real-time updating and semantic reasoning, and comprises map construction, a rule base and an incremental learning mechanism; the multi-modal legal document analysis module performs structured analysis on the heterogeneous document to generate a feature vector; the risk quantitative evaluation module is combined with Monte Carlo simulation and an analytic hierarchy process, quantifies the risk according to a compliance reference and analysis characteristics, and outputs a thermodynamic diagram and a report; a compliance verification workflow is driven by a finite-state machine, a verification module and a conflict detection module are integrated, a generative language large model is called to generate an improved scheme, and audit records are solidified and fed back for optimization. And the system runs according to the processes of analysis, supply rule, bias calculation and verification correction, so that the intelligence and accuracy of legal affair processing are improved.
Owner:邢嘉怡

Water conservancy and hydropower engineering construction safety supervision system and method based on multi-source data fusion

The invention belongs to the technical field of water conservancy and hydropower engineering, and discloses a water conservancy and hydropower engineering construction safety supervision system based on multi-source data fusion. The system comprises a multi-source sensing acquisition module, a heterogeneous data fusion processing module, a risk identification and early warning module, a safety behavior evaluation and feedback module, and a command scheduling and visualization module. According to the invention, by fusing multi-dimensional data such as image monitoring, environment sensing, personnel positioning, equipment state and the like, a space-air-ground three-dimensional sensing network is constructed, and in a high slope area, the distributed optical fiber strain sensors are linked with thermal imaging data of the unmanned aerial vehicle, so that millimeter-level deformation and temperature field abnormity can be captured in real time; a video stream is analyzed in real time by means of a YOLOv8 algorithm, illegal operation behaviors of personnel can be accurately identified, a cross-modal fusion model of a Transform architecture is combined, the system can dynamically capture potential correlation among data, and millisecond-level response to risks such as side slope landslide, equipment faults and personnel dangerous operation is achieved.
Owner:YUNNAN TUOMEI DECORATION ENGINEERING CO LTD

Classification of Image Data from Synthetic Aperture Radar Images and Electro-Optical Images with Multi-Modal Fusion

Systems and methods are disclosed for classifying objects using electro-optical and synthetic aperture radar images through multi-modal feature alignment and fusion. A computing system acquires and preprocesses image data, then aligns features across modalities using a multi-modal alignment engine. A cross-modal attention fusion network extracts and integrates complementary information using transformer-based attention mechanisms. A modality-specific feature extraction framework processes EO and SAR images through specialized branches, ensuring optimal feature representation. An adaptive fusion decision system dynamically determines the best fusion strategy based on image quality and confidence scores. A self-supervised consistency controller enforces alignment between EO and SAR features using contrastive learning. The fused representations are processed by a neural network to generate object classifications. This system improves accuracy and robustness in environments where one modality may be degraded or missing, enhancing applications such as remote sensing, surveillance, and autonomous navigation.
Owner:ATOMBEAM TECH INC

Apparatus for automatically setting measurement reference element and measuring geometric feature of image

InactiveUS20020057828A1automatic measurement of the geometric feature of the object image can be efficientlyefficient measurementImage enhancementImage analysisReference imageImaging data
In a measurement processing apparatus for measuring a geometric feature of an object image: a measurement-reference-element setting unit automatically sets at least one first measurement reference element for use in measurement of the geometric feature of the object image, at at least one first position on the object image based on first image data representing the object image and position information indicating at least one second position of at least one second measurement reference element which is set on a measurement reference image corresponding to the object image; and a geometric-feature measurement unit measures the geometric feature of the object image based on the at least one first position of the at least one first measurement reference element.
Owner:FUJIFILM CORP

PCBA board defect detection method and system based on image processing

The invention relates to the technical field of image detection, in particular to a PCBA board defect detection method and system based on image processing, and the method comprises the following steps: carrying out the meshing calculation of a gray scale deviation after a gray scale image is subjected to Gaussian filtering denoising, generating change rate data, carrying out the statistics of a frequency number, constructing a histogram, combining with an Otsu algorithm, and generating a candidate mask; extracting pixels based on a mask, calculating a gradient modulus, screening edge candidate points, carrying out gradient direction connection and morphological processing to generate a complete edge structure, expanding a connected domain through a region growing algorithm, aligning the connected domain with a template contour, and outputting defect coordinates. According to the method, the defect identification sensitivity is improved through combination of gray level image gridding processing and dynamic threshold calculation, a candidate mask is generated through grid gray level change rate statistics and an Otsu algorithm to avoid over-segmentation missing detection, and the contour precision is improved through combination of gradient modulus difference screening and morphological closed operation optimization. The region growing algorithm and template dynamic alignment reduce deformation misjudgment, and staged dimension reduction and feature enhancement reduce calculation complexity and solve resource waste.
Owner:广东德智矩阵科技有限公司

Construction scene prediction method and device fusing image, text and BIM mode

The invention provides a construction scene prediction method and device fusing an image, a text and a BIM modal, and relates to the technical field of intelligent construction prediction management. According to the method, the BIM semantic graph is constructed by extracting the BIM semantic information of the BIM model, and the target detection recognition of the construction site video is carried out in combination with the YOLO model to obtain the target detection result; performing cross-modal alignment with the CLIP to realize deep fusion of the multi-modal data of the image, the text and the BIM to obtain a multi-modal heterogeneous graph; and inputting the multi-modal heterogeneous graph into a space-time sequence model for prediction, outputting prediction results of construction scenes at a plurality of moments in the future, and dynamically mapping the prediction results to a digital twinborn platform to realize risk early warning and visual display. The construction dynamic change can be captured in real time, the construction progress and risk can be accurately predicted, and the intelligent level of construction management is improved.
Owner:XIAMEN UNIV OF TECH

Gait emotion recognition method, system, storage medium, and computer equipment based on spatiotemporal graph convolution.

This invention relates to a gait emotion recognition method, system, storage medium, and computer device based on spatiotemporal graph convolution. The method includes the following steps: S1, data augmentation by reversing the temporal direction of gait; S2, obtaining deep emotion features and prior emotion features respectively through a spatiotemporal graph convolutional network and prior feature statistical methods; S3, performing nonlinear mapping on the prior emotion features using a feature mapping layer; S4, inputting the fused features of the deep emotion features and prior emotion features into an emotion classifier to obtain the emotion category. The feature mapping layer of this invention achieves more effective feature fusion by performing nonlinear mapping on prior features; it also introduces causal temporal convolution to replace general temporal convolution, effectively extracting fine-grained temporal features by enhancing temporal correlation and cross-period feature fusion. Furthermore, a walking direction recognition auxiliary task is designed to accelerate the training and convergence speed of the model, enhancing the ability to extract temporal-dependent features and the performance of emotion recognition.
Owner:SOUTH CHINA UNIV OF TECH

Self-adaptive question-answering system and method based on knowledge distillation and multi-modal dynamic fusion

The invention discloses an adaptive question-answering system based on knowledge distillation and multi-modal dynamic fusion, and the system comprises a knowledge distillation module which is used for migrating knowledge of a teacher model pre-trained on corpora in the communication field to a lightweight student model, achieving model compression through optimizing a distillation loss function, and obtaining a multi-modal dynamic fusion model; the loss function comprises a soft label output by the teacher model and a KL divergence constraint output by the student model; the multi-modal knowledge fusion module comprises a feature extraction unit, a self-adaptive weighting unit and an attention fusion unit; the self-adaptive inference engine comprises a semantic analysis unit; according to the cross-modal reasoning method and system, semantic alignment of equipment parameters, protocol texts and topological graphs is achieved through the multi-modal dynamic fusion technology, and the cross-modal reasoning accuracy is improved; compared with an original model, the lightweight student model has the advantage that the reasoning speed is increased in a protocol analysis task.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Power monitoring system and method integrating image recognition and data analysis

The invention relates to the field of electric power monitoring, and discloses an electric power monitoring system and method fusing image recognition and data analysis, and the method comprises the steps: carrying out the visual angle coverage modeling of a target equipment group through a multi-type visual collection unit disposed at a transformer substation and a power distribution terminal; performing cross-frame fine-grained texture differential analysis on the equipment state image sequence, and constructing an image event time window in combination with synchronous disturbance characteristics of multi-source monitoring parameters; based on the high-vigilance candidate frame set, fusing the image structure variability index and the operation data multi-dimensional deviation vector by using a feature encoder, and constructing a multi-modal state coupling feature tensor; map mapping is carried out on the potential fault evolution trend, and semantic association is established between structural nodes with abnormal attributes in the image and frequently fluctuating parameter indexes in the monitoring data; and combining a node interference path in the local fault association subgraph with fault precursor distribution induced in a historical accident sample. The method has the advantage that the operation safety is improved.
Owner:HANGZHOU HOFF ELECTRICAL AUTOMATION

Intelligent forest pest and disease damage monitoring method and system based on unmanned aerial vehicle remote sensing

The invention relates to the technical field of remote sensing monitoring, in particular to an intelligent forest disease and pest monitoring method and system based on unmanned aerial vehicle remote sensing, and the method comprises the following steps: collecting multispectral data by an unmanned aerial vehicle, extracting reflectivity and smoothing the reflectivity, carrying out differential recognition on abnormal pixels, extracting curves and screening significant changes, segmenting scab boundaries, and classifying health states. And generating a pest and disease map layer prediction trend. According to the method, the reflectivity time sequence is constructed, differential processing is carried out, the vegetation change trend is dynamically captured, abnormal areas are identified by combining slope offset and persistence analysis, significant pixels are screened according to main peak wavelength offset, boundary information is extracted, the scab positioning precision is improved, and the recognition resolution of the lesion state is enhanced through reflectivity combined analysis; accurate description of disease spot dynamic changes is realized, static image dependence limitation is broken through, monitoring time continuity and space response capability are enhanced, and disease and insect pest change capture efficiency and state classification accuracy are effectively improved.
Owner:SHIHEZI UNIVERSITY

Automatic financial information processing method based on AI

The invention discloses an AI-based automatic financial information processing method, and relates to the field of financial automation, and the method comprises the steps: achieving the automatic collection and storage of structured and unstructured data through the access of enterprise multi-source financial data; systematic preprocessing is carried out on the collected multi-source heterogeneous financial data, and a unified and high-quality financial data set is constructed; based on natural language processing and a knowledge graph technology, performing text semantic understanding, transaction automatic classification, field standardization and label generation on the cleaned and integrated financial data; comprehensively quantifying enterprise operation and financial performance based on the structured transaction data and the semantic annotation result; based on historical financial indexes, establishing a multi-model architecture to predict key financial variables; and based on the structured data, the prediction result and the historical rule, identifying potential financial abnormity and risk behaviors, and realizing intelligent early warning. According to the method, the intelligence, the real-time performance and the accuracy of financial information processing can be remarkably improved.
Owner:CHANGSHA DILU DIGITAL TECH

Multi-modal data processing method and apparatus, electronic device, computer-readable storage medium, and computer program product

Disclosed in the present application are a multi-modal data processing method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a reference image and a reference text; extracting a reference visual feature of the reference image; by means of a multi-modal large language model, determining an embedding of the reference text, an embedding of a start mark of the reference visual feature, an embedding of the reference visual feature, and an embedding of an end mark of the reference visual feature; on the basis of the multi-modal large language model, splicing the embedding of the reference text, the embedding of the start mark, the embedding of the reference visual feature, and the embedding of the end mark into a target embedding sequence, performing attention processing on the basis of the embedding of the start mark, the embedding of the end mark, and an embedding selected by a sliding window in the target embedding sequence, and outputting a predicted sequence; and generating a predicted image and a predicted text on the basis of the predicted sequence.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD