Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

383 results about "Structure extraction" patented technology

Semantic segmentation method for low-resolution road scene

The invention discloses a semantic segmentation method for a low-resolution road scene, and aims to solve the problems of difficulty in small target recognition, fuzzy details, texture information loss and the like existing in a low-resolution image in the conventional semantic segmentation technology. The method comprises the following steps: (1) collecting a low-resolution road scene image and a corresponding semantic tag; (2) constructing a semantic segmentation model consisting of an edge guidance module (BGM), a double-domain feature decomposer (DDFD), a domain alignment attention fusion module (DAAFM) and a double-layer attention context aggregation module (HACAM); (3) designing a joint loss function to carry out multi-scale supervision on semantic regions, edges and middle features; (4) carrying out model training by utilizing the road scene image; and (5) outputting a semantic segmentation result map and an edge prediction map. The boundary perception capability is enhanced by introducing learnable pixel difference convolution, the extraction precision of a small target and a global structure is improved by combining frequency domain and spatial domain feature alignment, and context semantic relationship expression is optimized by fusing a channel and a spatial attention mechanism. The method effectively improves the semantic segmentation precision and boundary restoration capability of the model in a low-resolution complex road environment, and is suitable for intelligent analysis tasks of road images in scenes of automatic driving, intelligent traffic, severe weather and the like.
Owner:CENT SOUTH UNIV

Multi-mode fusion product document and source code association retrieval method based on knowledge graph

The invention discloses a multi-mode fusion product document and source code association retrieval method based on a knowledge graph, and relates to the technical field of software engineering and artificial intelligence. The method comprises the steps that source codes are preprocessed, and code structure information and business semantics are mapped in combination with a predefined business term dictionary; performing controlled induction on a code file and a product document by utilizing a large model, extracting business terms, logic intentions and a subject relationship, fusing with original codes, and establishing a vector retrieval index system; further analyzing a code structure by using an abstract syntax tree, and extracting an entity and a calling relationship; semantic enhancement and relation normalization are performed in combination with the large model, entities and relations are stored in a graph database, and a knowledge graph is formed; and performing parallel processing on user query based on a full-text retrieval index, a vector retrieval index system and a knowledge graph, and finally generating a product concept. According to the method, the retrieval speed, the semantic depth and the logical reasoning ability can be considered at the same time, and the retrieval accuracy is improved.
Owner:MARCO POLO TRAVEL TECH CO LTD

Multi-modal index knowledge base, construction method thereof and question and answer processing method

The invention discloses a multi-modal index knowledge base and a construction method thereof. The construction method comprises the following steps: processing a heterogeneous document to obtain a semi-structured document; identifying a title hierarchical relationship of the document to construct a document logic structure; carrying out minimum chapter blocking on the text of the semi-structured document to obtain logic blocks; performing semantic segmentation on each logic block to obtain text blocks; the method comprises the following steps: constructing text block nodes by meta-information of text blocks, constructing non-text block nodes by meta-information of non-text elements, extracting document nodes, chapter nodes and chapter-chapter inclusion relationships according to a document logic structure, respectively extracting semantic information from the text blocks and the non-text elements, and storing the semantic information in a database; recording the corresponding relationship between the text block nodes and the semantic information and between the non-text block nodes and the semantic information; constructing a knowledge graph based on each node and relationship and storing the knowledge graph into a graph database; and constructing semantic knowledge based on the semantic information and storing the semantic knowledge into a vector database. According to the scheme, lossless retention of multi-modal information and structured organization of document logic are realized, and efficient indexing and accurate recall are facilitated.
Owner:浙江泰隆商业银行股份有限公司

Water conservancy industry electronic dark bidding document enterprise internal examination method and system based on artificial intelligence

The invention provides a water conservancy industry electronic dark bidding document enterprise internal examination method and system based on artificial intelligence, and belongs to the field of water conservancy industry bidding. Performing automatic pre-auditing on the bidding file: performing compliance inspection and integrity inspection, identifying problems existing in the bidding file, and providing improvement suggestions to ensure that the problems conform to the basic requirements of electronic dark label review; carrying out data standardization and cleaning on the bidding file: carrying out semantic extension and ambiguity elimination, identifying images, tables and handwritten contents in the bidding file, identifying a main body structure of the water conservancy design drawing, extracting engineering quantity list data, and carrying out compliance verification; identifying an abnormal behavior in the bidding file through a machine learning model; and an internal examination decision tree is constructed according to preset key indexes and score weights of the water conservancy project, an interpretable artificial intelligence algorithm is utilized to carry out internal examination decision making on the bidding document, an internal examination score result is fed back, and specific score deduction reasons are explained. And the standardization and the accuracy of internal examination of the bidding document are improved.
Owner:POWERCHINA BEIJING ENG CORP

Highly dense broken ice image segmentation method based on iterative MGAC and SAM model

The invention discloses a highly dense broken ice image segmentation method based on an iteration MGAC and an SAM model, and the method comprises the steps: carrying out the sea ice instance segmentation of a preprocessing image based on the SAM model, and obtaining a sea ice mask image corresponding to the preprocessing image; according to the binary image and the sea ice mask image, obtaining an initial sea ice residual region image which is not identified by the SAM model; obtaining initial sea ice residual region images under different gray threshold values to obtain an initial seed mask graph; taking the initial sea ice residual region image and the initial seed mask image as inputs of a preset MGAC contour model, and obtaining an MGAC sea ice recognition result based on a multi-round iteration partitioning mechanism; and performing union operation on the MGAC sea ice identification result and the sea ice mask image to obtain a crushed ice segmentation mask result. The method solves the problem that the existing method is insufficient in structure extraction precision and boundary integrity of the dense broken ice area.
Owner:DALIAN MARITIME UNIVERSITY

New media AI marketing content creation method and device, equipment and medium

The invention relates to a new media AI marketing content creation method and device, equipment and a medium. The method comprises the steps of obtaining user demand configuration, and obtaining a creation intention vector through natural language analysis; based on the platform characteristic knowledge base, extracting a corresponding structure specification and a propagation mechanism according to the target platform, and encoding to obtain a platform characteristic vector; obtaining historical content interaction data corresponding to the audience group, and generating a user-content interaction vector by adopting collaborative filtering and a label similarity algorithm in combination with the content keyword; and calling an artificial intelligence large model, and performing content generation according to the creation intention vector, the platform feature vector and the user-content interaction vector to obtain content creation data. By adopting the method, the goal of automatic, high-quality and personalized new media marketing content creation in a multi-platform environment can be realized by means of natural language analysis, knowledge structure extraction, large model generation constraint and the like.
Owner:JIANGSU XUZHOU HIGHER VOCATIONAL & TECH SCHOOL OF FINANCE & ECONOMICS

Method and system based on NLP file analysis

The invention provides a method and system based on NLP file analysis, and relates to the technical field of natural language processing. According to the method, time and identifier unification and format and character set standardization are carried out on the multi-source file, layout segmentation, table structure extraction, reference analysis, term standardization and anaphora resolution are combined, semantic representation is constructed, a hierarchical index and a unique traceability identifier are generated, intention recognition, retrieval sorting, incremental updating and consistency verification are supported, and the method is suitable for large-scale popularization and application. Unification, semantization and traceability of the file analysis process are achieved, and the processing efficiency and accuracy are improved.
Owner:ZUNYI NORMAL COLLEGE

PCB defect real-time detection method based on multi-scale feature fusion

The invention discloses a PCB defect real-time detection method based on multi-scale feature fusion, and relates to the technical field of PCB defect real-time detection methods, and the method comprises the steps: obtaining a to-be-detected PCB image, carrying out the size normalization and pixel value standardization processing of the image, and obtaining a standardized image meeting the input requirements of a model; inputting the standardized image into a backbone network of a teacher detection model, and extracting a multi-scale primary feature map containing texture information in different directions through a grouping convolution structure; transmitting the multi-scale primary feature map to a neck network of a teacher detection model, and performing weighted fusion on feature maps of different scales by using a learnable weight to generate a multi-scale fusion feature map; and in an up-sampling path of the neck network, generating channel description information after global pooling is performed on the deep fusion feature map, generating a channel attention weight through nonlinear transformation, acting the weight on a primary feature map of a corresponding level, and outputting an enhanced feature map.
Owner:SHAANXI SCI TECH UNIV

Thin sheet type component performance rapid prediction method based on deep learning

ActiveCN120596856AFeature setAlgorithm
The invention relates to the technical field of artificial intelligence, in particular to a sheet part performance rapid prediction method based on deep learning, which comprises the following steps: collecting multi-working condition simulation data to generate a training sample, constructing and coding a grid topological structure to extract multi-dimensional features, and inputting a perceptron to predict stress and evaluate errors after feature fusion and self-attention mechanism processing. According to the method, a structured training sample set is constructed by introducing simulation information, a geometric structure feature set is formed by combining node space coordinates, boundary constraints and a connection relation, so that mutual positions and constraint conditions among nodes are completely expressed in a graph structure, and through node-level feature extraction and feature fusion processing, a graph structure is obtained. According to the method, deep embedding of node geometric layout and boundary interrelation is realized, learnable expression of a stress evolution path in a space structure is established through local subgraph and context analysis, a multi-layer feature aggregation and attention mechanism is introduced in a node graph embedding process, and feature response expression of a key area is enhanced.
Owner:CHONGQING HUIQIAN TECH CO LTD

Drainage basin distributed runoff prediction method and system based on graph neural network

The invention relates to the technical field of hydrological prediction, in particular to a drainage basin distributed runoff prediction method and system based on a graph neural network, and the method comprises the following steps: obtaining a drainage basin multi-source runoff data set to construct a multi-relation dynamic graph structure, and extracting node feature vectors; based on the node feature vector, obtaining a watershed evolution trend forward feature by establishing a watershed diffusion fitting architecture; constructing a distributed runoff probability prediction model, and taking the watershed trend forward features as model input to obtain runoff initial condition probability distribution of each sub-watershed in multiple periods in the future; establishing a mixed loss function, and performing physical constraint optimization on the runoff initial condition probability distribution to obtain distributed runoff optimization probability distribution; and performing uncertainty quantification on the distributed runoff optimization probability distribution to obtain a drainage basin distributed runoff prediction result. According to the method, the hydrological process simulation capability of the complex watershed is improved, and accurate prediction of the distributed runoff volume is realized.
Owner:HENAN UNIVERSITY

Protein compound model interface quality evaluation method based on multi-scale isotropic graph neural network

A protein complex model interface quality evaluation method based on a multi-scale isovariant graph neural network comprises the following steps: firstly, screening out a co-crystallized natural protein complex structure from a non-redundant protein interaction database PRISM, and generating a bait structure by using a HDock docking algorithm; the method comprises the following steps: firstly, extracting molecular surface interaction fingerprints, atomic-level features and residue-level features on the basis of each compound bait structure, obtaining graph representation of the compound bait structures, then fully capturing and fusing multi-scale information through a depth isotropic graph neural network, and finally obtaining an interface mass fraction through prototype comparison prediction. According to the method, the interface quality evaluation of the protein compound model can be accurately carried out, and the problems of low precision and poor generalization of the interface quality evaluation of the protein compound model are effectively solved.
Owner:ZHEJIANG UNIV OF TECH

Terrain change detection system based on unmanned aerial vehicle

The invention relates to the technical field of topographic change analysis, in particular to an unmanned aerial vehicle-based topographic change detection system, which comprises a slope direction sensing track control module, a texture structure extraction module, a crack evolution track construction module, a direction trend comparison module and a patrol recheck positioning module. According to the method, a continuous elevation point column of an unmanned aerial vehicle scanning area is extracted, laser reflection point coordinates are fused, a space relation of transition point distribution is constructed, dynamic adjustment of a ground-imitated flight path is achieved, and a texture structure area with continuous directivity is recognized in combination with a high-angle image boundary communication relation; texture boundary evolution is compared at different time nodes to form a crack path, the stability of the path and the slope direction is judged through an included angle sequence, recognition and sorting of areas with the consistent direction are completed, a space comparison result is registered in a three-dimensional coordinate system, terrain change areas are accurately marked, and rapid positioning and continuous tracking of high-risk areas are achieved.
Owner:SHANDONG TRAFFIC PLANNING DESIGN INST

Multi-level detail automatic simplification method for oblique photography live-action three-dimensional model

The invention discloses a multi-level detail automatic simplification method for an oblique photography live-action three-dimensional model, and relates to the technical field of three-dimensional model simplification and computer graphics, and the method comprises the steps: obtaining oblique photography original data and three-dimensional model basic information; preprocessing the model, performing adaptive Gaussian filtering denoising, improving RANSAC to remove outer points, compressing textures in a blocking manner, correcting mapping coordinates, and repairing a topological structure; extracting multi-scale features; constructing a simplified decision model, and determining a simplification rate and a priority by combining an observation distance, scene precision and hardware performance; performing hierarchical simplification, vertex hierarchical improved edge folding, patch hierarchical adaptive deletion and regional hierarchical grid reconstruction; performing multi-dimensional quality evaluation, and if the requirements are not met, performing backtracking adjustment; and outputting a simplified model stored according to the LOD hierarchy, wherein the simplified model comprises transition information and a simplified log. According to the method, the data quality is improved through refined preprocessing, the simplification pertinence is enhanced through multi-dimensional feature extraction and intelligent decision, and the application value of the model is improved.
Owner:HUNAN CHUANGXIN WEILI TECH CO LTD

Original script-oriented AI autonomous plot structure adaptive generation system

PendingCN121233763ASemantic analysisBiological modelsNeural oscillationAlgorithm
The invention discloses an original script-oriented AI autonomous plot structure adaptive generation system, and relates to the technical field of creation assistance, and the system specifically comprises the following modules: a structure extraction analysis module, an emotional role analysis module, a plot inference module, a scene generation optimization module, a conservation target generation module, an adaptive control module, and a constraint punishment module. According to the method, a multi-level narrative structure and a causal relationship graph are constructed to form an emotion vector and a trajectory curve, an optimal causal path is generated based on a graph neural network, a graph convolution / attention mechanism and a graph generation algorithm, and emotion toning and conservation target driven text generation are performed on scene and dialogue levels. Structural consistency and emotional arcs are optimized in real time in combination with self-adaptive control, plot path weighting and rewriting triggering are performed by utilizing a multi-dimensional emotional space and a neural oscillator network, intelligent structured management of a script is realized, and script creation efficiency and quality are improved.
Owner:GOLDEN TIMES CULTURE COMM

Intelligent enterprise compliance auditing method based on data driving

The invention discloses an enterprise intelligent compliance auditing method based on data driving. The method comprises the following steps: S1, automatically collecting auditing data of various heterogeneous data sources in an enterprise in real time through a cross-domain data access interface; s2, performing feature automatic identification and standardization processing on the audit data by adopting a semantic adaptive coding method; s3, a cross-domain collaborative characterization model is constructed by extracting and fusing shared features through a sub-domain multi-expert structure; s4, generating a visual feature heat map by using a hierarchical attention mechanism of the model, and outputting an anomaly detection result; s5, extracting long and short period correlation mode features, inputting the features into a contrast learning framework, and outputting an abnormal risk score; s6, constructing a dynamic enterprise compliance risk knowledge graph; s7, strategy training is carried out, and risk rating is output; and S8, generating an audit report according to the risk rating. According to the invention, efficient and accurate enterprise compliance risk identification and audit decision support are realized.
Owner:LIANYUNGANG JIRAN INFORMATION TECHNOLOGY CO LTD

Dynamic protection constant value cooperation method for 30-degree phase angle difference non-perception loop closing of power distribution network

The invention discloses a power distribution network 30-degree phase angle difference non-perception loop closing dynamic protection constant value cooperation method, and relates to the technical field of power grid automation, and the method comprises the following steps: extracting a protection equipment connection relation based on a power distribution network topology structure, and constructing a topology model containing a hierarchical membership relation; according to the real-time measurement data, generating an interval load distribution model by adopting a periodic entropy weight correction removing algorithm; simulating a load transfer process under a 30-degree phase angle difference based on the topology model and the interval load distribution model, and outputting current change time sequence data of the transfer path protection equipment; according to the hierarchical membership relationship and the current change time sequence data, calculating and checking a protection constant value, and dynamically adjusting the protection constant value by using an impact tolerance coefficient to generate a final protection constant value; and issuing the final protection constant value to the protection equipment, and triggering load transfer operation.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD BOZHOU POWER SUPPLY CO +1

Single tree trunk structure extraction method and system based on deep learning

The invention discloses a single tree trunk structure extraction method and system based on deep learning, and the method comprises the steps: obtaining two-dimensional image data of a single tree, marking the two-dimensional image data, and constructing a single tree trunk segmentation data set; constructing a spatial domain and frequency domain double-branch network based on the segmented data set, respectively extracting spatial domain features and frequency domain features, and fusing the double-branch features to generate a coding feature map; the coding feature map is decoded, in the decoding process, coordinate convolution CoordConv is adopted to enhance position perception, and mask segmentation and semantic label extraction are respectively carried out through a dynamic mask reconstruction branch DRMask Branch and an instance branch Inst Branch; based on a bipartite graph matching strategy, associating results of the mask segmentation and instance branches, and realizing segmentation of a single tree trunk and matching of instance-level labels; according to the method, the boundary precision and the detail reconstruction capability of trunk segmentation are remarkably improved, and the problem of feature loss of a traditional method in a complex under-forest environment is solved.
Owner:NANJING FORESTRY UNIV

Automatic driving environment sensing method and system based on multispectral fusion

The invention discloses an automatic driving environment sensing method and system based on multispectral fusion, and relates to the field of auxiliary driving, and the method comprises the steps: collecting environment data in a complex environment through a sensor; performing time-space synchronization and standardized preprocessing on the environment data to obtain preprocessed data; constructing a sensor confidence estimation network, and inputting the preprocessed data into the confidence estimation network; features are extracted and fused through a structure combining a shared encoder and a branch encoder, and a real-time confidence score vector of each sensor is output; a preset mapping function of confidence and weight is obtained, the confidence score vector is converted into a dynamic weight vector, and the sum of all components of the weight vector is 1; and integrating the sensing results of the sensors by adopting a weighted fusion strategy, verifying the sensing results of the high-weight sensors through a cross validation mechanism, and outputting final environment sensing data.
Owner:FAW JIEFANG AUTOMOTIVE CO

Plant protection field literature information batch structured extraction method based on AI large model

The invention discloses a plant protection field literature information batch structured extraction method based on an AI large model, and belongs to the technical field of computer data processing and artificial intelligence application. According to the method, a PDF format literature is converted into a Markdown format text through a PDF literature self-adaptive preprocessing module and a Markdown conversion module; then, a double-strategy self-adaptive AI extraction module is adopted, different modes are adopted for processing according to a text length threshold value, and JSON data are extracted in combination with a JSON format data restoration strategy; then, the CPU intensive conversion task and the I / O intensive AI analysis task are processed in parallel through a two-stage parallel scheduling module, and a structured database file is generated through aggregation according to a predefined mapping rule through a multi-dimensional aggregation export module. According to the method, end-to-end automation from literature acquisition to structured data output is realized, and the problem of low extraction efficiency of manual literature reading information is solved.
Owner:SANYA INSTITUTE OF NANJING AGRICULTURAL UNIVERSITY

Fabric cross-modal image-text retrieval method based on knowledge graph and storage medium

The invention relates to the technical field of knowledge maps, and particularly discloses a fabric cross-modal image-text retrieval method based on a knowledge map and a storage medium, and the method comprises the following steps: constructing a textile fabric structured knowledge map; performing semantic modeling on the structured knowledge graph of the textile fabric according to a TransR model; semantic structure extraction is carried out on the structured semantic model of the textile fabric according to the graph attention network; extracting deep semantic features of the textile fabric image; constructing an image and entity bidirectional contrast loss function according to the textile fabric entity semantic structure and the image deep semantic features; constructing a multi-task joint loss function according to the bidirectional comparison loss function and the structured semantic model; and training the spatial alignment matching degree of the entity semantic structure and the deep semantic features of the image according to a multi-task joint loss function to obtain a fabric cross-modal image-text retrieval result. The fabric cross-modal image-text retrieval method based on the knowledge graph can be suitable for fabric cross-modal image-text retrieval.
Owner:JIANGNAN UNIV

PDF (Portable Document Format) document structured extraction system based on multi-modal language model

The invention discloses a PDF (Portable Document Format) document structured extraction system based on a multi-modal language model, belongs to the technical field of document processing and optical character recognition, and aims at solving the technical problem of how to improve the existing OCR (Optical Character Recognition) technology to improve the analysis capability of a complex document structure and improve the recognition precision of handwritten forms and other non-standard fonts. According to the technical scheme, the system adopts a layered decoupling architecture and comprises an input layer, a preprocessing layer, a reasoning layer, an output layer and a monitoring and fault-tolerant module; wherein the output layer is used for multi-source data access and path management to realize a local file system or S3 cloud storage; the preprocessing layer is used for invalid document filtering and visual feature extraction; the reasoning layer is used for multi-modal model interaction and content processing; the output layer is used for outputting a content aggregation result; and the monitoring and fault-tolerant module is used for realizing real-time state monitoring, resource consumption analysis and exception handling.
Owner:JIANGSU HAIRUO INFORMATION TECHNOLOGY CO LTD

Method for detecting low-confidence small target in radar echo based on hybrid architecture

The invention belongs to the technical field of radar signal processing, and particularly relates to a low-confidence small target detection method in radar echoes based on a hybrid architecture, and the method comprises the steps: firstly carrying out the spectrum symmetric movement and dimension recombination of radar echo data, and generating five-dimensional tensors [B, T, C, H, W] containing time sequence features; then the tensor is input into a detection model formed by cascading a Hurglass 3D module and a YOLOv8 network, the Hurglass 3D module extracts multi-scale spatial-temporal features through a structure of three-dimensional convolution down-sampling, bottleneck layer and three-dimensional transposition convolution up-sampling, and feature fusion is achieved through jump connection; and finally, target detection is completed through a backbone network, a neck network and a decoupling detection head of the YOLOv8 network. According to the invention, through spatio-temporal feature combined extraction and small target feature enhancement, the detection accuracy and the positioning precision of the low signal-to-noise ratio small target in radar echoes are effectively improved.
Owner:ANHUI UNIV

A warning method for self-extubation behavior of ICU patients based on RGB video monitoring

The present invention discloses a method for early warning of ICU patient self-extubation behavior based on RGB video monitoring, which belongs to the field of computer vision behavior recognition. It includes constructing an ICU patient self-extubation behavior dataset, obtaining a patient self-extubation early warning classification model based on a neural network and based on optical flow behavior features in an ICU scenario based on the dataset, obtaining a video image texture feature classification result by the patient self-extubation early warning classification model in an ICU scenario based on a neural network, and obtaining a patient self-extubation optical flow behavior feature classification result by the patient self-extubation early warning classification model in an ICU scenario based on optical flow behavior features; the above two classification results are fused and used as input to build a dual-stream ICU scenario patient self-extubation early warning model, thereby realizing ICU patient self-extubation early warning work based on RGB monitoring video. The present invention solves the problem of incomplete feature extraction from a single-stream structure, and has the advantages of high real-time performance, high accuracy, full automation and contactlessness.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Detection method, device and equipment of liquid crystal display screen and storage medium

The invention provides a liquid crystal display screen detection method, device and equipment and a storage medium, and the method comprises the steps: carrying out the hierarchical scale feature extraction of an image collected by a liquid crystal display screen through a multistage object pyramid structure, and obtaining a multi-scale defect feature representation; encoding the multi-scale features into a unified high-dimensional space by adopting a comparison encoding algorithm, and calculating a feature vector distance relationship to obtain comparison encoding feature vectors; inputting the feature vector into a defect detection model, and performing task specific activation through a coordinate convolution positioning branch and a dynamic snakelike convolution class branch to obtain a decoupling detection result; and according to a decoupling detection result, identifying and classifying defect types through a detection head, and outputting a detection result. According to the method, defects of different scales are captured through multi-stage pyramid structure extraction, discriminative enhancement of feature representation is achieved through high-dimensional space mapping of a comparison coding algorithm, and positioning precision and classification accuracy are improved through decoupling processing of coordinate convolution and dynamic snakelike convolution.
Owner:JIANGXI SUPER SPEED TECHNOLOGY CO LTD

Document structure extraction and model training method and device, equipment and medium

The invention discloses a document structure extraction and model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: constructing a special training data set containing data of at least two document understanding tasks (including optical character recognition, layout analysis, text positioning, regional text extraction, image description and chart title generation), and a fine tuning data set for converting a document image into a machine-readable structured text format; constructing a multi-modal large model comprising a shape adaptive cutting module, a visual encoder, a visual token compression module, a modal connector and a language decoder; pre-training the model by using the special training data set to jointly learn various document understanding tasks; and performing fine tuning on the pre-trained model by using the fine tuning data set, and adapting to a document structure extraction task to obtain a document structure extraction model. By means of the technical scheme, efficient and accurate document structure extraction can be achieved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Power distribution equipment point cloud multi-target classification extraction method, system, equipment and medium

The invention discloses a power distribution equipment point cloud multi-target classification extraction method, system, equipment and medium, and belongs to the technical field of power distribution line three-dimensional visualization, and the method comprises the steps: obtaining three-dimensional laser point cloud data covering a power distribution line region, and carrying out the noise point elimination and spatial density resampling, and obtaining first point cloud data; extracting ground points in the first point cloud data, and constructing a digital elevation model; removing the ground points to obtain second point cloud data containing the power distribution equipment and the environment target; executing clustering processing, and extracting structural feature parameters of each cluster; inputting the extracted structural feature parameters into a pre-trained classification model to complete classification; and correcting a classification result, and outputting a power distribution equipment point cloud classification data set with a category label. The method has the beneficial effects that in the ground point identification and elimination stage, a sliding window area consistent with the distribution line layout characteristics is constructed, an initial ground point set is identified by using a height gradient trend, and a misjudgment point set with obvious terrain disturbance or local ground characteristics is screened out through joint verification of a fitting residual error and a ground structure template, so that the accuracy of the ground point identification and elimination is improved. And the stability of subsequent structure extraction and the precision of ground object separation are enhanced.
Owner:GUIZHOU POWER GRID CO LTD

High-proportion photovoltaic regional power distribution network state sensing method, system and device based on heterogeneous graph neural network and medium

The invention relates to the technical field of intelligent power grid operation monitoring and state recognition, and discloses a method, a system and equipment for sensing the state of a high-proportion photovoltaic regional power distribution network based on a heterogeneous graph neural network and a medium. The limitation of a traditional method in processing complex network topology, multi-source heterogeneous data and disturbance driving behaviors is broken through. In the state sensing process, from acquisition of multi-source data, a series of preprocessing, construction of a reasonable graph structure, extraction of node embedding vectors and output of state labels and abnormal scores are performed by using a preset state recognition model, and all links are closely connected and have innovativeness. And key nodes or regions with potential risks or abnormal states in the power distribution network are identified through the abnormal score function, and state online sensing and risk early warning of the power distribution network containing the photovoltaic region are realized.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD QUZHOU POWER SUPPLY CO

Multi-mode intelligent semantic understanding and abstract generation system and method based on HDMI (High Definition Multimedia Interface) stream

The invention provides a multi-mode intelligent semantic comprehension and abstract generation system and method based on an HDMI stream, and the system comprises an HDMI input module which is used for receiving a video signal outputted by external equipment; the image content streaming analysis module is configured to perform content segmentation, optical character recognition, layout structure extraction and graphic element detection on the video signals and output structured image data; the audio acquisition module is used for acquiring an audio signal and preprocessing the audio signal; the voice recognition module is used for transferring the preprocessed audio signal into a voice text; and the multi-modal fusion and semantic understanding module is configured to fuse the structured image data and the voice text, generate summary information and output the summary information to the result output module. According to the invention, real-time synchronous analysis of the multi-modal content and intelligent generation of the abstract are realized.
Owner:BEIJING YUNJIANXIN TECH CO LTD

Multi-modal phosphorylation site prediction method, device, equipment and medium

The invention relates to the technical field of biological information, and discloses a multi-mode phosphorylation site prediction method and device, equipment and a medium, and the method comprises the steps: extracting a subsequence of a candidate phosphorylation site from a to-be-predicted protein sequence, and generating a sequence feature of the subsequence through a deep learning model; constructing a graph structure based on the three-dimensional structure information of the candidate phosphorylation sites, and extracting structural features of the candidate sites through a graph neural network according to the graph structure; performing multi-modal feature fusion on the sequence features and the structural features to generate fusion features; and predicting the phosphorylation probability of the candidate site based on the fusion feature to obtain a phosphorylation probability prediction result. According to the method, the protein sequence language model and the multi-modal features of the three-dimensional structure diagram convolutional network are fused, so that the accuracy and comprehensiveness of phosphorylation site prediction can be effectively improved, and meanwhile, the calculation efficiency and the bioinformatics interpretation are considered.
Owner:SHENZHEN UNIV

Underwater DOA estimation method based on graph nerve and convolutional neural network

The invention relates to the field of underwater sound signal processing, in particular to an underwater DOA (direction of arrival) estimation method based on graph nerves and a convolutional neural network, which comprises the following steps: 1, establishing a linear array, and enabling narrow-band signals to simultaneously reach an underwater sound array; 2, performing signal preprocessing to obtain a signal covariance matrix, and performing normalization processing; 3, extracting correlation between array elements and spatial features of array signals, and performing data supplementation on sparse linear array information; 4, forming a double-branch structure, enhancing the information aggregation capability, and extracting features from a space path and a time domain path; and 5, constructing an adjacent matrix, filling node features of damaged array elements, adopting a double-branch structure, extracting spatial features and time domain features, carrying out feature integration, and outputting a DOA estimation result. The spatial correlation between array elements is extracted and the array sparsity problem is processed by using the graph neural network, and the time domain features of the signals are extracted in combination with the convolutional neural network, so that more accurate and more robust DOA estimation can be realized under the conditions of low signal-to-noise ratio and array sparsity.
Owner:QINGDAO UNIV OF SCI & TECH