Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

98 results about "Semi-structured data" patented technology

Semi-structured data is a form of structured data that does not obey the formal structure of data models associated with relational databases or other forms of data tables, but nonetheless contains tags or other markers to separate semantic elements and enforce hierarchies of records and fields within the data. Therefore, it is also known as self-describing structure.

High-performance AIOT real-time data processing method and device based on Flink, Kafka and Dores

The invention discloses a high-performance AIOT real-time data processing method and device based on Flink, Kafka and Dores, and the method comprises the steps: constructing an end-to-end hierarchical architecture, capturing multi-source database change data in real time through Flink CDC, and analyzing semi-structured data in combination with Apache NiFi, and generating a standardized data stream; the method comprises the following steps: realizing message routing based on KafkaTopic division of a service domain, ensuring the consistency of a data structure by using Schema Registry, and optimizing the transmission efficiency through a dynamic compression algorithm; an Flink real-time computing engine is adopted to implement window aggregation, asynchronous dimension table association and incremental state persistence strategies, and the resource fluctuation and state management problems in streaming computing are solved; time pre-partitioning and intelligent cold and hot layering are achieved through a Dores storage layer, the query performance is improved in combination with materialized views and multi-level indexes, and finally a full-link closed loop of collection-transmission-calculation-storage-management is formed. According to the invention, the problems that the real-time performance and the consistency are difficult to balance and the system complexity is high in the prior art are solved.
Owner:SHANGHAI QUZHI NETWORK TECH CO LTD

Hybrid expert multi-model task processing method and system based on AI Agent scene

The invention relates to a mixed expert multi-model task processing method and system based on an AI Agent scene. The method comprises the following steps: constructing a knowledge block set for unstructured data and semi-structured data; constructing an associated document vector set of topN similarity; user input and a reverse decoding document are spliced, and user tasks are divided by using LLM to generate subtask sets; constructing a task allocation data set # imgabs0 # based on the scoring model and the expert model pool; and the dependency aggregation # imgabs1 # for generating the subtask answers is used as a final user task result. According to the method and the system, fine-grained division and route selection of the model pool are carried out on user tasks, so that the accuracy, diversity and authenticity of answers required by users can be effectively improved, and the multi-model parallel utilization rate of an Agent system is enhanced.
Owner:FUJIAN TERTON SOFTWARE CO LTD

Personalized learning system based on multi-agent and retrieval enhancement generation

The invention discloses a personalized learning system based on multiple agents and retrieval enhancement generation. The system comprises a knowledge extraction agent, a retrieval enhancement generation module, a planning teaching agent, a knowledge consolidation agent and a test evaluation agent. The knowledge extraction agent extracts document information from the original learning materials and converts the document information into semi-structured data; the retrieval enhancement generation module is used for storing the semi-structured data into a vector library in a vector form, retrieving information in the vector library according to a request sent by a user, and inputting enhancement information combined by the retrieval information and the request into a downstream agent; and the planning teaching agent, the knowledge consolidation agent and the test evaluation agent are respectively used for constructing a knowledge graph according to the received information to generate a personalized teaching plan, generating exercise questions to consolidate learned knowledge, evaluating a learning effect and providing feedback suggestions. The learning efficiency, the learning effect and the learning experience of the learner can be remarkably improved, and personalized learning can be met.
Owner:SOUTH CHINA UNIV OF TECH

Resume matching method and device

The invention provides a resume matching method and device, and is suitable for the technical field of data processing. The method comprises the following steps: analyzing and mining resume unstructured data, resume semi-structured data, resume structured data, post demand unstructured data, post demand semi-structured data and post demand structured data to obtain resume implicit information and post demand implicit information; according to a resume matching calculation parameter, a resume matching parameter optimization step length, resume unstructured data, resume semi-structured data, resume structured data, resume implicit information, post demand unstructured data, post demand semi-structured data, post demand structured data and post demand implicit information; and generating resume matching information. The matching accuracy of the resumes and the post requirements is effectively improved, the resume matching reliability and rationality are enhanced, and therefore the recruitment efficiency and effectiveness are improved.
Owner:BEIJING ZHONGKE JIANYOU TECHNOLOGY CO LTD

Supplier behavior real-time sensing and monitoring method and device and electronic equipment

The invention relates to a supplier behavior real-time sensing and monitoring method and device and electronic equipment. The method comprises the steps of collecting and preprocessing supplier multi-source heterogeneous data in real time, dynamically capturing a Binlog of a MySQL database through a Flink CDC framework, obtaining service behavior data in real time, collecting text log data through a Flume framework, and transmitting network behavior data in real time. Integrating the structured data, the unstructured data and the semi-structured data, and constructing a unified data transmission channel; constructing a multi-dimensional user portrait; comparing the semantic similarity between the bidding file and the bidding requirement, and triggering compliance early warning when the matching degree is lower than a threshold value; analyzing a supplier cooperation network, and marking a cluster group with abnormal transaction frequency and irrelevant business fields as a cross-bidding risk; and analyzing the access behavior sequence, and detecting a malicious access pattern to generate a malicious access alarm. According to the supplier behavior monitoring system and method, real-time performance, accuracy and multidimensional performance of supplier behavior monitoring are achieved through the architecture of multi-source heterogeneous data real-time integration, multi-dimensional portrait construction and intelligent risk identification.
Owner:CHINA ACADEMY OF RAILWAY SCI CORP LTD +1

Large language model knowledge question-answering method and system fused with multi-modal knowledge graph

The invention relates to the technical field of knowledge questions and answers, in particular to a large language model knowledge question and answer method and system fused with a multi-modal knowledge graph. The method comprises the following steps: acquiring structured and semi-structured data and unstructured multi-modal data in an enterprise business scene; capturing data change through a real-time synchronization technology and executing targeted preprocessing to ensure the consistency of data quality and modal characteristics; constructing a dynamically updated multi-modal knowledge graph based on the preprocessed data, and integrating a multi-modal embedded model, a reordering model and a fine-tuned generative large model to obtain a knowledge storage-feature matching-semantic generation integrated architecture; according to the method, the defect of knowledge lag of a traditional system is overcome, and seamless fusion of the AI capability and the enterprise core business process is realized.
Owner:CPI INFORMATION TECH CO LTD +1

Multi-modal medical data fusion and analysis method

The invention discloses a multi-modal medical data fusion and analysis method, which is applied to the technical field of medical informatization and comprises the following steps: acquiring multi-modal medical data; wherein the multi-modal medical data comprises structured data, unstructured data and semi-structured data generated in the diagnosis and treatment process of the patient; structural data features are extracted by adopting an optimized Z-score standardization formula; using sparse attention CNN to extract unstructured data features; deepening diagnosis and treatment semantic association by adopting a Transform encoder and entity embedding fusion form, and extracting semi-structured data features; multi-modal data are fused through closed-loop design of coding dimension reduction integration, decoding reconstruction verification, exclusive loss optimization and feature standardization output, and a feature relevance verification mechanism is fused; and matching the fused multi-modal data with a data analysis task, and performing task analysis and mining. According to the invention, efficient and reliable technical support is provided for medical digital transformation and precise diagnosis and treatment.
Owner:CHINA JAPAN FRIENDSHIP HOSPITAL

Health risk grading evaluation method and system based on dynamic feature recognition

The invention provides a health risk grading evaluation method and system based on dynamic feature recognition, and relates to the technical field of data processing, and the method comprises the steps: 1, collecting multi-source heterogeneous data in real time through an information platform, and employing an automatic classification engine to analyze a data structure change to generate a feature grading tag; distributing the data to heterogeneous storage areas according to the feature classification labels, wherein the heterogeneous storage areas comprise a structured data area, a semi-structured data area and an unstructured data area; deploying a dynamic probe group in each type of storage partitions, and capturing a basic evolution index set in real time; and 2, based on the feature grading labels, matching encryption strategies for different grading data, dynamically adjusting a transmission protocol based on the basic evolution index set, and when the feature grading labels meet preset conditions, executing de-identification processing, and outputting a to-be-analyzed data stream. According to the method, through full-process dynamic treatment and closed-loop optimization, accurate recognition, hierarchical management and control and adaptive protection of health data risks are realized, and data security and utilization efficiency are both considered.
Owner:HANGZHOU JUNYUAN HEALTH TECHNOLOGY CO LTD

Low-code process automatic arrangement method and system

The invention discloses a low-code process automatic arrangement method and system, and particularly relates to the technical field of automatic arrangement, and the method comprises the steps: carrying out the unified collection and verification of an external system, resources and system boundaries before deployment, generating an operation environment fingerprint with a signature, and carrying out the strong binding of the fingerprint with a process instance; process DAG modeling is completed on the visual canvas, system and semantic double validity verification is carried out, and a post-curing parameter snapshot is used as a subsequent unique parameter source; in the online simulation period, observable measurement is collected, a deviation perception index is formed, and steady selection and switching of an execution kernel are achieved by combining hysteresis and a minimum residence strategy; in a production state, a canary release strategy, a degradation strategy and a rollback strategy are used for communicating simulation and real disk, and the SLA and the key KPI are continuously monitored; normalized collection and immutable archiving are carried out on structured and semi-structured data in the operation period, an evidence chain and a minimum reproduction experiment package are generated based on hash and signature, and the comparability and traceability of cross-version results are achieved.
Owner:HANGZHOU SANLIANG TECHNOLOGY CO LTD

Method and apparatus to create structured documents and generate content

A semantic diffusion model may generate semi-structured data using existing character image creations. Form image generation is one area of possible application. Embodiments include both training the diffusion model and using the diffusion model. The model can learn to permute and rearrange character features for different regions. Newly generated forms can be applied to train the semantic diffusion model to provide further improvements to the model's capability and generality. The model can generate high quality character-like images that incorporate geometric properties such as character locations and regions of similar meaning, which humans can check and / or interpret. Embodiments are suitable for semi-structured data such as forms, tables, and aligned keyword text generation, and resolve the issue of generating data for mixed and combined geometries and semantics. There also is applicability to hybrid or multimodality datasets, so long as the raw data can be interpreted and converted as character-like image tensors.
Owner:KONICA MINOLTA BUSINESS SOLUTIONS USA INC

Unified structured and semi-structured data types in database systems

The subject technology receives a first semi-structured object. The subject technology iterates through a list of fields specified by a target object type. The subject technology, for each field, determines whether a field with a same name is present in the first semi-structured object. The subject technology, in response to the field being found in the first semi-structured object, converts a value of the field to a target field type according to defined type conversion rules. The subject technology stores the converted value in a unified representation comprising a data structure that stores both structured and semi-structured data types. The subject technology processes a query using the unified representation.
Owner:SNOWFLAKE INC

Document processing method and device, equipment and storage medium

The invention discloses a document processing method and device, equipment and a storage medium, and relates to the technical field of data processing, the disclosed document processing method comprises the following steps: obtaining a to-be-processed document, and identifying a text region and an image region in the to-be-processed document; performing text extraction on the text area to obtain vector text content; performing character recognition on the image area through a plurality of optical character recognition models and obtaining a recognition result; determining a confidence coefficient weight in the recognition result according to the character recognition result corresponding to each optical character recognition model; fusing each character recognition result with the vector text content according to each confidence coefficient weight to obtain a fused text content; and outputting the structured data or the semi-structured data according to the fused text content, thereby being compatible with the input of the mixed type document, realizing the processing of the mixed type document containing the text region and the image region, and effectively meeting the user requirements.
Owner:QUANGDA (BEIJING) INFORMATION TECHNOLOGY RESEARCH INSTITUTE CO LTD

Zero sample learning-based multi-source bill call ticket data integration method, system and product

The invention discloses a zero sample learning multi-source bill call ticket data integration method, system and product, and the method comprises the following steps: importing multi-source bill call ticket data, and classifying the imported multi-source bill call ticket data according to the data format type, dividing the multi-source bill call ticket data into structured data, semi-structured data and unstructured data; extracting structured data features to obtain structured features so as to create a data field, and importing corresponding data; performing analysis and feature extraction on the semi-structured data to obtain semi-structured features, matching the semi-structured features with the structured features to form similar features, and importing corresponding data; extracting the unstructured data to obtain unstructured features, matching the unstructured features with the structured features or similar features, and importing corresponding data; and acquiring data in the data field according to feature matching. According to the invention, zero sample learning automatic data import of the multi-source bill call ticket can be realized.
Owner:SHANGHAI XINREN INFORMATION TECH CO LTD

Electric power material detection system based on improved association rule and knowledge graph

The invention provides an electric power material detection system based on an improved association rule and a knowledge graph, and belongs to the field of electric power material detection. According to the system, an association analysis model for equipment detection is established by using an association rule algorithm, a hidden association relationship between unqualified detection items is mined, and a judgment basis for the disqualification of the related items is provided. In addition, in combination with information such as a material and equipment detection report and expert knowledge, knowledge such as entities, attributes and relationships is extracted from semi-structured data by adopting a natural language processing technology, a material and equipment detection management knowledge graph is constructed in a graph database Neo4j, a systematic material detection knowledge base is provided, network visualization is carried out, and the material and equipment detection management knowledge graph is established. And information query and management are facilitated. According to the invention, the detection efficiency of the electric power material equipment is improved, and increasing detection requirements are met.
Owner:GUANGXI ZHUANG AUTONOMOUS REGION INFORMATION CENT (GUANGXI ZHUANG AUTONOMOUS REGION BIG DATA RES INST)

Efficient storage and querying of schema-less data

A method (300) of storing semi-structured data (12U) includes receiving user data (12) comprising semi-structured user data from a user (10) of a query system (150). The method includes receiving an indication (14) that the semi-structured user data fails to include a fixed schema. In response, the method further includes parsing the semi-structured user data into a plurality of data paths (210) and extracting a data type (220) associated with each respective data path of the plurality of data paths. The method additionally includes storing the semi-structured user data as a row entry in a table (204) of a database in communication with the query system, wherein each column value associated with the row entry corresponds to a respective one of the plurality of data paths and the data type associated with the respective data path.
Owner:GOOGLE LLC

A method for recommending labels and label instances

The present invention relates to a method for recommending labels and label instances, and belongs to the field of computer data processing technology. In response to the problems of label management and intelligent label management of massive test data, the present invention proposes a method for recommending labels and label instances. The method mainly includes: collecting structured, unstructured and semi-structured data, constructing ontology concepts according to the types of test data, and forming a label library; then, using different methods to extract equipment entities and entity relationships from multimodal data such as images, texts, audio, video, and paper, and constructing a label instance library; using methods such as rule mapping and natural language processing to map the relationship between labels and label instances; finally, mining user personal information and label usage information, combining personal information and label information, and forming intelligent recommendations based on labels and label instances.
Owner:NAT UNIV OF DEFENSE TECH

Intelligent pneumonia data collection method and device based on big data

The invention relates to a pneumonia data intelligent collection method based on big data. The method comprises the following steps: obtaining pneumonia-related data to obtain to-be-collected data; performing data preprocessing on the to-be-collected data to obtain processed data; classifying the processed data through a pre-trained CNN neural network, and dividing the processed data into structured data, unstructured data and semi-structured data; and storing the structured data in a relational database, and storing the unstructured data and the semi-structured data in a non-relational database. According to the pneumonia data intelligent collection method and device based on the big data, on one hand, the problem of high labor cost of manual processing is solved, and on the other hand, due to the fact that the CNN neural network model is adopted for classification, the problem that final summarized and collected data are disordered, and follow-up analysis is affected is solved.
Owner:MACAU UNIV OF SCI & TECH

Artificial intelligence-based cardiovascular chronic disease data management method

The application discloses a cardiovascular chronic disease data management method based on artificial intelligence, relates to the technical field of artificial intelligence medical management, and comprises the following steps: collecting patient multi-element heterogeneous medical data, normalizing and converting structured and semi-structured data, extracting key features of unstructured medical image data and vectorizing and encoding the key features, and generating a standardized cardiovascular chronic disease comprehensive data cube; calling a pre-trained multi-modal cardiovascular risk assessment model to analyze the data cube, and generating a comprehensive risk assessment report containing a risk level, a contribution factor, an evolution trend and an individualized early warning threshold; combining a clinical path knowledge base to generate an individualized health management plan, continuously collecting plan execution feedback data, dynamically updating the data cube, and iteratively optimizing the assessment report and the management plan. The method realizes multi-element heterogeneous data regularization, enriches risk assessment dimensions, realizes dynamic adaptation of a health management plan, and improves the standardization and individualization level of cardiovascular chronic disease data management.
Owner:FUJIAN PROVINCIAL HOSPITAL

Special material safety standard-oriented domain knowledge extraction method and system

The invention relates to the technical field of knowledge extraction, in particular to a domain knowledge extraction method and system for special material safety standards. The method comprises the following steps: collecting special material safety standard data, preprocessing the data, and manually classifying standard terms; carrying out ontology design; extracting structured and semi-structured data knowledge; extracting knowledge of the unstructured data; performing preference fine tuning and optimization on the entity type; the domain knowledge extraction system for the special material safety standard comprises a data collection module, a classification module, an ontology design module, a first data knowledge extraction module and a second data knowledge extraction and preference fine adjustment module. Through the mode, the corpus tagging efficiency of the professional field can be improved, the dependence on artificial resources is reduced, and the conditions that local deployment is difficult, the professional term understanding of the model professional field is insufficient, and the knowledge extraction ability is insufficient are relieved.
Owner:SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Data warehouse construction and application method based on multi-source data fusion

The invention discloses a data warehouse construction and application method based on multi-source data fusion, and relates to the technical field of data fusion, and the method comprises the steps: multi-source data integration layer construction: obtaining structured, unstructured and semi-structured data through a heterogeneous data collection module, carrying out the automatic preprocessing and standardization processing of the data, and obtaining a multi-source data integration layer; the standardization processing comprises a unified naming rule, a coding rule and a data type; designing a core model of the data warehouse: constructing a star model based on a business theme, the star model comprising at least one fact table and a plurality of associated dimension tables, the fact table storing business index data, and the dimension tables storing descriptive attribute data; implementing an ETL automatic process: configuring a data extraction and conversion rule and a scheduling strategy through a visual ETL tool, automatically loading the data processed by the integration layer to a star model, and monitoring a task state in real time; and intelligent query and application recommendation.
Owner:INSPUR SMART TECH INNOVATION (SHANDONG) CO LTD

Analysis system and analysis method for automatically detecting semi-structured data quality problems

The application discloses an analysis system and an analysis method for automatically detecting semi-structured data quality problems, the analysis method takes JSON data in semi-structured data as a specific research object, and comprises the following steps: a parsing mode module is used for parsing semi-structured data into an aggregated schema tree; a data quality problem monitoring module is used for automatically detecting potential data quality problems of data according to the aggregated schema tree and a data quality space; a visualization generation module is used for visualizing the aggregated schema tree and the data quality problems, helping users to interactively find and view the data quality problems; and a data cleaning module is used for user configuration and solution of the data quality problems. The system can help users to quickly and effectively locate and clean the data quality problems of semi-structured data.
Owner:ZHEJIANG UNIV

A patrol management system and method for long-distance pipelines

This invention relates to the field of pipeline inspection and management technology, specifically a long-distance pipeline inspection and management system and method, including a centralized control center, edge processing equipment, cameras, a cloud platform, and manual inspection terminals; the cameras are installed at designated locations for real-time monitoring of target objects. The advantages of this invention are: by introducing a data governance module and adding three-dimensional tags containing timestamps, location coordinates, and associated object IDs to multi-source data, and simultaneously establishing a core index master ledger and associated ledgers bound together by unique data IDs, this system achieves automatic association and fusion of unstructured, structured, and semi-structured data generated during inspections. This solves the "data silo" problem caused by inconsistent data formats and scattered storage in existing technologies, providing a complete and traceable data chain for each potential hazard or equipment status, supporting full-process traceability and analysis.
Owner:SOUTH CHINA BLUESKY AVIATION OIL & GAS CO LTD

Adaptive lda topic model training system based on public opinion real-time data stream

The application provides a self-adaptive LDA topic model training system based on public opinion real-time data flow, comprising a data gathering module, a data preprocessing module, a self-adaptive LDA model training module and an incremental LDA model fusion module; the data gathering module is used for extracting and converting and loading the structured and semi-structured data and inputting into a distributed message bus kafka; the data preprocessing module is used for preprocessing the data in the message bus kafka and finally forming a weighted word vector; the self-adaptive LDA model training module is used for training to obtain an LDA model result and merging the training result; the incremental LDA model fusion module is used for fusion training to generate a new round of LDA model. The application is superior to the traditional LDA topic analysis method in accuracy and performance, and is applied to network public opinion field event detection, recommendation, word cloud and retrieval and other practical engineering projects, and creates commercial value.
Owner:NANJING LES CYBERSECURITY & INFORMATION TECH RES INST CO LTD +1

Knowledge graph construction method, fault query method, electronic equipment and storage medium

The embodiment of the invention provides a knowledge graph construction method, a fault query method, electronic equipment and a storage medium. The knowledge graph construction method comprises the steps of obtaining a first knowledge graph of a target system, wherein the first knowledge graph is formed based on an association relationship between first entities of structured data and an association relationship between second entities of semi-structured data in the target system; extracting an association relationship between third entities corresponding to the non-structural data in the target system based on an extraction model, wherein the extraction model is an entity relationship extraction model trained based on the first entities, the association relationship between the first entities and the association relationship between the second entities; and updating the first knowledge graph based on the third entities and the association relationship between the third entities to obtain a second knowledge graph. According to the scheme of the embodiment, the knowledge graph is automatically constructed, the human input is reduced, and the manual annotation cost is saved; and a technical basis is provided for quickly positioning fault nodes in the system.
Owner:ZTE CORP

Column type partition storage method and device for variant type, equipment and medium

The embodiment of the invention discloses a column type partition storage method and device for variant types, equipment and a medium. A specific embodiment of the method comprises the following steps: performing grammar embedding processing on a variant type; performing data extraction processing on the mode template statement; generating mode template metadata; performing data analysis processing on the initial semi-structured data; performing data grading matching processing on the initial key value pair data; performing conversion verification processing on each unmatched key value pair data and each composite key value pair data; performing data screening processing on each composite conversion sub-column; performing partition coding processing on each unmatched sub-column, each target conversion sub-column and each to-be-processed conversion sub-column; and performing determinant partition storage on the dynamic sub-column region data, the sparse column region data and the typed sub-column region data. According to the embodiment, waste of storage resources can be reduced.
Owner:BEIJING FLYWHEEL DATA TECH CO LTD

Method and system for AI-based interactive searches

A system for interactive searches based on user queries data and a plurality of Large Language Models (LLMs) including a processor of a Human-Machine Interface (HMI) server node configured to host a network of LLMs and at least one machine learning module (ML) and connected to at least one user-entity node over a network and a memory on which are stored machine-readable instructions that when executed by the processor, cause the processor to: receive a search request input data from the at least one user-entity node; evaluate, by a first dedicated LLM, relevance of the search request input data by discerning between primary and secondary information; responsive to evaluation by the first dedicated LLM, derive classifying features from the primary and secondary information and generate a feature vector based on the classifying features; ingest the feature vector into the ML module configured to extract additional search parameters from a predictive search model based on historical search data associated with the at least one user-entity node; dissect, by a second dedicated LLM, the search request input data and the additional search parameters to separate the data into qualitative and quantitative criteria elements based on the primary and the secondary information; transform, by a third dedicated LLM, the quantitative criteria elements into structured queries for a database; process, by a fourth dedicated LLM, the qualitative criteria elements by searching through a semi-structured data repository; and synthesize, by the fifth LLM, processed search findings into a succinct human-language summary.
Owner:INGOLD DAN +2

Ship plate high-strength steel material composition design method based on multi-objective black box optimization

The application discloses a ship plate high-strength steel material composition design method based on multi-target black box optimization, relates to the field of metal material design, and collects ship plate high-strength steel historical data to obtain a standard semi-structured data set through alignment and cleaning treatment; based on the data set, optimization targets, decision variables and constraint conditions are determined, and a multi-target optimization model is constructed; a mechanical property prediction model integrating multiple performances is obtained through machine learning algorithm training; an optimal solution set is obtained through iterative calculation of a reference vector guided multi-target optimization algorithm, and a final composition design scheme is selected by combining target steel grade service scenarios and assigning weights, so that the complex mapping relationship of composition-process-performance is accurately quantified, the strength, toughness, weldability and corrosion resistance are simultaneously optimized to match service requirements, the optimization algorithm considers global exploration and local development, the optimal solution under multiple constraints is efficiently obtained, manual experience derivation is not needed, and the composition design can be popularized to other high-strength steels or special steels.
Owner:NORTHEASTERN UNIV CHINA

Knowledge graph construction method based on Bi-GRU relation extraction and vector variance algorithm

The invention relates to a knowledge graph construction method based on Bi-GRU relation extraction and a vector variance algorithm. The method comprises the following steps: constructing an initial knowledge graph hierarchical structure from structured knowledge; extracting an implicit classification relationship from the semi-structured open tag by adopting a co-word analysis algorithm; constructing a Bi-GRU remote supervision relation extraction model containing a multi-scale attention mechanism and an improved cross entropy loss function, and extracting a non-classification relation from an unstructured text; calculating node membership and variance through a vector variance algorithm, and determining a threshold to delete irrelevant nodes and associated edges in the field; the method aims at solving the technical problems that an existing domain knowledge graph depends on structured / semi-structured data, so that knowledge is incomplete, non-structured text relation extraction is interfered by noise, and domain-independent node redundancy exists in the graph.
Owner:GUIZHOU QIANCHENG HUITONG TECH DEV CO LTD

A review expert library random extraction evaluation method

The application discloses a kind of review expert database random extraction evaluation methods.The method relates to intelligent matching extraction technical field, including the following steps: multi-source data acquisition and vectorization processing, matching accuracy dynamic adjustment, data assetization closed loop transformation and regulation and expert random extraction execution.The application is by collecting review related data and document and carries out vectorization processing, it is to unstructured, semi-structured data into structured vector data, obtains expert and review matter matching accuracy, and judges whether to carry out dynamic adjustment, then carries out expert and review project correlation analysis.Based on correlation analysis acquisition data assetization closed loop transformation degree, and judges whether to carry out closed loop adjustment, finally completes expert random extraction, realizes review matching precision, data transformation efficiency, extraction process is stable and reliable, improve a kind of review expert database random extraction evaluation reliability, solve the low reliability of a kind of review expert database random extraction evaluation in prior art.
Owner:BOWENDE (BEIJING) TECHNOLOGY CO LTD +1

Named entity recognition method for high-speed railway technical improvement and overhaul project text

The invention provides a named entity recognition method for a high-speed railway technical improvement and overhaul project text, which comprises the following steps of: preprocessing unstructured data and semi-structured data in a picture and / or a text, and converting character-level positions of entity labels by adopting BIOES (Basic Input / Output Element Specification) labeling to obtain a technical improvement and overhaul data set; performing data enhancement on the technical improvement and overhaul data set by adopting a data enhancement strategy to obtain a data enhancement result; and inputting the text data subjected to data enhancement into an RBC fusion model, converting the text data into character vectors, carrying out a sequence labeling task to obtain a prediction label result of each character, and finally realizing an entity of a technical renovation and overhaul text through a large model prompt project method. Through preprocessing, data enhancement and an RBC fusion model, non-structured and semi-structured data in the high-speed rail technical improvement and overhaul project are efficiently identified, the problem that key information cannot be dynamically extracted through a traditional method is solved, and the method has the advantage that the information extraction efficiency and accuracy of the high-speed rail technical improvement and overhaul project are improved.
Owner:BEIJING TECH & BUSINESS UNIV +3