Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

70 results about "Text mining" patented technology

Text mining, also referred to as text data mining, roughly equivalent to text analytics, is the process of deriving high-quality information from text. High-quality information is typically derived through the devising of patterns and trends through means such as statistical pattern learning. Text mining usually involves the process of structuring the input text (usually parsing, along with the addition of some derived linguistic features and the removal of others, and subsequent insertion into a database), deriving patterns within the structured data, and finally evaluation and interpretation of the output. 'High quality' in text mining usually refers to some combination of relevance, novelty, and interest. Typical text mining tasks include text categorization, text clustering, concept/entity extraction, production of granular taxonomies, sentiment analysis, document summarization, and entity relation modeling (i.e., learning relations between named entities).

Multi-document key phrase extraction method based on graph structure node influence

The invention provides a multi-document key phrase extraction method based on graph structure node influence, and relates to the technical field of natural language processing and text mining. Firstly, a candidate phrase set is generated through noun phrase extraction and standardization; secondly, a semantic relation between phrases is captured through local subgraph construction and a sliding window mechanism, and the semantic relation is integrated into a global phrase co-occurrence graph; thirdly, dynamically dividing theme communities based on two-dimensional structure entropy minimization and a potential game model, and identifying phrase groups with high semantic aggregation; then, cross-topic nodes are processed through a structure entropy heuristic function, and flexibility of topic division is enhanced; and finally, in combination with node influence sorting, extracting key phrases with theme representativeness and propagation capability. The method does not need to label data, is suitable for multiple fields of academic literatures, news texts and the like, has high efficiency, accuracy and universality, and provides an innovative solution for multi-document key phrase extraction.
Owner:YUNNAN POWER GRID CO LTD +1

Dialogue text mining method and system for code scanning consumption cash return based on big data

The invention relates to the technical field of artificial intelligence, and discloses a dialogue text mining method and system based on code scanning consumption cash back of big data, and the system comprises a consumption data collection module, a multi-modal analysis module, a strategy decision engine and a dynamic execution module. When code scanning consumption cash-back strategy decision making is carried out, real-time dialogue text semantics and consumption time-space behavior data are fused, a deep association mechanism of consumption intentions and payment scenes is established, the matching deviation of user language expression and actual consumption behaviors can be eliminated in real time, the consistency of a cash-back strategy and real consumption appeals is ensured, and the payment efficiency is improved. The problem of strategy mismatching caused by semantic understanding isolation in a traditional system is effectively avoided, the precision of personalized marketing is improved, double feedback of dialogue emotion tendency and historical strategy effects is dynamically analyzed, the deviation value of strategy feature matching is corrected in real time, and the system can actively adjust weight distribution of feature combinations in a strategy library.
Owner:YIBIN DIGITAL CHAIN INTELLIGENT TECH CO LTD

Telecommunication fraud prevention-oriented record text mining method

The invention relates to a record text mining method for telecommunication fraud prevention, which belongs to the field of clue mining, and comprises the following steps: inputting telecommunication fraud records formed in a certain region within a period of time into a KnowLM-13b-i e model, extracting 9 types of entities and relationships, and generating structured data; importing the generated structured data into a Neo4j database, defining three layers of entities of "victim-attribute-structuring" and a corresponding relationship, performing structured analysis and centrality calculation, and further constructing a knowledge graph; initializing the constructed knowledge graph by using a reasoning technology, generating a data set and optimizing a loss function, performing model training on a reasoning model by using the data set, performing multi-task fine tuning in combination with the reasoning model, and enhancing the fraud recognition capability; and inputting a case description, and outputting a fraud type and a subsequent fraud means by the reasoning model. According to the method, the electric fraud record text can be combined with the pre-trained large model to realize knowledge reasoning of the electric fraud record.
Owner:CHINESE PEOPLE'S PUBLIC SECURITY UNIVERSITY

Method and system of converting unstructured digital documents to a structure format using a secure API

In one aspect, a computerized method for document extraction workflow for unstructured documents includes the steps of implementing a text mining operation on a set of digital documents the incoming documents. This is done by defining a document type of each digital document. Based on the document type, the method defines a set of data dictionaries to extract any data from each digital document. The method uses the defined set of data dictionaries to extract any data from each digital document.
Owner:YERRAMSETTY VENKATA SAI RAMAN +1

Sensitive data identification method and apparatus, device, and computer storage medium

The present application discloses a sensitive data identification method and apparatus, a device, and a computer storage medium. A text mining technology is used to mine a plurality of sensitive data rules from a data security specification file of a target industry to form a sensitive data rule base, the rule base is continuously augmented by using technologies such as NLP and NER, and after data to be identified of the target industry is obtained, a sensitivity class and a sensitivity level of the data to be identified can be identified by matching the sensitive data rules in the sensitive data rule base corresponding to the target industry with the data to be identified.
Owner:CHINA UNIONPAY

Adaptive discovery and mixed-variable optimization of next generation synthesizable microelectronic materials

This invention relates to systems and methods for adaptive discovery and mixed-variable optimization of synthesizable microelectronic materials, and applications of the same. Specifically, an exemplary system includes a virtual screening (VS) module to extract information from literatures of a knowledge base by text mining, a ML-assisted conceptual exploration (CE) module to identify candidate material families for the specific class of compound materials based on the extracted information via a combination of ML models and to generate exogenous models of objective functions f(x, y) and constraint functions g(x, y), and an adaptive discovery (AD) engine to generate and optimize design of the newly discovered compound materials. The AD engine includes a mixed-variable ML module, a mixed-integer optimization (MIO) module, and a high-fidelity evaluation (HFE) module, which are iteratively and sequentially executed.
Owner:NORTHWESTERN UNIV +1

Method for determining importance degree of seepage monitoring points of earth and rockfill dam and application of method

PendingCN121144726AMonitoring siteText mining
The invention discloses a method for determining the importance degree of seepage monitoring points of an earth and rockfill dam and application of the method. The method comprises the following steps that six seepage safety parts of the earth and rockfill dam are recognized, wherein the six seepage safety parts comprise a dam body, an anti-seepage body area, an inverted filtration drainage area, a dam penetrating building area, a dam foundation and a bank slope area; constructing an earth and rockfill dam seepage Bayesian network according to a historical monitoring text, and calculating importance degrees of six seepage safety parts; acquiring three-dimensional space coordinates of the seepage monitoring points: stake numbers, wheelbases and elevations; calculating the comprehensive importance degree of each part to each section of the monitoring point; and superposing the comprehensive importance degrees of the cross sections in the transverse, longitudinal and horizontal directions to obtain the importance degrees of the monitoring points. According to the method, historical monitoring text mining, a Bayesian network probability model and a three-dimensional space attenuation algorithm are coupled, and the importance degree of a quantitative measuring point is calculated, so that the problem that a traditional method depends on experience and neglects local structure defects is solved; and the high-weight section automatically triggers special inspection of the corresponding seepage safety part, so that timely further disposal is facilitated, and the safe operation life of the dam is effectively prolonged.
Owner:NANJING COLLEGE OF CHEM TECH +1

Graph model-based state grid document collaborative clustering analysis method, system and equipment and medium

The invention relates to the technical field of text mining and data clustering, and discloses a state grid document collaborative clustering analysis method, system and device based on a graph model and a medium, comprising: constructing a joint graph model, the joint graph model comprising a document node subset and a word node subset, the edge weight between document nodes being determined based on the diversity between documents, and the word node subset being determined based on the difference between the document nodes; the edge weight between the word nodes is determined based on the diversity between the words, and the document nodes and the word nodes are not directly connected; on the basis of the joint graph model, constructing and solving a collaborative clustering target function to synchronously perform graph division on the document node subset and the word node subset so as to obtain a clustering result of the document and a clustering result of the word at the same time; wherein the collaborative clustering objective function is configured to maximize the sum of the cut value of the document sub-graph division and the cut value of the word sub-graph division. According to the method, the complex dissimilar relationship between the documents and between the words can be visually and effectively captured by utilizing the natural relationship representation capability of the graph model.
Owner:GUANGAN POWER SUPPLY COMPANY STATE GRID SICHUANELECTRIC POWER

Sensitive data identification method and apparatus, device, and computer storage medium

The present application discloses a sensitive data identification method and apparatus, a device, and a computer storage medium. A text mining technology is used to mine a plurality of sensitive data rules from a data security specification file of a target industry to form a sensitive data rule base, the rule base is continuously augmented by using technologies such as NLP and NER, and after data to be identified of the target industry is obtained, a sensitivity class and a sensitivity level of the data to be identified can be identified by matching the sensitive data rules in the sensitive data rule base corresponding to the target industry with the data to be identified.
Owner:CHINA UNIONPAY

Automatic identification, classification and development trend analysis method of net red villages based on multi-source data fusion and natural language processing

The method for automatic identification, classification and development trend analysis of net red villages based on multi-source data fusion and natural language processing comprises the following steps: UGC data is crawled from Xiaohongshu and Douyin through a distributed master-slave architecture, de-duplicated based on SimHash, and normalized in time and coding format; a text semantic fingerprint is generated, and multi-level semantic cache fingerprint matching is performed; for unassigned text, its complexity is calculated, and a large language model API is adaptively called to automatically complete and extract five-level administrative divisions; weights are determined based on the analytic hierarchy process, interaction indicators such as likes, comments, collections and forwards are integrated, and a comprehensive network heat index of the village is obtained; an external text mining tool is connected, and batch word frequency analysis, semantic network analysis and sentiment tendency evaluation are performed; a document-term matrix is constructed, TF-IDF weighting is performed, and unsupervised clustering algorithm is used for clustering analysis of village characteristics; cross-dimension analysis is performed on the clustering results, and a development portrait, advantage mining and operation suggestion warning are automatically generated in combination with the SWOT model.
Owner:ZHEJIANG UNIV OF TECH

Automated method for virtual technical assistance in the correction of computer vulnerabilities through combined usage of software automation and artificial intelligence technics

The invention relates to an automatic method for technical assistance in the correction of vulnerabilities of a computer system, where the aforementioned method comprises at least the following steps: a) receiving information relative to computer system vulnerabilities by text, voice, file input or through data exchange; b) launching tools based on artificial intelligence algorithms and machine learning to understand the information and requests entered relative to vulnerabilities; c) performing parsing and text mining of the information received and process it through intelligent data matching with information relative to the computer vulnerabilities present in a database or through interaction with external online vulnerability databases; d) in case of unsatisfactory results, starting an extended online search by web scraping using deep learning algorithms and unsupervised machine learning; e) updating an internal vulnerability database with the additional information found; and f) generating a vulnerability report accompanied by relative remediation.
Owner:CYLOCK SRL

Urban land utilization identification method based on BERT model text classification algorithm

The invention relates to an urban land utilization recognition method based on a BERT model text classification algorithm, and the method is characterized in that the method comprises the following steps: 1, determining an urban land utilization recognition type, determining an urban land utilization recognition region range, obtaining POI data in the region range, and carrying out the data screening; step 2, establishing an urban land utilization identification unit in an urban land utilization identification model based on a BERT model text classification algorithm, and associating the screened POI data with the urban land utilization identification unit; step 3, carrying out geographic text mining on the POI data; and step 4, performing urban land utilization identification of the urban land utilization identification model based on the BERT model text classification algorithm to obtain a high-precision urban land utilization identification result. According to the method, high-precision urban land utilization identification can be timely and accurately carried out by utilizing the available POI data.
Owner:TIANJIN UNIV

Knowledge graph construction method and device for intellectual property retrieval and storage medium

The invention provides an intellectual property retrieval-oriented knowledge graph construction method and device and a storage medium, and the method comprises the steps: reading a target intellectual property text data set and a parallel corpus and a reference associated text in the target intellectual property text data set, mining synonymous mapping clues, reference traceability clues and technical theme associated clues implied in the texts, and constructing the intellectual property retrieval-oriented knowledge graph by the synonymous mapping clues, the reference traceability clues and the technical theme associated clues; sorting to obtain a weak supervision signal set, extracting synonymous expression pairs in parallel corpora and semantic association pairs in a reference association text, carrying out cross validation and duplicate removal to obtain a text alignment reference set, and carrying out global association matching on the text alignment reference set and the target intellectual property text data set to obtain an initial structured knowledge unit set; the method comprises the steps of obtaining a standardized knowledge unit set, performing credibility regularization to obtain a standardized knowledge unit set, performing entity classification and relation association organization according to hierarchical requirements of intellectual property retrieval, and constructing to obtain the intellectual property retrieval-oriented knowledge graph. The intellectual property retrieval accuracy and response efficiency can be effectively improved through the intellectual property retrieval method and device.
Owner:HENAN UNIV OF ANIMAL HUSBANDRY & ECONOMY

Text mining-based aviation accident cause intelligent identification method

The invention provides an aviation accident cause intelligent identification method based on text mining, relates to the technical field of aviation accident cause identification, and aims to improve the accuracy of aviation accident cause identification. The method comprises the following steps: acquiring aviation accident report data; extracting a plurality of candidate accident cause factors from the aviation accident report data; determining a confidence coefficient between any two candidate accident cause factors in the plurality of candidate accident cause factors; determining a direct influence matrix based on the confidence degree between any two candidate accident cause factors in the plurality of candidate accident cause factors; based on the direct influence matrix, measuring parameters of all candidate accident cause factors are determined, and the measuring parameters comprise centrality and / or cause degree; the centrality is used for representing the cause importance degree of the candidate accident cause factors; the cause degrees are used for representing cause guiding degrees of the candidate accident cause factors; and determining a target accident cause factor from all the candidate accident cause factors based on the centrality and the cause degree.
Owner:CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD +1

A college scientific research hotspot mining method and system based on improved BERTopic

PendingCN122286698AText miningEngineering
This invention relates to a method and system for mining research hotspots in universities based on an improved BERTopic, belonging to the field of text mining technology. It includes the following steps: Step S1, obtaining a list of effective word segments from university research papers; Step S2, capturing the core semantic information of academic texts; Step S3, obtaining a 3D low-dimensional vector that retains the core semantic features; Step S4, identifying potential research hotspot topic clusters; Step S5, generating a set of university research hotspot topics containing core keywords, weights, temporal attributes, and dual-dimensional labels of "topic + keyword"; Step S6, outputting multi-dimensional visualization results and hotspot prediction results, iteratively adjusting model parameters based on a feedback optimization mechanism, and finally generating a structured report and research decision-making suggestions. This application has the effect of improving the quality of mining research hotspots in universities.
Owner:NANTONG UNIV

An interactive self-service analysis retrieval system and method based on a large model

PendingCN122432203AText miningData retrieval
The application discloses an interactive self-service analysis and retrieval system based on a large model, which comprises a data source management module, a retrieval library construction module, an algorithm development module, a task scheduling management module and a service publishing module; the data source management module pre-processes initial information data to obtain structured information data; the retrieval library construction module constructs a retrieval library according to the structured information data; the algorithm development module extracts data features in the structured information data and recommends an adaptive retrieval algorithm according to the data features; the task scheduling management module performs retrieval analysis in the retrieval library according to the adaptive retrieval algorithm and an input retrieval instruction; and the service publishing module is used for converting the retrieval analysis result into an API service that can be called. The application aims to solve the pain points of the existing data analysis technology, such as fragmented functions, high operation threshold and lack of special text processing modules. The application provides an efficient solution for data modeling, text mining and service deployment.
Owner:NAVAL UNIV OF ENG PLA

Sensitive data detection method and device, electronic equipment and storage medium

The invention discloses a sensitive data detection method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining text data; performing word segmentation processing on the text data to obtain target text data; performing feature extraction on the target text data by using a first model to obtain text features; performing semantic analysis on the text features to obtain target text features; and inputting the target text features into a second model and a third model, and outputting sensitive data. According to the technical scheme, sensitive data features can be deeply mined, the method has higher adaptability to complex and changeable data forms, the recognition accuracy is greatly improved, and the probability of misjudgment and missed judgment is reduced. And deep fusion of natural language processing and text mining enhances semantic comprehension and further guarantees recognition precision.
Owner:AGRICULTURAL BANK OF CHINA

Recommendation method and system based on large language model and time series link prediction

The application discloses a recommendation method and system based on a large language model and time sequence link prediction, and the method comprises the following steps: obtaining relevant patent data of an industry field where a target company is located and inputting the large language model after training, and outputting patent division results of each stage of the industry field patent industry chain; adopting a sliding window strategy to construct a time sequence network from patent data of each stage of the industry chain, and extracting time sequence features; inputting the extracted time sequence features into a CNN-GRU model with a multi-head attention mechanism optimized by a Lingri optimization algorithm, outputting a prediction result, and performing potential technical fusion opportunity screening and verification; and recommending a potential partner of a next stage according to social influence and technical focus degree of the selected technical fusion opportunity. The application aims to accurately identify and recommend potential technical partners from the time dimension and dynamic change by combining deep learning text mining, a knowledge graph, a multi-dimensional adjacent attribute network and other methods.
Owner:WUHAN UNIV

Text mining and network analysis-based carbon and pollution reduction measure collaboration evaluation method, system and equipment and medium

The invention discloses a carbon reduction and pollution reduction measure collaboration evaluation method, system and equipment based on text mining and network analysis and a medium, and belongs to the technical field of climate and environmental governance measure analysis and evaluation.The method comprises the steps that carbon reduction and pollution reduction related measure files are obtained, and a personalized external corpus is constructed; performing duplicate removal and word segmentation on the text data; initially selecting keywords based on a TF-IDF method, and checking and summarizing the keywords according to a keyword three-level classification system; calculating the co-occurrence frequency between every two keywords, and constructing a word co-occurrence network; and analyzing the structure characteristics of the co-word network from different levels by using indexes such as network average degree, edge density, average shortest path and the like, and summarizing the collaboration of carbon and pollution reduction measures. Based on real and reliable first-hand data and a more targeted personalized corpus, keywords are accurately extracted and classified for extraction and classification, and in combination with a network analysis method, internal collaboration of carbon and pollution reduction measures is analyzed.
Owner:BEIJING INST OF TECH

Student personalized skill formation evaluation method and system based on text mining

The present application provides a kind of student individualized skill formation evaluation method and system based on text mining, the method includes step 1: obtaining the original text of formative evaluation homework, and it is preprocessed, obtain the homework text after preprocessing;Step 2: the homework text after preprocessing is classified, and the homework category to which homework text belongs is obtained;Step 3: for the homework category to which homework text belongs, analysis engine follows corresponding analysis rule set to analyze homework text;Step 4: output evaluation report.The present application can formative evaluation on student's experimental report, project homework and other formative evaluation homework such as this kind of unstructured, involves professional field knowledge, comprehensive and engineering text homework, to identify the professional skill defects of student in technical thinking, problem solving and professional expression ability and other aspects embodied in formative evaluation homework.
Owner:BEIJING VOCATIONAL COLLEGE OF ECONOMICS & MANAGEMENT (BEIJING MANAGER COLLEGE)

Chemical accident risk analysis and causal network construction method based on text mining and GNN-Apriori association rule

The invention belongs to the field of industrial safety and artificial intelligence, and discloses a chemical accident risk analysis method and system based on text mining and GNN-Apriori association rules. The method comprises the following steps: cleaning and structuring a chemical accident text, fusing TextRank, BM25 and BERT algorithms, and extracting a high-quality keyword set through a weighted model; constructing a graph structure based on a keyword co-occurrence relationship, and learning node semantic embedding by using GNN; inputting the embedded vector into an improved Apriori algorithm, and mining a semantic enhanced association rule in combination with semantic similarity, support degree, confidence and lifting degree; and finally, constructing a chemical accident causal complex network, and evaluating key nodes by using node centrality by taking the improvement degree as an edge weight to realize risk propagation path analysis and visual display. According to the method, chemical accident causes can be intelligently identified, semantic association is mined, a causal network is modeled, and interpretable decision support is provided for accident prevention and the like.
Owner:JILIN INST OF CHEM TECH

A traditional Chinese medicine knowledge graph construction method and system, and a storage medium

The embodiment of the application discloses a traditional Chinese medicine knowledge graph construction method and system and a storage medium. The method comprises the following steps: performing regular extraction on the obtained original text to obtain initial entity and relationship data converted into traditional Chinese medicine, and performing auditing on part of the data to obtain text training data; training based on the text training data to obtain a text mining model; then, the remaining part of the initial entity and relationship data is sent into the text mining model for prediction to obtain predicted entity and relationship data; a Neo4j graph database is used, and corresponding traditional Chinese medicine knowledge graphs are generated according to the predicted entity and relationship data; the effect is that the relationship and attribute between each entity are embodied; and therefore, the application of the user is more convenient, and the defects that the technology is too professional and there is no association between Chinese and Western medical terms in the prior art are overcome.
Owner:HANGZHOU PULSE HEALTH TECH CO LTD

Urban highway tunnel operation period toughness safety evaluation method

The invention discloses an urban highway tunnel operation period toughness safety evaluation method, and relates to the field of traffic infrastructure safety evaluation, and the method comprises the steps: constructing a standardized toughness safety database of an urban highway tunnel operation period; finely adjusting the large language model based on domain ontology knowledge to form a semantic analysis tool; building a knowledge graph through entity relationship mining; establishing a multi-scene index system in combination with a text mining technology; and dynamic weight calculation and safety scoring are realized by using an analytic hierarchy process. The method has the advantages that the objectivity of an evaluation index system is improved, the multi-scene adaptive capacity is enhanced, and the operation process is simplified.
Owner:TONGJI UNIV +1

Text mining-based ship fire risk factor identification method

The invention discloses a ship fire risk factor identification method based on text mining. The ship fire risk factor identification method comprises the following steps: 1) obtaining ship fire risk text data and establishing a text database; 2) performing word segmentation on the ship fire risk text data to obtain a ship fire risk word segmentation feature item list; 3) obtaining a ship fire risk keyword list; 4) extracting related phrases, and constructing a keyword related phrase set; 5) clustering the keyword related phrase set, and performing semantic analysis on a clustering result to obtain a ship fire risk factor list; 6) calculating the occurrence probability of the risk factors, and obtaining a ship fire risk factor probability table; according to the ship fire hazard risk factor risk assessment method, the importance of the risk factors and the fire hazard scene is quantitatively assessed through the text mining technology and the semantic analysis model.
Owner:WUHAN UNIV OF TECH

A career assessment method and system based on recruitment big data

The application discloses a kind of career assessment method and system based on recruitment big data, comprising: collecting relevant recruitment information;The recruitment information is carried out data preprocessing, obtains the data set corresponding to the recruitment information;Text mining is carried out to the data set, obtains the keyword that has influence on average salary and the average salary corresponding to the keyword;The information of the job applicant is substituted into the Lasso regression model for salary prediction constructed in advance, and the salary expected value of the job applicant is obtained;According to the keyword and the regression coefficient of the Lasso regression model, the career assessment scheme of the job applicant is determined;The career assessment scheme is used to carry out career assessment to the job applicant. Solve the problem that the job seeker cannot reasonably position oneself and cannot understand market demand in time.
Owner:AISINO CORPORATION

Threat entity relationship identification method and device in threat intelligence, equipment and medium

The embodiment of the invention provides a threat entity relation identification method and device in threat intelligence, equipment and a medium, and the method comprises the steps: carrying out the feature extraction of the threat intelligence according to a text feature extraction model, and obtaining a text feature; obtaining a domain knowledge graph according to the threat intelligence; processing the domain knowledge graph according to a graph feature extraction model to obtain text structure information features; fusing the text features and the text structure information features to obtain fused features; processing the fusion features according to a multilayer feedback neural network to obtain a threat intelligence threat entity expression matrix; and obtaining a threat entity relationship of the threat intelligence according to the threat entity expression matrix of the threat intelligence. The threat intelligence document is subjected to refined splitting, fine-grained recognition of threat intelligence is achieved, and the problem that threat entity relation recognition is not accurate due to sentence-level text mining in the prior art is solved.
Owner:孙雄韬 +1

Catering take-out consumption information prediction method based on large language model and text mining

The invention provides a catering take-out consumption information prediction method based on a large language model and text mining. The method comprises the steps of obtaining catering take-out consumption data, social economic data and environmental data of a target area in a target historical time period; based on a text mining model, performing vectorization processing on the catering take-out consumption data to determine a catering take-out consumption preference vector corresponding to the catering take-out consumption data; on the basis of a large language model, according to the social economic data and the environmental data, predicting the catering take-out consumption amount of the target area; and based on a catering take-out consumption information prediction model, according to the catering take-out consumption data, the catering take-out consumption preference vector and the catering take-out consumption amount, determining catering take-out consumption information of the to-be-predicted area. According to the invention, the prediction accuracy and fineness of the catering take-out consumption information can be improved.
Owner:GUANGDONG UNIV OF TECH

Semantic network-based document relationship analysis apparatus and method

The present invention relates to a semantic network-based document relationship analysis apparatus that analyzes a relationship between documents by utilizing artificial intelligence and natural language processing so as to identify an association between the documents. The semantic network-based document relationship analysis apparatus comprises: a document data reception unit that receives a plurality of document data from a user terminal; a document data analysis unit that analyzes text content of each document data by using a natural language processing technique and a text mining method, and extracts main concepts and keywords of the document data; a semantic network generation unit that analyzes the main concepts and keywords extracted from each document data, and generates a semantic network between the corresponding document data when the main concepts and keywords are determined to have a synonym relationship, a hierarchical relationship, and a near-synonym relationship; and a result data generation unit that generates result data in which the relationship between the document data connected by the generated semantic network is reflected in an image format, and displays the result data on the user terminal, wherein a user may visually check connection relationships according to the main concepts and keywords between the document data by means of the result data, and may explore document data related to the main concepts and keywords of the corresponding document data by selecting specific document data via the user terminal.
Owner:ALLBIGDAT INC

Text mining-based air-riding service case semantic extraction and network construction method

The invention belongs to the technical field of natural language processing, particularly relates to a text mining-based air-taking service case semantic extraction and network construction method, and aims to solve the problem of how to intelligently extract effective information of air-taking service cases. The method comprises the following steps: acquiring an air-ride service case text, and preprocessing the text to form case text data; analyzing a syntactic structure relationship among words in the case text data by utilizing a HanLP algorithm, and screening out a semantic core structure relationship; taking part-of-speech and word frequency information into consideration, improving a feature extraction algorithm, and extracting keyword items; based on adjacent positions of words in the feature data, analyzing to obtain directed association features among the words, and obtaining directed association phrases; constructing a co-occurrence network based on the directed association features and the directed association phrases, and performing modular clustering analysis; and key node links and high-influence node groups are extracted from the co-occurrence network, so that efficient and accurate extraction of service case work key points and core elements is realized.
Owner:CHINA EASTERN TECH APPL RES & DEV CENT CO LTD