Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

113 results about "Text mining" patented technology

Text mining, also referred to as text data mining, roughly equivalent to text analytics, is the process of deriving high-quality information from text. High-quality information is typically derived through the devising of patterns and trends through means such as statistical pattern learning. Text mining usually involves the process of structuring the input text (usually parsing, along with the addition of some derived linguistic features and the removal of others, and subsequent insertion into a database), deriving patterns within the structured data, and finally evaluation and interpretation of the output. 'High quality' in text mining usually refers to some combination of relevance, novelty, and interest. Typical text mining tasks include text categorization, text clustering, concept/entity extraction, production of granular taxonomies, sentiment analysis, document summarization, and entity relation modeling (i.e., learning relations between named entities).

Risk assessment model based on artificial intelligence in financial big data analysis

The invention relates to the field of financial science and technology, and discloses a financial risk dynamic assessment system and method based on artificial intelligence. The system comprises a multi-source heterogeneous data acquisition module which acquires transaction data, public opinion texts and association maps in real time; the adaptive feature engineering module dynamically screens key risk factors; the dynamic risk map construction module calculates a risk conduction coefficient through a map neural network; the multi-modal AI analysis engine cooperatively runs a time sequence prediction model, a text mining model and a graph calculation model; a risk conduction simulator quantifies a systematic risk path. The problems of data splitting processing, model static solidification and correlation risk quantification deficiency in the prior art are solved, the false alarm rate is reduced to 12%, the response speed reaches 90 seconds, the prediction deviation is reduced to 22%, and an interpretable supervision report is generated.
Owner:BEIJING CREDIT MANAGEMENT CO LTD

User demand comprehensive analysis method based on online user comment data

The invention discloses a user demand comprehensive analysis method based on online user comment data. The method comprises: calculating a user demand intensity value; calculating a user demand emotion score; calculating a user demand weight value by adopting an IGR-AHP model based on the user demand emotion score; and respectively endowing the user demand intensity value and the user demand weight value with weight coefficients by using a weighted average method, thereby calculating a user comprehensive evaluation index value, and finally carrying out priority ranking on the user demands according to the user comprehensive evaluation index value. The method has the advantages that the LDA topic model is used for text mining, and the problem of how to effectively extract user demand information from a large number of user comments is solved; a user demand emotion score is calculated by fine tuning the BERT model, and the problem that a traditional emotion dictionary analysis method is tedious in process is solved; and finally, combining and considering the strength value and the weight value of the user demand through a weighted average method, and carrying out priority ranking on the user demand.
Owner:CIVIL AVIATION UNIV OF CHINA

Data analysis decision method and device, equipment and medium

The invention discloses a data analysis decision-making method, device and equipment and a medium, and is suitable for data analysis decision-making in the financial and medical fields. The data analysis decision-making method comprises the following steps: firstly, performing distributed data acquisition from a target data source by combining a passive collection mode and an active collection mode to obtain original data; and then performing efficient cleaning and processing on the original data through a preset segmentation rule and a pre-trained detection model. And through text mining analysis and association rule mining technologies, key modes and potential associations in a large amount of data are extracted, so that an accurate target portrait is constructed, and dynamic adjustment is performed according to real-time updated original data. And finally, through a preset decision-making algorithm, based on the target portrait, carrying out deep analysis and generating a decision-making report. According to the method, the efficiency and accuracy of data processing and decision support are remarkably improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-document key phrase extraction method based on graph structure node influence

The invention provides a multi-document key phrase extraction method based on graph structure node influence, and relates to the technical field of natural language processing and text mining. Firstly, a candidate phrase set is generated through noun phrase extraction and standardization; secondly, a semantic relation between phrases is captured through local subgraph construction and a sliding window mechanism, and the semantic relation is integrated into a global phrase co-occurrence graph; thirdly, dynamically dividing theme communities based on two-dimensional structure entropy minimization and a potential game model, and identifying phrase groups with high semantic aggregation; then, cross-topic nodes are processed through a structure entropy heuristic function, and flexibility of topic division is enhanced; and finally, in combination with node influence sorting, extracting key phrases with theme representativeness and propagation capability. The method does not need to label data, is suitable for multiple fields of academic literatures, news texts and the like, has high efficiency, accuracy and universality, and provides an innovative solution for multi-document key phrase extraction.
Owner:YUNNAN POWER GRID CO LTD +1

Dialogue text mining method and system for code scanning consumption cash return based on big data

The invention relates to the technical field of artificial intelligence, and discloses a dialogue text mining method and system based on code scanning consumption cash back of big data, and the system comprises a consumption data collection module, a multi-modal analysis module, a strategy decision engine and a dynamic execution module. When code scanning consumption cash-back strategy decision making is carried out, real-time dialogue text semantics and consumption time-space behavior data are fused, a deep association mechanism of consumption intentions and payment scenes is established, the matching deviation of user language expression and actual consumption behaviors can be eliminated in real time, the consistency of a cash-back strategy and real consumption appeals is ensured, and the payment efficiency is improved. The problem of strategy mismatching caused by semantic understanding isolation in a traditional system is effectively avoided, the precision of personalized marketing is improved, double feedback of dialogue emotion tendency and historical strategy effects is dynamically analyzed, the deviation value of strategy feature matching is corrected in real time, and the system can actively adjust weight distribution of feature combinations in a strategy library.
Owner:YIBIN DIGITAL CHAIN INTELLIGENT TECH CO LTD

Production safety accident key cause link identification method based on complex network

The invention relates to the technical field of safety production management and intelligent analysis, in particular to a production safety accident key cause link identification method based on a complex network, which comprises the following steps of: acquiring structured production safety accident report data from a safety production supervision platform by utilizing a web crawler technology; a TF-IDF algorithm is combined with an industrial standard term library to optimize Jieba word segmentation, and a chi-square statistical method is adopted to calculate the correlation between each risk keyword and an accident category. By fusing text mining, statistical analysis and complex network theories and combining field dictionary optimization word segmentation, multi-accident model adaptation and triple centrality index analysis, the problems that a traditional method is single in analysis dimension, depends on subjective experience and is insufficient in technology integration are solved; efficient utilization of unstructured data, scientific layering of risk factors and accurate identification of key cause links are realized, and objectivity and effectiveness of safety management of a complex production system are effectively improved.
Owner:CIVIL AVIATION UNIV OF CHINA

Telecommunication fraud prevention-oriented record text mining method

The invention relates to a record text mining method for telecommunication fraud prevention, which belongs to the field of clue mining, and comprises the following steps: inputting telecommunication fraud records formed in a certain region within a period of time into a KnowLM-13b-i e model, extracting 9 types of entities and relationships, and generating structured data; importing the generated structured data into a Neo4j database, defining three layers of entities of "victim-attribute-structuring" and a corresponding relationship, performing structured analysis and centrality calculation, and further constructing a knowledge graph; initializing the constructed knowledge graph by using a reasoning technology, generating a data set and optimizing a loss function, performing model training on a reasoning model by using the data set, performing multi-task fine tuning in combination with the reasoning model, and enhancing the fraud recognition capability; and inputting a case description, and outputting a fraud type and a subsequent fraud means by the reasoning model. According to the method, the electric fraud record text can be combined with the pre-trained large model to realize knowledge reasoning of the electric fraud record.
Owner:CHINESE PEOPLE'S PUBLIC SECURITY UNIVERSITY

Method and system of converting unstructured digital documents to a structure format using a secure API

In one aspect, a computerized method for document extraction workflow for unstructured documents includes the steps of implementing a text mining operation on a set of digital documents the incoming documents. This is done by defining a document type of each digital document. Based on the document type, the method defines a set of data dictionaries to extract any data from each digital document. The method uses the defined set of data dictionaries to extract any data from each digital document.
Owner:YERRAMSETTY VENKATA SAI RAMAN +1

Sensitive data identification method and apparatus, device, and computer storage medium

The present application discloses a sensitive data identification method and apparatus, a device, and a computer storage medium. A text mining technology is used to mine a plurality of sensitive data rules from a data security specification file of a target industry to form a sensitive data rule base, the rule base is continuously augmented by using technologies such as NLP and NER, and after data to be identified of the target industry is obtained, a sensitivity class and a sensitivity level of the data to be identified can be identified by matching the sensitive data rules in the sensitive data rule base corresponding to the target industry with the data to be identified.
Owner:CHINA UNIONPAY

Multi-modal medical text extraction method based on adaptive domain knowledge fusion

The invention relates to a multi-modal medical text extraction method based on adaptive domain knowledge fusion, and belongs to the technical field of medical text mining, and the method comprises the following steps: text entity-relationship mining: extracting related entities and semantic relationships from clinical structured and unstructured texts, and generating a preliminary text disease network; professional knowledge linking: performing knowledge alignment by using an external medical knowledge graph; and extracting a medical spatio-temporal event chain: extracting an event sequence with a time sequence and a spatial position from a clinical text, constructing a dynamic text disease knowledge network, and extracting a multi-modal medical text.
Owner:THE SECOND AFFILIATED HOSPITAL ARMY MEDICAL UNIV

Online forum-oriented low-resource topic key topic extraction method

The invention belongs to the technical field of natural language processing and text mining, and discloses an online forum-oriented low-resource topic key topic extraction method, which comprises the following steps of: performing semantic-preserving data enhancement on an original text through a large language model to generate an enhanced document set; utilizing a pre-training language model to extract context-aware semantic representation of the document; constructing a learnable topic embedding matrix, and calculating and generating topic distribution; designing a semantic perception contrast learning framework, and optimizing theme diversity by adopting a dynamic negative sample screening strategy; and meanwhile, priori alignment loss is used for ensuring theme consistency. According to the invention, an LLM enhanced data expansion mechanism and a lightweight theme coding architecture are creatively fused, and through dual optimization of contrast learning regularization and prior distribution matching, three technical problems of data sparsity, model over-fitting and noise sensitivity in a low-resource scene are effectively solved; and an efficient and reliable theme modeling solution is provided for social media public opinion analysis.
Owner:NANJING UNIV OF POSTS & TELECOMM

Systems and methods for intelligent content filtering and persistence

A source content processor receives content from a crawler and calls a text mining engine. The text mining engine mines the content and provides metadata about the content. The source content processor applies a source content filtering rule to the content utilizing the metadata from the text mining engine. The source content filtering rule is previously built based on at least one of a named entity, a category, or a sentiment. The source content processor determines whether to persist the content according to a result from applying the source content filtering rule to the content and either stores the content in a data store or deletes the contents from the data ingestion pipeline such that the content is not persisted anywhere. Embodiments disclosed herein can significantly reduce the amount of irrelevant content through the data ingestion pipeline, prior to data persistence.
Owner:OPEN TEXT SA ULC

Adaptive discovery and mixed-variable optimization of next generation synthesizable microelectronic materials

This invention relates to systems and methods for adaptive discovery and mixed-variable optimization of synthesizable microelectronic materials, and applications of the same. Specifically, an exemplary system includes a virtual screening (VS) module to extract information from literatures of a knowledge base by text mining, a ML-assisted conceptual exploration (CE) module to identify candidate material families for the specific class of compound materials based on the extracted information via a combination of ML models and to generate exogenous models of objective functions f(x, y) and constraint functions g(x, y), and an adaptive discovery (AD) engine to generate and optimize design of the newly discovered compound materials. The AD engine includes a mixed-variable ML module, a mixed-integer optimization (MIO) module, and a high-fidelity evaluation (HFE) module, which are iteratively and sequentially executed.
Owner:NORTHWESTERN UNIV +1

E-commerce user portrait construction method based on reinforcement and increase learning

The invention discloses an e-commerce user portrait construction method based on reinforcement and increase learning, and the system comprises a data collection module which collects the information of a user, and provides a data source for the comprehensive analysis of user behaviors; the data processing and analyzing module is used for cleaning and preprocessing the data and analyzing the behavior pattern and preference of the user according to the processed data; and the label generating and grading module is used for automatically generating labels for the user according to the result output by the data processing and analyzing module in combination with a preset label rule base, and calculating the comprehensive score of the user portrait. According to the method, browsing behavior characteristics are calculated through a specific formula, interest preferences are analyzed by applying a text mining technology and a deep learning model, and user behavior modes and interest points can be accurately grasped; a reward mechanism is constructed based on feedback, a push strategy is evaluated by using a reinforcement learning algorithm, an optimal strategy is continuously selected, and the push effect and the accuracy of a user portrait are improved.
Owner:MOUTAI INST

Method for determining importance degree of seepage monitoring points of earth and rockfill dam and application of method

PendingCN121144726AMonitoring siteText mining
The invention discloses a method for determining the importance degree of seepage monitoring points of an earth and rockfill dam and application of the method. The method comprises the following steps that six seepage safety parts of the earth and rockfill dam are recognized, wherein the six seepage safety parts comprise a dam body, an anti-seepage body area, an inverted filtration drainage area, a dam penetrating building area, a dam foundation and a bank slope area; constructing an earth and rockfill dam seepage Bayesian network according to a historical monitoring text, and calculating importance degrees of six seepage safety parts; acquiring three-dimensional space coordinates of the seepage monitoring points: stake numbers, wheelbases and elevations; calculating the comprehensive importance degree of each part to each section of the monitoring point; and superposing the comprehensive importance degrees of the cross sections in the transverse, longitudinal and horizontal directions to obtain the importance degrees of the monitoring points. According to the method, historical monitoring text mining, a Bayesian network probability model and a three-dimensional space attenuation algorithm are coupled, and the importance degree of a quantitative measuring point is calculated, so that the problem that a traditional method depends on experience and neglects local structure defects is solved; and the high-weight section automatically triggers special inspection of the corresponding seepage safety part, so that timely further disposal is facilitated, and the safe operation life of the dam is effectively prolonged.
Owner:NANJING COLLEGE OF CHEM TECH +1

Graph model-based state grid document collaborative clustering analysis method, system and equipment and medium

The invention relates to the technical field of text mining and data clustering, and discloses a state grid document collaborative clustering analysis method, system and device based on a graph model and a medium, comprising: constructing a joint graph model, the joint graph model comprising a document node subset and a word node subset, the edge weight between document nodes being determined based on the diversity between documents, and the word node subset being determined based on the difference between the document nodes; the edge weight between the word nodes is determined based on the diversity between the words, and the document nodes and the word nodes are not directly connected; on the basis of the joint graph model, constructing and solving a collaborative clustering target function to synchronously perform graph division on the document node subset and the word node subset so as to obtain a clustering result of the document and a clustering result of the word at the same time; wherein the collaborative clustering objective function is configured to maximize the sum of the cut value of the document sub-graph division and the cut value of the word sub-graph division. According to the method, the complex dissimilar relationship between the documents and between the words can be visually and effectively captured by utilizing the natural relationship representation capability of the graph model.
Owner:GUANGAN POWER SUPPLY COMPANY STATE GRID SICHUANELECTRIC POWER

Sensitive data identification method and apparatus, device, and computer storage medium

The present application discloses a sensitive data identification method and apparatus, a device, and a computer storage medium. A text mining technology is used to mine a plurality of sensitive data rules from a data security specification file of a target industry to form a sensitive data rule base, the rule base is continuously augmented by using technologies such as NLP and NER, and after data to be identified of the target industry is obtained, a sensitivity class and a sensitivity level of the data to be identified can be identified by matching the sensitive data rules in the sensitive data rule base corresponding to the target industry with the data to be identified.
Owner:CHINA UNIONPAY

Automatic identification, classification and development trend analysis method of net red villages based on multi-source data fusion and natural language processing

The method for automatic identification, classification and development trend analysis of net red villages based on multi-source data fusion and natural language processing comprises the following steps: UGC data is crawled from Xiaohongshu and Douyin through a distributed master-slave architecture, de-duplicated based on SimHash, and normalized in time and coding format; a text semantic fingerprint is generated, and multi-level semantic cache fingerprint matching is performed; for unassigned text, its complexity is calculated, and a large language model API is adaptively called to automatically complete and extract five-level administrative divisions; weights are determined based on the analytic hierarchy process, interaction indicators such as likes, comments, collections and forwards are integrated, and a comprehensive network heat index of the village is obtained; an external text mining tool is connected, and batch word frequency analysis, semantic network analysis and sentiment tendency evaluation are performed; a document-term matrix is constructed, TF-IDF weighting is performed, and unsupervised clustering algorithm is used for clustering analysis of village characteristics; cross-dimension analysis is performed on the clustering results, and a development portrait, advantage mining and operation suggestion warning are automatically generated in combination with the SWOT model.
Owner:ZHEJIANG UNIV OF TECH

Text mining method and system for identifying chemical formula of material

The invention relates to a text mining method and system for identifying a material chemical formula. The method comprises the steps of firstly obtaining a text set and preprocessing to obtain a target text set; marking the positions of the chemical elements, correcting errors to obtain a mark vector, and obtaining a difference vector through exponential offset and difference processing; and then calculating the maximum rising subsequence length of the outlier sequence, setting the cluster number, performing k-means clustering on the offset vector, and finally marking the text according to the k-means clustering result. Compared with the prior art, the method has the advantages of high text chemical formula extraction accuracy, high efficiency and the like.
Owner:TONGJI UNIV

Automated method for virtual technical assistance in the correction of computer vulnerabilities through combined usage of software automation and artificial intelligence technics

The invention relates to an automatic method for technical assistance in the correction of vulnerabilities of a computer system, where the aforementioned method comprises at least the following steps: a) receiving information relative to computer system vulnerabilities by text, voice, file input or through data exchange; b) launching tools based on artificial intelligence algorithms and machine learning to understand the information and requests entered relative to vulnerabilities; c) performing parsing and text mining of the information received and process it through intelligent data matching with information relative to the computer vulnerabilities present in a database or through interaction with external online vulnerability databases; d) in case of unsatisfactory results, starting an extended online search by web scraping using deep learning algorithms and unsupervised machine learning; e) updating an internal vulnerability database with the additional information found; and f) generating a vulnerability report accompanied by relative remediation.
Owner:CYLOCK SRL

Urban land utilization identification method based on BERT model text classification algorithm

The invention relates to an urban land utilization recognition method based on a BERT model text classification algorithm, and the method is characterized in that the method comprises the following steps: 1, determining an urban land utilization recognition type, determining an urban land utilization recognition region range, obtaining POI data in the region range, and carrying out the data screening; step 2, establishing an urban land utilization identification unit in an urban land utilization identification model based on a BERT model text classification algorithm, and associating the screened POI data with the urban land utilization identification unit; step 3, carrying out geographic text mining on the POI data; and step 4, performing urban land utilization identification of the urban land utilization identification model based on the BERT model text classification algorithm to obtain a high-precision urban land utilization identification result. According to the method, high-precision urban land utilization identification can be timely and accurately carried out by utilizing the available POI data.
Owner:TIANJIN UNIV

A keyword-related topic analysis method based on local word embedding technology

The present invention discloses a keyword-related topic analysis method based on local word embedding technology. When a user enters a keyword, the method uses global word embedding technology to analyze the keyword in text big data, sorts the cosine similarity between the vocabulary and the keyword in the data, and selects the subject words. The subject dictionary is constructed based on the subject words, and the text big database is segmented by time, author, article type, etc. The local word embedding analysis is performed on the keywords and the subject dictionary within the segment, and the local cosine similarity in each segment is used to obtain the theme of each segment. Then, the theme change of the keyword in the text big database can be analyzed. The present invention can be widely used in the fields of social science, public opinion monitoring, historical document analysis, etc., and provides an efficient and accurate theme dynamic analysis tool for text mining.
Owner:ZHEJIANG UNIV

Smart Government Service Platform

The present invention relates to the field of task management technology, specifically to a smart government service platform, which includes a priority adjustment module, a process adjustment module, an intelligent text mining module, and an intelligent matching optimization module. The present invention realizes the accurate sorting of task priorities by analyzing the historical data and user behavior of the completion of government service tasks, effectively improves the efficiency of task processing, and can adjust the workflow in real time when responding to public health emergencies or natural disasters, ensuring the immediacy and adaptability of government actions, deeply mining government text data and extracting key information, optimizing the decision support system, and enhancing the accuracy of services. By accurately analyzing the degree of matching between text content and work requirements and optimizing information display, the consistency and efficiency of the workflow are significantly improved, processing time and potential errors are reduced, thereby improving the overall response speed and quality of government services.
Owner:TIANJIN YITIAN DIGITAL SERVICE CO LTD

Multi-mode scientific and technological innovation resource data intelligent screening method

The invention discloses a multi-mode scientific and technological innovation resource data intelligent screening method. The method comprises the following steps: firstly, setting admission thresholds for research and development investment, the number of researchers, the number of patents and the proportion of new product sales income, so as to preliminarily screen out candidate enterprises meeting the lowest requirements; then, unified preprocessing is conducted on enterprise data from different sources, the multi-source and multi-format problem is solved, and enterprise entity analysis and merging are achieved in combination with methods such as editing distance and cosine similarity; on the basis, text semantic features such as enterprise brief introduction and related news are extracted by means of text mining, word frequency-inverse document frequency and the like, and multi-modal feature vectors are constructed together with numerical features; and finally, performing K-means clustering analysis by adopting comprehensive distance measurement, iteratively calculating a clustering center, and dividing the enterprises into clusters with the highest similarity to obtain a multi-modal clustering result.
Owner:CHONGQING ACADEMY OF SCI & TECH

Knowledge graph construction method and device for intellectual property retrieval and storage medium

The invention provides an intellectual property retrieval-oriented knowledge graph construction method and device and a storage medium, and the method comprises the steps: reading a target intellectual property text data set and a parallel corpus and a reference associated text in the target intellectual property text data set, mining synonymous mapping clues, reference traceability clues and technical theme associated clues implied in the texts, and constructing the intellectual property retrieval-oriented knowledge graph by the synonymous mapping clues, the reference traceability clues and the technical theme associated clues; sorting to obtain a weak supervision signal set, extracting synonymous expression pairs in parallel corpora and semantic association pairs in a reference association text, carrying out cross validation and duplicate removal to obtain a text alignment reference set, and carrying out global association matching on the text alignment reference set and the target intellectual property text data set to obtain an initial structured knowledge unit set; the method comprises the steps of obtaining a standardized knowledge unit set, performing credibility regularization to obtain a standardized knowledge unit set, performing entity classification and relation association organization according to hierarchical requirements of intellectual property retrieval, and constructing to obtain the intellectual property retrieval-oriented knowledge graph. The intellectual property retrieval accuracy and response efficiency can be effectively improved through the intellectual property retrieval method and device.
Owner:HENAN UNIV OF ANIMAL HUSBANDRY & ECONOMY

Active demand learning-based few-sample text mining method and apparatus, and electronic device

The invention relates to the technical field of news mining, and provides a few-sample text mining method and device based on active demand learning and electronic equipment. Comprising the steps of generating a structured prompt according to a predefined task template; inputting the structured prompt into a large-scale pre-training model for reasoning to obtain a mining result; calculating an uncertainty score of a specified sample in the mining result, and if the uncertainty score is in a descending trend, returning to execute the step of generating the structured prompt according to the predefined task template; if the uncertainty score is in a rising trend, the number of the difficult samples is increased; receiving annotation information of a user on the difficult sample, updating task label description based on the annotation information, and returning to execute the step of generating the structured prompt according to the predefined task template; and when it is determined that the optimization termination condition is met, outputting a mining result. The demand of a news analysis task can be accurately captured, and the news text mining capability is effectively enhanced.
Owner:WUHAN UNIV OF TECH +1

Text mining-based aviation accident cause intelligent identification method

The invention provides an aviation accident cause intelligent identification method based on text mining, relates to the technical field of aviation accident cause identification, and aims to improve the accuracy of aviation accident cause identification. The method comprises the following steps: acquiring aviation accident report data; extracting a plurality of candidate accident cause factors from the aviation accident report data; determining a confidence coefficient between any two candidate accident cause factors in the plurality of candidate accident cause factors; determining a direct influence matrix based on the confidence degree between any two candidate accident cause factors in the plurality of candidate accident cause factors; based on the direct influence matrix, measuring parameters of all candidate accident cause factors are determined, and the measuring parameters comprise centrality and / or cause degree; the centrality is used for representing the cause importance degree of the candidate accident cause factors; the cause degrees are used for representing cause guiding degrees of the candidate accident cause factors; and determining a target accident cause factor from all the candidate accident cause factors based on the centrality and the cause degree.
Owner:CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD +1

A college scientific research hotspot mining method and system based on improved BERTopic

PendingCN122286698AText miningEngineering
This invention relates to a method and system for mining research hotspots in universities based on an improved BERTopic, belonging to the field of text mining technology. It includes the following steps: Step S1, obtaining a list of effective word segments from university research papers; Step S2, capturing the core semantic information of academic texts; Step S3, obtaining a 3D low-dimensional vector that retains the core semantic features; Step S4, identifying potential research hotspot topic clusters; Step S5, generating a set of university research hotspot topics containing core keywords, weights, temporal attributes, and dual-dimensional labels of "topic + keyword"; Step S6, outputting multi-dimensional visualization results and hotspot prediction results, iteratively adjusting model parameters based on a feedback optimization mechanism, and finally generating a structured report and research decision-making suggestions. This application has the effect of improving the quality of mining research hotspots in universities.
Owner:NANTONG UNIV

An interactive self-service analysis retrieval system and method based on a large model

PendingCN122432203AText miningData retrieval
The application discloses an interactive self-service analysis and retrieval system based on a large model, which comprises a data source management module, a retrieval library construction module, an algorithm development module, a task scheduling management module and a service publishing module; the data source management module pre-processes initial information data to obtain structured information data; the retrieval library construction module constructs a retrieval library according to the structured information data; the algorithm development module extracts data features in the structured information data and recommends an adaptive retrieval algorithm according to the data features; the task scheduling management module performs retrieval analysis in the retrieval library according to the adaptive retrieval algorithm and an input retrieval instruction; and the service publishing module is used for converting the retrieval analysis result into an API service that can be called. The application aims to solve the pain points of the existing data analysis technology, such as fragmented functions, high operation threshold and lack of special text processing modules. The application provides an efficient solution for data modeling, text mining and service deployment.
Owner:NAVAL UNIV OF ENG PLA

Sensitive data detection method and device, electronic equipment and storage medium

The invention discloses a sensitive data detection method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining text data; performing word segmentation processing on the text data to obtain target text data; performing feature extraction on the target text data by using a first model to obtain text features; performing semantic analysis on the text features to obtain target text features; and inputting the target text features into a second model and a third model, and outputting sensitive data. According to the technical scheme, sensitive data features can be deeply mined, the method has higher adaptability to complex and changeable data forms, the recognition accuracy is greatly improved, and the probability of misjudgment and missed judgment is reduced. And deep fusion of natural language processing and text mining enhances semantic comprehension and further guarantees recognition precision.
Owner:AGRICULTURAL BANK OF CHINA