Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Latent Dirichlet allocation" patented technology

In natural language processing, latent Dirichlet allocation (LDA) is a generative statistical model that allows sets of observations to be explained by unobserved groups that explain why some parts of the data are similar. For example, if observations are words collected into documents, it posits that each document is a mixture of a small number of topics and that each word's presence is attributable to one of the document's topics. LDA is an example of a topic model.

Text block dynamic segmentation method and system based on RAG

The invention provides an RAG-based text block dynamic segmentation method and system, and the method comprises the steps: analyzing a multi-level directory structure of a document, and segmenting a text block by taking a directory node as a reference; for a document without a directory or with an incomplete directory, a mixed segmentation mode of rules and semantics is switched; performing potential Dirichlet distribution topic modeling on the continuous text stream, and calculating a topic distribution vector of each text segment in real time; performing adaptive segmentation on the detected topic boundary, inserting a hard segmentation mark at a topic mutation point, and performing soft segmentation on a gradual change topic area; the initial window size of the sliding window is set according to the document type, and the semantic density in the window is monitored in real time; and layering, blocking and recombining. The system comprises a segmentation mode module, a mark confirmation module and a block recombination module. According to the method, the RAG retrieval accuracy is improved, the memory occupation is reduced, and meanwhile, the streaming throughput is supported.
Owner:北京三维天地科技股份有限公司

Political-oriented network public opinion manipulator detection model training method, detection method and device

The application provides a political orientation network public opinion manipulator detection model training method, a detection method and a device, comprising: constructing a training data set, each sample of which contains tweets collected from a social network platform; using a latent Dirichlet distribution to model the theme of the training data set, screening out political direction data to construct a political data set; constructing an initial detection model, which includes a feature selection module and a classification module; inputting the political data set, extracting features by the feature selection module, classifying by the classification module, and outputting a predicted classification result; training the initial detection model using the political data set, constructing a loss optimization model of the predicted classification result and the true classification; evaluating and selecting the best classification algorithm in terms of classification performance as the final classification algorithm of the classification model to obtain a detection model. The training method of the application guarantees the class balance ratio of the training set, has high training efficiency, and the detection model obtained by training has high precision and high performance.
Owner:CHINA ELECTRONICS CYBERSPACE RESEARCH INSTITUTE CO LTD

Automated content tagging with latent dirichlet allocation of contextual word embeddings

PendingEP4675510A2Mathematical modelsNatural language analysisLatent Dirichlet allocationDocumentation
Dynamic content tags are generated as content is received by a dynamic content tagging system. A natural language processor (NLP) tokenizes the content and extracts contextual N-grams based on local or global context for the tokens in each document in the content. The contextual N-grams are used as input to a generative model that computes a weighted vector of likelihood values that each contextual N-gram corresponds to one of a set of unlabeled topics. A tag is generated for each unlabeled topic comprising the contextual N-gram having a highest likelihood to correspond to that unlabeled topic. Topic-based deep learning models having tag predictions below a threshold confidence level are retrained using the generated tags, and the retrained topic-based deep learning models dynamically tag the content.
Owner:PALO ALTO NETWORKS INC

A conversation recommendation method based on cross-category heterogeneous hypergraph multi-intent representation

This invention belongs to the field of conversation recommendation technology, and specifically provides a conversation recommendation method based on multi-intent expression using cross-category heterogeneous hypergraphs. The method includes preprocessed data and the following steps: S1: The preprocessed data is input to a category recognition processing module, which uses a Latent Dirichlet Allocation (LDA) model to mine the latent category distribution of items and outputs category-classified items; S2: The preprocessed data and category-classified items are respectively input to a self-loop star graph module and a cross-category heterogeneous hypergraph module to construct the graph structure, outputting a self-loop star graph and a cross-category heterogeneous hypergraph. This invention combines the conversational intra-conversational structure modeling of self-loop star graphs with the semantic association expression of cross-category heterogeneous hypergraphs to construct a multi-intent modeling framework that can simultaneously capture local behavioral dependencies and global interest transfers. It dynamically identifies latent intent patterns and multi-category preference features even without long-term user history information, thereby significantly improving the accuracy, personalization, and robustness of conversation recommendations.
Owner:CHANGCHUN NORMAL UNIV

Fault diagnosis method and system based on potential Dirichlet allocation

The invention relates to the technical field of fault diagnosis, and provides a fault diagnosis method and system based on potential Dirichlet allocation. The method comprises the following steps: acquiring historical sensor signals, constructing a state word at each moment, and obtaining a corpus constructed by all cases; based on the state of the equipment corresponding to the state word at each moment, setting a theme tag, and constructing a training data set; training an LDA topic model based on the training data set, and extracting potential topics in the data to obtain a trained LDA topic model; obtaining a sensor signal, constructing a state word at each moment, adopting a trained LDA topic model, obtaining topic distribution, calculating the similarity between the topic distribution and each topic distribution in a training data set, and if the similarity with a certain topic distribution is the highest, considering that a fault type corresponding to the topic distribution occurs in the equipment; wherein the normal state is regarded as a special fault type. According to the invention, the fault detection accuracy can be improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Cold-chain logistics file analysis method based on multi-mode and dynamic association

The invention discloses a cold-chain logistics file analysis method based on multimodality and dynamic association, which comprises the following steps of: constructing a cold-chain logistics file lexicon, and carrying out manual annotation on theme keywords of preprocessed cold-chain logistics files; constructing an analysis model, wherein the analysis model comprises a keyword extraction module based on a potential Dirichlet allocation word vector model, a sorting module, a time sequence enhancement module, a business feature module and a technical feature module; based on the manually annotated theme keywords, acquiring theme keywords fusing the time sequence features, the business features and the technical features; constructing a loss function, and optimizing parameters for constructing the analysis model by presetting a label to minimize the loss function; and constructing dynamic association analysis of the business theme and the technical theme. According to the method, quantitative decision support is provided for technical investment priority judgment, supply chain risk early warning and business innovation directions in the field of cold-chain logistics, and the resource allocation efficiency and the policy response speed are remarkably improved.
Owner:EAST CHINA JIAOTONG UNIVERSITY

A biomedical literature long query content retrieval method, device and computer equipment

ActiveCN116414946BEmbody the central ideaImprove precisionContent retrievalSubject matter
The present application relates to biomedical literature content retrieval technology, in particular to a biomedical literature long query content retrieval method, device and computer equipment, the method comprises the following steps: constructing a biomedical literature hierarchical tree, and selecting a storage document segment based on the relevance of a corresponding subject word of a child node to a document segment of a retrieval library document; preprocessing a long query text input by a user to obtain to-be-queried content; performing subject reasoning on the to-be-queried content based on a latent Dirichlet allocation method on the biomedical literature hierarchical tree from top to bottom and layer by layer, that is, calculating the subject relevance of the to-be-queried content to other nodes except the root node from top to bottom, if the relevance is greater than a set threshold, continue to find the next layer of nodes of the child node; if the queried child node is a leaf node, end the subject reasoning, and find N documents closest to the to-be-queried content according to the document subject relevance distribution of the document segment on the leaf node; the present application improves the retrieval accuracy of biomedical literature.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

A Cold Chain Logistics Document Analysis Method Based on Multimodal and Dynamic Association

ActiveCN121351805BCold chainDocument analysis
This invention discloses a cold chain logistics document analysis method based on multimodal and dynamic correlation, comprising the following steps: constructing a cold chain logistics document thesaurus and manually annotating the subject keywords of preprocessed cold chain logistics documents; constructing an analysis model, which includes a keyword extraction module, a ranking module, a temporal enhancement module, a business feature module, and a technical feature module based on a latent Dirichlet allocation word vector model; obtaining subject keywords that integrate temporal, business, and technical features based on manually annotated subject keywords; constructing a loss function and optimizing the parameters of the constructed analysis model by using a preset label minimization loss function; and constructing a dynamic correlation analysis between business and technical subjects. This method provides quantitative decision support for prioritizing technology investment, providing supply chain risk warnings, and determining business innovation directions in the cold chain logistics field, significantly improving resource allocation efficiency and policy response speed.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Device and method for reviewing literature by using Latent Dirichlet Allocation

A device and method for reviewing literature by using Latent Dirichlet Allocation (LDA) is proposed. The device may include a pre-processing unit extracting text data for modeling, and a modeling unit automatically classifying topics as many as a set number (K) and generating a probability distribution of the topics by literature and a probability distribution for words by topic. The device may also include a clustering unit updating the number (K) of the topics, an interest analysis unit confirming trends by topic over time, and a generality analysis unit quantitatively confirming a research scope of each specific topic. The device may further include a similarity analysis unit quantitatively confirming a similarity between the topics, a network analysis unit quantitatively confirming a correlation between the topics, and a display displaying the trends by topic over time, research scope of each specific topic, similarity between the topics, and correlation between the topics.
Owner:PUKYONG NAT UNIV IND ACADEMIC COOPERATION FOUND

A social media public participation prediction method, medium and computer device

The application discloses a social media public participation degree prediction method, medium and computer equipment, wherein the method comprises the following steps: selecting a social media platform, obtaining a green travel related original data set on the social media platform, and obtaining high-quality samples after cleaning and screening; text preprocessing is performed on the high-quality samples to obtain effective samples, and a latent Dirichlet allocation model is selected for training and prediction, and the clustering effect of the latent Dirichlet allocation model is highly dependent on the selection of the number of themes; the number of themes is determined by comprehensively considering the perplexity and consistency evaluation indexes; text themes are extracted according to the theme recognition result of the latent Dirichlet allocation model, a prediction variable set is constructed by fusing multi-modal features, a social media public participation degree prediction model is established, and a participation degree prediction result and key driving factors are output. The application discloses the mechanism of the social media in spreading green travel, and provides a scientific and direct basis for customizing and optimizing a spreading strategy.
Owner:HEFEI UNIV OF TECH

Manufacturing intelligent quality management core element extraction method and electronic device

The application discloses a manufacturing industry digital quality management core element extraction method and electronic equipment, and belongs to the technical field of data processing. The method comprises the following steps: obtaining raw material purchasing data and usage data from a database according to a first text; obtaining a first text sequence based on the purchasing data and the usage data; extracting a named entity in the first text sequence by using a first large language model to obtain a first named entity sequence; processing the first text sequence by using an implicit Dirichlet distribution to obtain a first element list; performing vectorization on the first named entity sequence to obtain a first vector set corresponding to the first named entity sequence; obtaining a second vector set corresponding to the first named entity sequence according to the correlation degree between words in the first element list and a quality management theme; and determining the correlation degree between elements in the first element list based on the first vector set and the second vector set. The application can automatically discover complex correlations between elements.
Owner:CHINA UNIV OF MINING & TECH (BEIJING)

Interest identification method, device, equipment and storage medium

The present application provides an interest identification method, apparatus, device and storage medium, which relates to the field of big data. The method includes: obtaining voice call data, converting the voice call data into text data; constructing a heterogeneous graph based on the text data, and determining the high-frequency words in the heterogeneous graph by statistical methods; classifying the heterogeneous graph by means of a graph neural network to obtain at least one category corresponding to the text data; performing topic analysis on the documents in each category by means of latent Dirichlet distribution to obtain the topic of each category and the keywords in each topic; determining user interests based on the keywords and high-frequency words in each topic. The method of the present application realizes the rapid and accurate identification of user interests, avoids the inaccuracy of manual identification of user interests and the corresponding waste of personnel, reduces the waste of manpower, reduces the time of user interest analysis, and improves the efficiency of user interest identification.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

A jeans appearance emotional preference analysis method

ActiveCN120561274BMarket predictionsWeb data indexingData setLatent Dirichlet allocation
The application provides a jeans appearance emotional preference analysis method, and belongs to the technical field of emotional analysis. Massive real and dynamic user online comment data is collected to obtain original comment data sets, and structured text corpus is formed after preprocessing. A latent Dirichlet allocation model is used to model the theme of the preprocessed text corpus, and a domain knowledge filtering mechanism is introduced to establish a jeans appearance optimal emotional keyword set. Then, the weight values of the emotional keywords in each theme are calculated. Finally, based on the weight calculation results of the emotional keywords, fuzzy comprehensive evaluation method is used to perform fuzzy operation on the candidate design scheme to obtain the emotional preference scores of different design schemes, so that quantitative evaluation and optimization of the design scheme are realized. The application can improve the efficiency and accuracy of emotional keyword recognition, clearly define the priority of user multi-emotional needs, output the emotional preference scores of the user on the product design scheme, and provide the optimal design scheme that meets the emotional needs of the user.
Owner:TIANJIN POLYTECHNIC UNIV +1

A text and topic matching method, system, device and storage medium

The application discloses a text and theme matching method, system, device and storage medium, and the method comprises the following steps: obtaining a text information pair to be matched, wherein the text information pair to be matched comprises a theme keyword group to be matched and text to be matched; performing word segmentation processing on the text information pair to be matched to obtain word segmentation information; inputting the word segmentation information into a keyword recognition model to obtain a text keyword group; using an implicit Dirichlet distribution theme recognition model to obtain a theme keyword weight corresponding to each text keyword group, and continuing optimization through matching model training; according to the theme keyword weight, performing weighted calculation on the theme keyword group to be matched and the text keyword group, or adopting a regular expression to judge a matching degree, and obtaining a matching result. The application has higher recognition rate and recall rate, has high text matching success rate for texts of different lengths, has good recognition effect, and has shorter reasoning time.
Owner:HAINA CLOUD IOT TECH CO LTD +2

Method for evaluating driving risk level in tunnel based on vehicle bus data and system therefor

A method for evaluating a driving risk level in a tunnel based on vehicle bus data and a system therefor are provided. The method uses the Controller Area Network (CAN) bus data collected in a vehicle driving process, designs and extracts a driving risk characteristic feature index reflecting the driving behavior of a driver through a sliding time window method, writes a feature codebook to symbolize an extracted sequence feature, and then randomly samples all of the samples, and based on the sampled symbolic data, using a Latent Dirichlet Allocation (LDA) theme model to evaluate the driving risk level. The training method of the model is to acquire the optimal number of risk levels by evaluating the perplexity and the coherence scores, and to analyze the driving risk of the driving data in all of the samples.
Owner:ZHEJIANG UNIV ZHONGYUAN INST +1

Construction method and device of theme keyword dictionary, equipment and storage medium

The invention discloses a theme keyword dictionary construction method and device, equipment and a storage medium. The method comprises the following steps: acquiring a training corpus text; based on a to-be-trained topic-keyword dictionary, performing latent Dirichlet allocation LDA clustering on the training corpus text to obtain a clustering result, the clustering result at least comprising a topic corresponding to each sentence in the training corpus text; and updating a to-be-trained topic-keyword dictionary on the basis of sentences with topics being known topics in the training corpus text to obtain an updated topic-keyword dictionary. According to the method, the topic-keyword dictionary is constructed, so that information in corpora can be mined through the topic-keyword dictionary.
Owner:PIPECHINA SOUTH CHINA CO +1

Method and device for generating event evolution relationship tree

This disclosure provides a method, apparatus, device, and storage medium for generating an event evolution relationship tree, which can be applied to the field of natural language processing technology. The method includes: processing a text dataset to be processed based on a word frequency-inverse document frequency algorithm to obtain a first text feature matrix; obtaining a density boundary threshold based on the Euclidean distance between each element in the first text feature matrix and other elements; clustering the elements in the first text feature matrix based on the density boundary threshold to obtain multiple target text datasets corresponding to multiple topic events; processing the target text dataset corresponding to each topic event based on a hidden Dirichlet distribution algorithm to obtain multiple related sub-events; and generating an event evolution relationship tree corresponding to each topic event based on the occurrence times of the multiple related sub-events.
Owner:SUZHOU AEROSPACE INFORMATION RES INST

Manufacturing industry digital intelligence quality management core element extraction method and electronic equipment

The invention discloses a manufacturing industry digital intelligent quality management core element extraction method and electronic equipment, and belongs to the technical field of data processing. The method comprises the following steps: retrieving from a database according to a first text to obtain purchase data and use data of raw materials; obtaining a first text sequence based on the purchase data and the use data; using a first large language model to extract named entities in the first text sequence to obtain a first named entity sequence; processing the first text sequence by using implicit Dirichlet distribution to obtain a first element list; vectorizing the first named entity sequence to obtain a first vector set corresponding to the first named entity sequence; obtaining a second vector set corresponding to the first named entity sequence according to the association degree of words in the first element list and the quality management theme; and determining the correlation degree between the elements in the first element list based on the first vector set and the second vector set. According to the method, complex association between elements can be automatically found.
Owner:CHINA UNIV OF MINING & TECH (BEIJING)

Method for evaluating driving risk level in tunnel based on vehicle bus data and system therefor

A method for evaluating a driving risk level in a tunnel based on vehicle bus data and a system therefor are provided. The method uses the Controller Area Network (CAN) bus data collected in a vehicle driving process, designs and extracts a driving risk characteristic feature index reflecting the driving behavior of a driver through a sliding time window method, writes a feature codebook to symbolize an extracted sequence feature, and then randomly samples all of the samples, and based on the sampled symbolic data, using a Latent Dirichlet Allocation (LDA) theme model to evaluate the driving risk level. The training method of the model is to acquire the optimal number of risk levels by evaluating the perplexity and the coherence scores, and to analyze the driving risk of the driving data in all of the samples.
Owner:ZHEJIANG UNIV ZHONGYUAN INST +1

RAG-based text block dynamic segmentation method and system

The present invention provides a method and system for dynamic text block segmentation based on RAG. The method comprises parsing a document's multi-level directory structure and segmenting text blocks based on directory nodes; switching to a hybrid segmentation mode based on rules and semantics for documents without or with incomplete directories; performing latent Dirichlet allocation topic modeling on continuous text streams and calculating the topic distribution vector for each paragraph in real time; performing adaptive segmentation on detected topic boundaries, inserting hard segmentation markers at topic mutation points, and employing soft segmentation for gradually changing topic areas; setting the initial window size of the sliding window based on the document type and monitoring the semantic density within the window in real time; and hierarchical block reorganization. The system comprises a segmentation mode module, a marker confirmation module, and a block reorganization module. The present invention improves the accuracy of RAG retrieval, reduces memory usage, and supports streaming throughput.
Owner:北京三维天地科技股份有限公司

Automated content tagging with latent Dirichlet allocation of contextual word embeddings

ActiveUS12417353B2Mathematical modelsNatural language analysisConfidence metricLatent Dirichlet allocation
Dynamic content tags are generated as content is received by a dynamic content tagging system. A natural language processor (NLP) tokenizes the content and extracts contextual N-grams based on local or global context for the tokens in each document in the content. The contextual N-grams are used as input to a generative model that computes a weighted vector of likelihood values that each contextual N-gram corresponds to one of a set of unlabeled topics. A tag is generated for each unlabeled topic comprising the contextual N-gram having a highest likelihood to correspond to that unlabeled topic. Topic-based deep learning models having tag predictions below a threshold confidence level are retrained using the generated tags, and the retrained topic-based deep learning models dynamically tag the content.
Owner:PALO ALTO NETWORKS INC

Method and system for maintaining latent Dirichlet allocation model accuracy

A method and a system for counteracting model drift in a Latent Dirichlet Allocation (LDA) model by maintaining freshness and accuracy in the LDA model by leveraging LDA classification and vector algebra to detect potential degradations are provided. The method entails determining whether an LDA model has drifted and therefore requires retraining based on measurements of cosine similarities of respective topics that correspond to two LDA models and a topic match entropy between the two LDA models that is determined based on the measurements.
Owner:JPMORGAN CHASE BANK NA

Automobile public opinion analysis method

The invention discloses an automobile public opinion analysis method, and relates to the technical field of public opinion analysis, and the method comprises the steps: obtaining automobile public opinion data; inputting the automobile public opinion data into a trained sentiment classification model to obtain a sentiment classification result of the automobile public opinion information; mining themes in the automobile public opinion data by using a potential Dirichlet allocation model to obtain theme keywords; constructing a co-occurrence matrix based on the subject keywords, generating a text network graph and analyzing the text network graph to obtain a text network analysis result; and obtaining a public opinion analysis result based on the sentiment classification result, the theme keyword and the text network analysis result. According to the scheme, the internal relation between each theme and the keyword is captured, the automobile manufacturer is helped to grasp the opinions of the public on the brand in real time, and a targeted strategy is formulated.
Owner:ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD

Multi-field marketing topic robot detection method based on feature fusion

The invention relates to a multi-field marketing topic robot detection method based on feature fusion, and belongs to the technical field of social network data analysis and machine learning model detection. Constructing weighted metadata feature representation, vectorizing metadata such as a user introduction, a concerned number and a concerned number through an embedded layer, and introducing an attention mechanism to calculate an attribute weight through multi-layer perceptron nonlinear transformation to obtain weighted metadata features; performing domain division on the text topics based on a potential Dirichlet allocation model, constructing a multi-domain relation network diagram by taking likes, collection, forwarding and comments of users in different domain topics as edges, and fusing weighted metadata features into the diagram; extracting multi-domain social relation features in the domain relation network by using fine granularity of a graph neural network; a long-short-term memory network model based on multiple fields is constructed, user multi-time slice behavior features are extracted to form a time sequence feature matrix, and the marketing topic robot is effectively detected and recognized in combination with social relation features.
Owner:CHONGQING UNIV OF POSTS & TELECOMM +1