Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Bigram" patented technology

A bigram or digram is a sequence of two adjacent elements from a string of tokens, which are typically letters, syllables, or words. A bigram is an n-gram for n=2. The frequency distribution of every bigram in a string is commonly used for simple statistical analysis of text in many applications, including in computational linguistics, cryptography, speech recognition, and so on. Gappy bigrams or skipping bigrams are word pairs which allow gaps (perhaps avoiding connecting words, or allowing some simulation of dependencies, as in a dependency grammar).

Hard tag text countermeasure attack sample generation method

The invention discloses a hard tag text adversarial attack sample generation method, which belongs to the technical field of text adversarial sample generation, and comprises the following steps: determining a text adversarial sample target; binary phrases are constructed for the original samples, and initial confrontation samples are generated by preferentially replacing the binary phrases; replacing unnecessary synonyms with original words; searching synonym iteration replacement optimization confrontation samples by determining transition replacement words; according to the hard tag text adversarial attack sample generation method, in an adversarial sample initialization stage, firstly, a high-frequency binary phrase is replaced to improve the quality of a generated initial adversarial sample; in the confrontation sample optimization stage, a relatively simple and effective mode is adopted to carry out iterative optimization, a generative method is not adopted any more, consumption of too much query times is avoided, the quality and efficiency of generating the confrontation sample are improved, the quality of the confrontation sample is remarkably improved in an attack scene closer to reality, and meanwhile, the query cost is reduced. And the query cost is reduced while the generation quality of the adversarial sample is improved.
Owner:LIYANG RES INST OF SOUTHEAST UNIV +2

A bad review classification method, device and equipment based on a bert model and a medium

The present application relates to big data technology, and discloses a bad review classification method based on a BERT model, which comprises the following steps: data integration and data cleaning are performed on evaluation information to obtain effective evaluation data, and manual sentiment labeling is performed on the effective evaluation data to obtain labeled data; the labeled data is used to perform parameter tuning on a pre-trained BERT model, the trained BERT sentiment classification model is used to classify the evaluation data, positive sentiment data is removed, and data with an emotional category that is not equal to a star level in the evaluation is extracted, screened and labeled; the labeled data and the effective evaluation data are used to retrain the BERT model; negative sentiment data is subjected to word segmentation, a BTM (Bigram Topic Model) is used to extract a bad review theme of the negative word segmentation, and the evaluation information is classified according to the bad review theme. The present application also provides a bad review classification device, equipment and medium based on the BERT model. The present application can improve the accuracy of customer bad review information extraction.
Owner:CHINA PETROLEUM & CHEMICAL CORP +1

Character grouping method based on binary character frequencies and security character library construction method

The present invention particularly relates to a character grouping method based on bigram frequencies and a method for constructing a secure character library. The character grouping method includes the following steps: traversing a corpus to statistically obtain a bigram frequency matrix of the occurrences of any two characters among N characters to be grouped; traversing the characters one by one in descending order of character frequency, and calculating the weight of the character c to be assigned to the k-th group according to a formula; adding the character c to be assigned to the group with the largest weight, and so on until all characters are grouped. The bigram frequency matrix reflects the frequency of two characters appearing together. Through the weight calculation formula, the weight increases when two characters that often appear together are in different groups. In this way, we can make the characters that appear together be in different groups as much as possible by selecting the group with the largest weight, thereby realizing reasonable grouping of characters. This grouping method does not limit the number of characters in each group, making it more reasonable.
Owner:HEFEI HIGH DIMENSIONAL DATA TECH CO LTD

Natural language short text-oriented topic clustering processing method and device

The invention provides a natural language short text-oriented topic clustering processing method and device, relates to the technical field of natural language processing, and aims to solve the technical problems that an existing topic clustering method is low in timeliness and accuracy and difficult to meet actual requirements. The method comprises the steps of obtaining a natural language text to be processed, and performing word segmentation processing on the natural language text; on the basis of the natural language text subjected to word segmentation processing, stop words are removed, and the stop words represent words which do not contribute to semantics of the natural language text; based on the natural language text after the stop words are removed, introducing associated words from a pre-constructed knowledge graph to perform semantic enhancement on the natural language text; based on the natural language text after semantic enhancement, any two keywords are extracted for disordered combination, and binary phrases are generated; and inputting the binary phrases into a pre-trained topic clustering model for topic clustering to obtain topic distribution of the natural language text.
Owner:齐鲁空天信息研究院

AI service window background display method and system

The invention relates to the technical field of screening display, and particularly discloses an AI service window background display method and system, and the method comprises the steps: obtaining dialogue information, and converting the dialogue information into a bigram sequence; counting all bigram sequences in a preset time period, determining a TF-IDF value of each bigram in each bigram sequence based on the counted bigram sequences, and converting each bigram sequence into a vector; on the basis of the vectors, clustering is carried out on all the bigram sequences, and for each type of bigram sequences, an identification sequence is selected from the type of bigram sequences; dialogue information corresponding to the identification sequence is read, the difference rate between the identification sequence and the related sequence is obtained, and the dialogue information and the difference rate are counted to serve as display content and sent to the background; according to the method, multiple pieces of dialogue information are simplified into a limited number of texts and the difference rate through a feature extraction scheme and a clustering scheme, the background display content is greatly simplified, and the work pressure of background workers is reduced.
Owner:国网山西省电力有限公司阳泉供电分公司 +1

A method for Chinese speech enhancement recognition and text correction correction

The application belongs to the field of speech and text processing, and particularly relates to a Chinese speech enhancement recognition and text error correction method, which comprises the following steps: preprocessing the audio to be recognized, extracting features through a voiceprint model, and establishing an initial coarse dialect identification model; establishing an initial network model to train the initial coarse dialect identification model to obtain a dialect identification model; determining an error correction candidate word segmentation set based on an N-gram language model; and outputting the text after error correction and correction through a Bigram 2-gram language model and the N-gram language model. The application preprocesses the audio to be recognized, reduces speech recognition interference factors, improves the recognition performance by adopting a GMM-SVM model, has a faster and better training fitting effect by adopting a combination model of a GMM-UBM model and an LSTM model to establish an initial network model, effectively reduces the error rate by processing the text through an N-gram language model and a Bigram 2-gram language model, and optimizes the result of converting the audio to be recognized into text information.
Owner:NANTONG UNIV