Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

109 results about "Lexical frequency" patented technology

Description Lexical usage frequency is known to influence the application rate of some variable processes. Specifically, variable lenition processes typically affect frequent lexical items more often than infrequent lexical items. For instance, variable t/d-deletion in English is more likely to apply to a frequent word (just)...

Lightweight knowledge graph rapid construction method and system based on NLP technology

The invention discloses a lightweight knowledge graph rapid construction method and system based on an NLP technology. The method comprises the steps of preprocessing an unstructured text and segmenting the unstructured text into semantic segments; unsupervised clustering is combined with the contour coefficient to determine the optimal clustering number, and a semantic association fragment cluster is obtained; performing word segmentation, part-of-speech tagging, NER and entity linking on the fragment cluster, extracting an entity and initial relationship, and fusing semantic similarity, TF-IDF word frequency collaboration degree and co-occurrence frequency to calculate a relationship edge weight; constructing a lightweight knowledge graph; in the question and answer stage, questions are disassembled through a process engine, related sub-graphs are retrieved, and answers with reasoning links are generated. The system correspondingly comprises a text preprocessing module, a clustering module, an entity relation processing module, a graph construction module, a question and answer reasoning module and a storage module. According to the method, the construction cost is reduced, the interpretability and the module coupling degree are improved, multiple scenes such as government and enterprise public opinions and medical assistance are adapted, and the problems of weak generalization, poor real-time performance and'black box 'in the traditional technology are solved.
Owner:XIAMEN MEIYA PICO INFORMATION CO LTD +1

Method and system for generating intelligent insight report based on AI large model

The invention relates to the field of intelligent report generation, in particular to an intelligent insight report generation method and system based on an AI large model, and the method comprises the steps: inputting an insight demand, and generating an insight data package comprising insight contents, associated data and industry labels; extracting a basic keyword set of the insight data packet to form a mixed feature code; after mixed feature coding preprocessing, weight distribution is carried out; and after weight distribution of the AI large model, injecting a high-weight feature vector into a semantic understanding core layer, injecting a low-weight feature vector into a logical reasoning layer, and outputting analysis data to form an intelligent insight report. According to the method, word embedding parameters are optimized according to field semantic characteristics, a parameter verification mechanism is introduced, field text characteristics are adapted by adjusting vector dimensions, context windows and low-frequency vocabulary filtering threshold values, window parameter validity is verified through cosine similarity, word frequency threshold value reasonability is verified through standard deviation, and field text characteristic matching is achieved. And it is ensured that the feature vectors can accurately capture domain-specific semantics.
Owner:SUZHOU YINGTIANDI INFORMATION TECH CO LTD

Semantic understanding-based completion environmental protection acceptance survey report auxiliary auditing system

The invention relates to the technical field of natural language processing, in particular to a semantic understanding-based completion environmental protection acceptance survey report auxiliary auditing system, which comprises a structure recognition module, a semantic chain construction module, a semantic reconstruction module, an image-text coordinate fusion module and a spatial offset judgment module. In the method, through extracting subjects, terms and facility words in facility paragraphs and establishing a combination relationship, ordered labeling of semantic elements is realized, statement structure definition is enhanced, a semantic chain group is constructed and facility classification is associated, semantic consistency comparison is supported, semantic sequences and word frequency distribution in design and reply contents are compared, and description consistency is verified; facility coordinates in an image are extracted, character paragraphs are matched, image-text content anchoring, orientation, terrain and distance word extraction and conversion into space parameters to be compared with an image path are achieved, space consistency verification is completed, and the intelligent level of auditing on semantic comprehension, image-text fusion and space logic is overall improved.
Owner:SHANGHAI RUIDUN INFORMATION TECHNOLOGY CO LTD +1

Speed Up Methods and Systems for Large Language Model Training

A method initializes and accelerates training of neural network based large language model, including by: (i) accessing a corpora for training a neural-network based large language model having word embeddings and word projections in respective word embedding and word projection layers and at least one hidden layer; (ii) counting raw token frequencies associated with content within the corpora; (iii) smoothing the raw token frequencies into a series of vector norms based on log or scaled log functions parameterized by maximum norm information; and (iv) injecting vector norm information into word embeddings and / or word projections based on norm-angle reparameterization to prepare the large language model for training.
Owner:APPL TECH APPTEK

Description generation method based on continuous zero sample

The invention provides a description generation method based on continuous zero samples. The description generation method comprises the following steps: sequencing plain text training corpora according to the difficulty of continuous learning; setting that each training step t only can access the t-th corpus, and for the t-th learning step, generating a synthetic image for an input text of the t-th corpus by using a diffusion model; extracting image embedding of the synthesized image by using a CLIP image encoder; retrieving a group of text description embedding with similar semantics in a text corpus by using image embedding; image embedding and text description embedding are input into a fusion module network, and the output of a fusion module is defined as a soft prompt of a large language model; for an input text, key entities are extracted from the text through word frequency statistics so as to construct a soft prompt and a hard prompt of the large language model; hard prompts and soft prompts are input into a large language model network, and key parts in input are concerned through an attention mechanism.
Owner:JIANGSU VOCATIONAL & TECHNICAL UNIVERSITY OF ARCHITECTURE

Analysis method and device for credit granting approval process

The invention provides a credit granting approval process analysis method and device. Comprising the steps of performing word segmentation on text content in a credit granting approval document set, and combining the obtained segmented words to obtain combined phrases; calculating weighted word frequency-inverse document word frequency of the combined phrases; target phrases with weighted word frequency-inverse document word frequency larger than a word frequency threshold value are screened out from the combined phrases, and an amplified word list is generated according to the target phrases and the original word list; training a target word segmentation device based on the amplified word list; finely adjusting the basic vector model according to the training corpus to obtain a term vector model; analyzing the auditing opinion information based on a large language model to obtain an initial risk analysis rule containing a preset dimension; based on a similarity algorithm, a target word segmentation device and a term vector model, grouping and integrating the initial risk analysis rules to obtain risk analysis rules; and processing the information of each examination and approval stage by using a big language model and adopting a risk analysis rule to obtain examination and approval suggestion information.
Owner:MINSHENG BANKING CORP

Automatic identification, classification and development trend analysis method of net red villages based on multi-source data fusion and natural language processing

The method for automatic identification, classification and development trend analysis of net red villages based on multi-source data fusion and natural language processing comprises the following steps: UGC data is crawled from Xiaohongshu and Douyin through a distributed master-slave architecture, de-duplicated based on SimHash, and normalized in time and coding format; a text semantic fingerprint is generated, and multi-level semantic cache fingerprint matching is performed; for unassigned text, its complexity is calculated, and a large language model API is adaptively called to automatically complete and extract five-level administrative divisions; weights are determined based on the analytic hierarchy process, interaction indicators such as likes, comments, collections and forwards are integrated, and a comprehensive network heat index of the village is obtained; an external text mining tool is connected, and batch word frequency analysis, semantic network analysis and sentiment tendency evaluation are performed; a document-term matrix is constructed, TF-IDF weighting is performed, and unsupervised clustering algorithm is used for clustering analysis of village characteristics; cross-dimension analysis is performed on the clustering results, and a development portrait, advantage mining and operation suggestion warning are automatically generated in combination with the SWOT model.
Owner:ZHEJIANG UNIV OF TECH

Word-tag-based language system for sentence acceptability judgment

Methods and systems for sentence acceptability judgment. A word frequency distribution for a predetermined textual data set is obtained. A replacement rate is determined based on an obtained word frequency distribution for a predetermined textual data set and every occurrence of a word having a frequency lower than the replacement rate is replaced in text of a training data set with a corresponding tag to generate revised text of the training data set (the training data set comprising at least a portion of the predetermined textual data set). A plurality of language models are trained with the revised text and a best performing trained language model of the plurality of trained language models is selected. An acceptability of each sentence of a plurality of candidate sentences is rated using the selected trained language model and a best sentence of the plurality of candidate sentences is selected based on the ratings.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

A domain knowledge triple extraction method, system, medium and device

The application discloses a domain knowledge triple extraction method, system, medium and equipment, and relates to the technical field of knowledge engineering.The domain knowledge triple extraction method and system provided by the application can provide high-quality and efficient data support for triple extraction by analyzing and converting the obtained domain professional text into a structured vector knowledge base which can be efficiently searched, then the domain professional text is subjected to quantitative calculation and fusion screening, the initial entities of domain core high-frequency words which have global high frequency and cross-text universality are mined, the problem of "local optimum and global deviation" caused by traditional single word frequency is avoided, non-core term interference is effectively avoided, and the stability, relevance and efficiency of iterative extraction are improved, and further, the high-frequency word initial entities are subjected to closed-loop iteration through iterative RAG algorithm combined with the structured vector knowledge base, accurate extraction of the large model and iterative correlation entities, the explicit and implicit semantic correlation triples in the text can be mined layer by layer, the strong relevance and structural integrity of the extracted triples are ensured, and the professionalism and accuracy of the model in extracting triples are improved.
Owner:XIAN UNIV OF TECH

Method and device for fine-tuning bert model, equipment and storage medium

Embodiments of the present application provide a BERT model fine-tuning method and device, equipment and a storage medium, and relate to the technical field of computers. The method comprises: generating an original token list according to a text corpus, and obtaining word frequency statistical information of the original token list; determining a mask generation strategy based on the word frequency statistical information, and performing mask processing on the original token list according to the mask generation strategy to obtain a mask processing result; and performing iterative learning on a pre-trained BERT model based on the mask processing result until a preset condition is met, to obtain a fine-tuned BERT model. The present application statistically obtains word frequency information corresponding to a text corpus, and adaptively adjusts a mask generation probability according to the word frequency, so that the mask distribution generated by tokens of different frequencies is more uniform, and the semantic information of the text corpus is better preserved, thereby improving the learning ability of the BERT model for masked corpus.
Owner:BEIJING TOPSEC NETWORK SECURITY TECH +2

A multi-directional text alignment method

The application provides a multi-directional text comparison method, comprising: exporting a packaging design drawing into a PDF format and parsing text content and corresponding position information; splitting the text parsed from the PDF according to intervals, and judging whether the split text is consistent; calculating 2-gram word frequency by using a large amount of Chinese corpus, calculating the probability of text in a normal order and a reverse order, and taking the probability as a basis to judge whether the text is in a reverse order; processing all the text according to position coordinates, judging the direction of the text block according to the normal order and the reverse order of the text, then sorting and merging the text in the text block; comparing the PDF text content with the standard text content for examination, matching similar lines, and marking the differences of the similar lines. The processing result can greatly reduce the structured difference between the parsed text and the actual text, improve the detection accuracy of the parsed text and the actual text, reduce the workload of manual intervention, thereby reducing the packaging design cost and precision.
Owner:PU HUA KE JI YOU XIAN GONG SI

Topic model generation method and device, storage medium and program product

The embodiment of the invention provides a topic model generation method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: preprocessing an obtained user comment text and comment metadata, and correspondingly obtaining a word frequency vector and a text label; inputting the word frequency vector into a preset topic model, converting the word frequency vector into potential space representation, and performing normalization processing to obtain topic distribution corresponding to a user comment text output by the topic model; inputting the user comment text into a preset attention mechanism model to obtain a text embedding vector corresponding to the user comment text; and based on the topic distribution, the text embedding vector and the text label, carrying out joint training, and determining a trained target model. The attention mechanism and the comment metadata are combined in the modeling process, the trained target model is determined, and the problems that in the prior art, a topic model is insufficient in real scene data processing capacity and low in efficiency are solved.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

A method for generating a pinyin bucket word library of a four-level index

PendingCN122452553ADigital dataData integrity
The present application relates to the technical field of electronic digital data processing, and discloses a four-level index pinyin bucket word library generation method, acquires a Chinese word library text, extracts the first Chinese character and the first pinyin of each line of words as a classification key, and generates a triple; the triple is classified into a corresponding pinyin bucket according to the pinyin, the words in the bucket are sorted in descending order of word frequency, an independent Chinese character block is generated for each unique Chinese character, and the words are solidified according to the word length by using a first character multiplexing mechanism; a four-level static index of an initial letter statistical area and a pinyin index area and a Chinese character word index area and a word storage area is sequentially constructed, each index area uses fixed-length entries and absolute offset addressing; a metadata area containing a file length double-semantic field is constructed, a high-frequency-first deterministic truncation is performed on the word library according to a preset threshold, each data area is spliced and a data integrity check value is appended, and an embedded binary word library is generated. The problems of high storage redundancy and uncontrollable memory are solved, and the purposes of deterministic analysis, resource adaptation and high security are achieved.
Owner:SICHUAN HAIGE HENGTONG PRIVATE NETWORK TECH CO LTD

An intelligent customer service system, method, medium and device based on emotion recognition

The present disclosure relates to an intelligent customer service system and method based on emotion recognition, equipment and medium, the system comprises: a voice data extraction module for extracting voice data of a user based on a conversation between the system and the user and storing; a word segmentation processing module for performing word segmentation processing on the voice data to obtain word vector data; an entity recognition module for extracting entity keywords in the word vector data and performing first sorting on the word frequency of the entity keywords; an emotion recognition module for extracting emotion keywords in the word vector data and performing second sorting according to the word frequency of the emotion keywords; a decision module for selecting a corresponding response strategy based on a decision algorithm according to the first sorting and the second sorting; a voice module for generating a reply voice statement based on the response strategy. The system and method of the present disclosure recognize user emotions, assist in decision-making, and ultimately improve user service quality.
Owner:PING AN PAY ELECTRONIC PAYMENT CO LTD

A language processing-based word vector representation method, device and terminal equipment

The application is suitable for the field of artificial intelligence technology, and provides a word vector representation method and device based on language processing and a terminal device, the method comprising: constructing an input word table; setting word frequency information of each word in the input word table; classifying the words into N clusters according to the word frequency information, words with the same word frequency information being classified into the same set to obtain a clustering set with N subsets; defining a first word vector dimension of the words in the first cluster through an experience factor, and obtaining a second word vector dimension of the words in each subset according to the subset where the words are located and the first word vector dimension; defining an nth linear mapping matrix for the nth cluster; and processing a target word vector through parameters shared by the input word table as an input of a natural language processing model. Through the application, the problem that the calculation amount is large, the efficiency is low and the word distribution features cannot be considered when converting words into word vectors in natural language processing can be solved.
Owner:SICHUAN LAN-BRIDGE INFORMATION TECHNOLOGY CO LTD

A semantic recognition model based on NLP and a training method thereof

The application discloses a semantic recognition model based on NPL and a training method thereof, relates to the technical field of semantic recognition model training, and aims to solve the problem of low recognition accuracy of a semantic recognition model. The current whole sentence is searched, multiple segmented sentences are screened out, words in the segmented sentences are recognized, the part-of-speech information of the words is collected, the part-of-speech weight under the corresponding part-of-speech information is called, the number of repeated occurrences of each word in the current whole sentence is counted and the repeated word frequency is calculated, the processing level of each word is calculated by comprehensively considering the repeated word frequency and the part-of-speech weight, the target words are screened out and recognized, the semantics of the target words are acquired and the number of semantics is counted under the current interactive window, the interactive duration of the target words is analyzed by calling historical data, the number of word reference sources is set as a reference, the number of semantics of the target words and the interactive duration are comprehensively considered to adjust the number of word reference sources of each target word, and the semantic analysis accuracy of each target word in the interactive sentence is improved.
Owner:SHANGHAI WUKE INTERNET TECHNOLOGY CO LTD

Document retrieval method and device based on AI large model

The invention provides a document retrieval method and device based on an AI large model, and belongs to the technical field of artificial intelligence. The document retrieval method comprises the steps that semantic extension is conducted on keywords in a query statement, and a plurality of extension words are obtained; obtaining a word frequency and an inverse document frequency of each keyword relative to each candidate document, and a word frequency and an inverse document frequency of each expansion word relative to each candidate document; determining a correlation score of each candidate document according to each word frequency and each inverse document frequency; and outputting a plurality of candidate documents of which the correlation scores are ranked in the top. According to the method, semantic enhanced sparse retrieval is realized, the method has the advantage of high efficiency of traditional sparse retrieval, and the retrieval accuracy is also improved.
Owner:BEIJING XIN LI FANG TECH INC

Urban scenic spot recommendation management method and system based on artificial intelligence

The invention relates to the technical field of scenic spot recommendation management, in particular to a city scenic spot recommendation management method and system based on artificial intelligence. The method comprises the following steps: receiving a scenic spot demand text input by a user, extracting a keyword data set in the scenic spot demand text, and performing word frequency evolution analysis to obtain keyword frequency evolution data; thirdly, deducing the playing tendency intention of the user according to the keyword frequency evolution data, obtaining tendency intention clustering data through clustering processing, and evaluating the adaptation degree of the urban scenic spots based on the data; and finally, extracting adaptive judgment logic, performing optimal recommendation strategy learning, outputting an adaptive recommendation strategy, and embedding the adaptive recommendation strategy into a terminal. The scenic spot recommendation management technology is optimized, so that the scenic spot recommendation management technology is more perfect.
Owner:BEIJING NORMAL UNIVERSITY

An entity matching method, system, device and medium based on word frequency

The application relates to an entity matching method, system, device and medium based on word frequency, wherein the method comprises the following steps: performing word segmentation on aliases in a plurality of entity data to obtain a first word segmentation list and an alias word set, counting word frequency data of words in the alias word set, removing city words existing in the first word segmentation list to obtain a second word segmentation list, judging whether the words in the second word segmentation list are common words according to the word frequency data, processing the words in the second word segmentation list according to the judgment result, obtaining alias keywords corresponding to the aliases, and querying whether the alias keywords exist in a text corpus; if the alias keywords exist, an entity corresponding to the alias keywords is recognized in the text corpus. Through the application, the problems of missed recognition caused by alias simplification and alias synonym misjudgment are solved. The alias keywords are obtained based on the word frequency of the words in the aliases, entity matching is more accurate, and the accuracy of entity matching is improved.
Owner:火石创造科技有限公司

A data mining method and system for social robots

The application discloses a data mining method and system for a social robot. The method comprises the following steps: determining an interactive object currently communicating with the social robot and collecting social media data; performing word segmentation and vectorization processing on the text through a natural language processing algorithm, and extracting an emotional tendency value through an emotional analysis model, and dividing the data into multiple emotional grades according to a preset critical value; extracting keywords and topic fields based on word frequency statistics and topic clustering, and deleting the data if the number of keywords is lower than a threshold value; constructing time series data according to social behaviors and time stamps, predicting an active period and a behavior trend by using an LSTM model, and adjusting the model through a sliding window mechanism; and combining multi-dimensional feature vectors, and assigning personalized labels to the interactive object through clustering results. The application realizes accurate analysis of behavior and emotional characteristics of the interactive object, and provides an effective personalized interaction strategy for the social robot.
Owner:THE FIRST RES INST OF MIN OF PUBLIC SECURITY

Rumor detection methods, devices, equipment and storage media

This application relates to the field of text processing technology, and provides a method, apparatus, device, and storage medium for rumor detection. The method includes: acquiring text information to be detected; extracting text features and word frequency statistical features from the text information; determining the credibility of the text information based on the text features, and determining the fusion degree of the text features and the word frequency statistical features based on the credibility; fusing the text features and the word frequency statistical features based on the fusion degree to obtain fused features; and detecting whether the text information contains rumor information based on the fused features. The rumor detection method provided by this application determines the fusion degree of word frequency statistical features and text features through the credibility of the text information, achieving adaptive feature fusion. This effectively prevents overfitting caused by excessive use of auxiliary features, and adaptively provides necessary auxiliary features for rumor detection based on the credibility of the text information, thereby improving the accuracy of rumor detection.
Owner:CHINA THREE GORGES UNIV

An exaggeration representation word extraction method for Chinese irony text

The application discloses an exaggeration representation word extraction method for Chinese irony text, and belongs to the natural language processing technology, and comprises the following steps: step 1: after the irony data set is preprocessed, a bidirectional maximum matching method is used for word segmentation; step 2: the word frequency of the segmented text is calculated by using TF-IDF to construct a candidate word set; step 3: the chi-square statistic is used to measure the correlation degree between the irony text and the exaggeration representation, and the best threshold is set by the chi-square test method to select the strongly correlated exaggeration representation word, so as to construct an exaggeration representation seed word set; step 4: based on the WoBERT semantic similarity calculation framework, the dynamic word vector semantic similarity of the irony text and the seed word set is calculated, and the threshold is set to select the exaggeration representation word with high similarity, so as to construct an exaggeration representation word set. The application aims to extract the words containing the exaggerated expressions in the Chinese irony text to mine the characteristics of the irony sentences, so as to provide technical support for the Chinese irony text recognition task.
Owner:ANHUI UNIV OF SCI & TECH

Examination and approval person determination method and device, equipment, medium and program product

The embodiment of the invention discloses an approver determination method and device, equipment, a medium and a program product. Comprising the steps of obtaining a to-be-approved document; determining an entropy value of each to-be-selected title and dispersion and word frequency of each to-be-selected title relative to each title dictionary; according to the entropy value of each to-be-selected title and the dispersion and word frequency of each to-be-selected title relative to each title dictionary, determining a title optimization weight of each to-be-selected title; performing weighted optimization on the name to be selected and each candidate official document content feature data in the historical official document database through the name optimization weight, determining a first feature vector and a second feature vector, and determining a cosine similarity between the first feature vector and each second feature vector; and determining the approver associated with the target historical official document corresponding to the candidate official document content feature data associated with the maximum cosine similarity as the target approver of the official document to be approved. The selection accuracy of the approver is improved, and the approval process is more intelligent.
Owner:CHINA MOBILE (XIONGAN) ICT CO LTD +3

Multi-source heterogeneous log anomaly analysis method based on large language model

The invention discloses a multi-source heterogeneous log anomaly analysis method based on a large language model. Multi-source heterogeneous logs from different systems, formats and structures can be processed at the same time. Firstly, an automatic log analyzer is adopted, data of various different systems can be accepted and processed, log noise is removed through data cleaning, variables are replaced, original semantics are reserved, grouping and sorting are conducted based on the length, statistics is conducted on the frequency of tokens in each log, tuple vectors and word frequency vectors are constructed, threshold values are calculated, and unstable tokens are replaced; and carrying out similar template combination by adopting a mode of combining similarity and a prefix tree. And semantic, counting and sequence features are extracted and fused to uniformly represent multi-source heterogeneous logs. And performing anomaly detection by adopting a CNN-LSTM model of a double attention mechanism, mapping an anomaly result back to an original log, and generating an anomaly analysis result by utilizing an instruction-based large language model, thereby realizing multi-source heterogeneous log anomaly analysis processing based on the large language model.
Owner:BEIJING UNIV OF POSTS & TELECOMM

An automatic text classification method based on automobile community

ActiveCN117312561BEasy to analyze and manageImplement classificationFeature vectorData set
The application relates to an automatic text classification method based on a car community, which comprises the following steps: obtaining a data set of car community texts; obtaining word vectors and text feature vectors as input values of a double-layer clustering model; performing clustering calculation on the word vectors and the text feature vectors to generate word classification and text classification respectively, so as to form the double-layer clustering model. When a new text enters: calculating the word vectors and the text feature vectors of the new text; calculating the word classification of each word vector of the new text and the word frequency under each classification; calculating the text classification of the new text; performing dynamic analysis according to the word classification, the word frequency and the text classification of each word vector of the new text; when the number of words and the word frequency of the new text generated outside the existing word classification reach a threshold value, the double-layer clustering model is updated. Therefore, the application can realize automatic classification of car community texts in the whole process, improve the classification accuracy and efficiency, and form a closed-loop management.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

An ai semantic analysis and data processing method based on a large language model

The application relates to the field of data processing, in particular to an AI semantic analysis and data processing method based on a large language model. First, professional terms and context windows in vertical field texts are extracted, hidden layer vectors are extracted by using the large language model, and spatial dispersion is calculated, and the terms are divided into three-state working modes according to semantic stability. For polysemous terms, independent semantic branches are identified through vector clustering, co-occurrence word matching and vector distance determination are used to realize accurate semantic attribution in a new context, and dynamic segmented adjustment of static term frequency (TF-IDF) weight is carried out accordingly, and finally the improved weight is injected into the downstream analysis process. The application gives the same term different weight expressions in different contexts, and the low-overhead cascade disambiguation mechanism efficiently solves the polysemy problem of terms, avoids confusion of structured data and error clustering of documents, and significantly improves the accuracy of text matching, information extraction and document classification and the like.
Owner:ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1

Question detection model for call transcript

Disclosed are some implementations of systems, apparatus, methods and computer program products for categorizing a sentence as a question. Rather than using a single model, several different models are leveraged to determine whether a sentence is a question. For example, the models can include an inverse text normalization (ITN) model, a sentence embeddings model, and a Term frequency inverse document frequency (TFIDF) model. The output of an ITN model is processed using a finite state transducer (FST) while the output of the sentence embeddings model and TFIDF model are processed using logistics regression (LR) models. A support vector machine (SVM) is then applied to the output of the FST and LR models to determine whether the sentence is a question.
Owner:SALESFORCE INC

A clustering-based text supervision checking method, device and computer equipment

The application relates to a clustering-based text supervision checking method and device and a computer device. Feature word sets corresponding to a plurality of question description texts in a question list are obtained by respectively extracting feature words from the question description texts. The word frequency of each feature word in the question description text, the first inverse document frequency in all question description texts and the second inverse document frequency in all checked units are respectively calculated to obtain a feature word weight set of the question description text of each checked unit. A plurality of question description text clusters and a plurality of initial high-correlation-degree feature word sets corresponding to the question description text clusters are obtained through clustering. Finally, the initial high-correlation-degree feature word set is optimized according to a chi-square statistic to obtain an accurate high-correlation-degree feature word set. The method saves the calculation resources while ensuring the accuracy of the results by optimizing the feature word weight calculation, adopting a clustering algorithm for preliminary clustering and optimizing the clustering results through chi-square statistics.
Owner:NAT UNIV OF DEFENSE TECH