Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Measure word" patented technology

In linguistics, measure words are words (or morphemes) that are used in combination with a numeral to indicate an amount of something represented by some noun.

A Chinese measure word correction method, system, computer device and storage medium

PendingCN122113908ANatural language data processingMeasure wordAlgorithm
The application relates to the technical field of natural language processing, and discloses a Chinese numeral quantifier correction method and system, computer equipment and a storage medium. The application constructs a dynamically expandable collocation word table and a numeral dictionary; performs sentence segmentation, word segmentation and part-of-speech tagging on input text, and identifies a three-element structure composed of numerals, quantifiers and nouns; excludes misjudgments on inherent expressions through a fixed phrase filtering mechanism; performs legality verification on the structure based on the collocation word table, and generates processing information for semantic correction if the verification fails; when collocation abnormalities are confirmed and the dictionary is insufficient, a large language model is triggered to replace the quantifier and adapt to the context based on the processing information, so that a final corrected sentence is generated. Through multi-layer cooperation of rule matching, word table verification and large model verification, the application effectively reduces the mis-correction rate while ensuring high-precision correction, and improves the accuracy and practicality of automatic correction of Chinese quantifiers.
Owner:山东齐鲁壹点传媒有限公司 +1

Text sensitive word library extraction method, device and equipment based on neural network model

The application relates to a text sensitive word library extraction method, device and equipment based on a neural network model. The method comprises the following steps: constructing a sensitive word library extraction model; the sensitive word library extraction model comprises a self-defined rule algorithm, a character-by-character segmentation algorithm, a Chinese word segmentation algorithm and an N-gram algorithm; the sensitive word library extraction model is pre-trained according to a training data set; the pre-trained sensitive word library extraction model is used for word segmentation extraction and word library quantitative analysis on the text to be extracted; the full-amount word library is subjected to field structure design; the pre-trained sensitive word library extraction model is iteratively trained according to the designed word library; after the sensitive word extraction is carried out by using the trained sensitive word library extraction model, data analysis is carried out according to a word segmentation frequency filtering rule, a word segmentation character quantity filtering rule and a word segmentation type filtering rule, and a sensitive word library is obtained. The method can improve the sensitive word extraction accuracy and efficiency.
Owner:HUNAN INKE INTERACTIVE ENTERTAINMENT NETWORK INFORMATION CO

Training data optimization methods and related devices for multi-vertical models

ActiveCN121859001Breduce vocabularysave spaceSemantic analysisMeasure wordData space
This application provides a method and related apparatus for optimizing training data of a multi-vertical category model. The method includes: constructing a full vocabulary and a word attribute library for each vertical category; performing multi-dimensional value evaluation on each word in the full vocabulary based on the word attribute library to obtain the comprehensive value of each word; optimizing the full vocabulary based on the comprehensive value of each word to obtain a high-value vocabulary for the vertical category; and merging and deduplicating the high-value vocabulary for each vertical category to obtain a hybrid vocabulary. In this way, the comprehensive value of words is quantified through their attribute information, thereby optimizing the independent vocabulary for each vertical category, filtering out low-value words, retaining high-value words, and finally merging and deduplicating the high-value vocabulary for each vertical category to obtain a hybrid vocabulary. This reduces the vocabulary size required for model training, significantly compresses the training data space, and lowers training costs.
Owner:SHENZHEN XISHIMA DATA TECH CO LTD

Rail transit event extraction method fusing semantic features and local dependency features

The rail transit event extraction method fuses the semantic features and the local dependency features, and comprises the following steps: step 1, obtaining original design specification texts for preprocessing to obtain event text data sets; step 2, enhancing the local dependency information of the event text data sets to obtain enhanced local dependency features; step 3, enhancing the semantic information of the event text data sets to obtain enhanced semantic features; step 4, fusing the enhanced local dependency features in step 2 and the enhanced semantic features in step 3 to generate a fused feature vector; and step 5, inputting the fused feature vector into a CRF model for prediction, and evaluating the model performance through various evaluation indexes. The method enhances the event semantic information by supplementing the measuring words and mining new trigger words, and fully utilizes the local dependency features of the events by constructing a dependency graph and an RGAT model, so that the precision of event extraction is significantly improved.
Owner:XIAN UNIV OF TECH

Bidding procurement text violation identification method, device and equipment and medium

PendingCN121457441ASemantic analysisBiological modelsMeasure wordText entry
The embodiment of the invention provides a bid procurement text violation recognition method and device, equipment and a medium. The method comprises the steps that a violation recognition large model is acquired, and a to-be-recognized text is determined; the violation identification large model comprises an encoder and a decoder; a to-be-recognized text is input into the violation recognition large model, multi-level embedding vectors of the to-be-recognized text are determined through an encoder, and the multi-level embedding vectors comprise a word-level embedding vector, a word-level embedding vector, a sentence-level embedding vector and a paragraph-level embedding vector; inputting the multi-stage embedded vector into a decoder; determining a hidden violation recognition result of the to-be-recognized text through a decoder according to the multi-level embedded vector; and according to the preset violation identification rule, determining the dominant violation identification result of the to-be-identified text, and according to the implicit violation identification result and the dominant violation identification result, outputting the violation identification result of the to-be-identified text, thereby effectively improving the identification efficiency, and further supporting the fairness and normalization of bid inviting and purchasing activities.
Owner:THREE GORGES HI TECH INFORMATION TECH CO LTD

Information processing method and device and information processing system

The invention provides an information processing method and device, relates to the technical field of artificial intelligence, in particular to the technical fields of natural language processing, voice processing, deep learning, large models and the like, and can be applied to scenes such as content generation of artificial intelligence. According to the specific implementation scheme, dialogue text data are determined based on obtained user dialogue information; matching the dialogue text data with template information in the dialogue template to obtain a matching result, the template information comprising at least one disease evaluation word; in response to a matching result that the dialogue text data is matched with the disease evaluation word in the template information, performing sentiment analysis on the dialogue text data based on a large language model to obtain a current sentiment parameter of the dialogue text data; based on the current emotion parameters and the historical emotion parameters of the user, the behavior prediction result of the user is predicted, and the safety and health of the patient are guaranteed.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Sentiment classification method and device based on prompt learning

The invention provides an emotion classification method and device based on prompt learning, and belongs to the technical field of artificial intelligence. The method comprises the steps that after a serialized first prompt template is obtained, the first prompt template is updated based on an input sample, and a second prompt template is obtained; on the basis of a second prompt template, the word embedding vector is coded to obtain a word vector, and the word embedding vector is obtained by coding the input sample; on the basis of a grammar dependency relationship matrix and an emotion polarity adjacency matrix, feature extraction is conducted on the word vectors, feature vectors are obtained, the grammar dependency relationship matrix reflects the dependency relationship between the words of the input samples, and the emotion polarity adjacency matrix reflects the emotion polarity between the words of the input samples; and on the basis of the feature vector, emotion category probability distribution corresponding to the input sample is calculated. According to the method, the iteration time of the prompt template can be shortened, the model fine tuning efficiency is improved, and the sentiment analysis accuracy is improved.
Owner:CHINA MOBILE M2M +1

Method and device for measuring word semantic similarity by integrating SAO and Bayesian model

ActiveCN119692358BSemantic analysisMeasure wordData mining
The present invention provides a method and apparatus for measuring the semantic similarity of words by integrating the SAO and Bayesian models, which relates to the technical field of calculating the semantic similarity of words. This method is an innovative method for measuring the semantic similarity of words, which integrates the extraction of the SAO (subject-action-object) structure and the Bayesian model. First, the SAO structure is extracted from the text, and then the occurrence frequencies of similar words are statistically counted based on these structures, which is used as a measure of the similarity of words. At the same time, semantic knowledge bases such as WordNet are also utilized to obtain the semantic similarity parameters of words, and these parameters and statistical results are integrated through the Bayesian model to calculate the posterior probability of the semantic similarity of words. The innovation of this method lies in that it not only utilizes the advantages of the statistical-based method but also combines the knowledge-based method, providing a new perspective to understand and quantify the similarity between words.
Owner:XIAMEN UNIV OF TECH

A morphologically enhanced tensorized word embedding compression system

The application discloses a morphological enhancement based tensorized word embedding compression system, which comprises a morpheme segmentation module, a morpheme index and embedding module and a word embedding generation module; the morpheme segmentation module divides each word in a word table of a text task into morphemes; the morpheme index and embedding module firstly generates a morpheme table by counting the division result of the morpheme segmentation module, then defines a morpheme index matrix and a plurality of trainable morpheme embedding matrices, each row of the morpheme index matrix represents the position of the morpheme of a corresponding word in the word table in the morpheme table, and each row of the morpheme embedding matrix represents an embedding vector of a corresponding morpheme in the morpheme table; the word embedding generation module indexes the morpheme vector from the morpheme embedding matrix and performs a tensor product on each word in the word table, adds the results of a plurality of tensor products to generate a word embedding vector; and the application overcomes the problems of a large amount of parameters and storage space occupation in general word embedding technology, and the problem of task effect loss when high compression word embedding is performed.
Owner:TIANJIN UNIV +1

A method, system and device for blind translation of Chinese combining rules

The application relates to the field of computer word processing, in particular to a Chinese blind translation method, system and device combined with rules. The method comprises the following steps: obtaining Chinese character data to be translated; inputting the Chinese character data to be translated into a large language model to obtain Braille translation; the large language model comprises a word embedding layer, a learnable rule layer, an encoding layer and a decoding layer; the Chinese character data to be translated is input into the word embedding layer to generate a word vector, and is input into the learnable rule layer to generate a rule vector; the word vector and the rule vector are fused to obtain a constraint enhanced feature vector; the constraint enhanced feature vector is input into the encoding layer to be encoded, and then is decoded through the decoding layer to obtain the Braille translation. The application combines Braille conversion rules, can solve the problem of converting Braille for multi-sound Chinese characters, and has good application value.
Owner:THE EYE HOSPITAL OF WENZHOU MEDICAL UNIVERSITY +1

Attribute category representation method and apparatus, terminal device, and storage medium

The application provides a representation method and device of an attribute category, terminal equipment and a storage medium. Attribute category data to be identified is obtained, the attribute category data to be identified comprises at least one attribute category; the encoding of each attribute category is found from a pre-configured attribute category encoding table; the encoding of each attribute category is input into a pre-established word vector table to obtain a representation vector of each attribute category; the word vector table is learned from a word vector embedding model, and the word vector embedding model is pre-trained on an Embedding module based on a self-attention mechanism method by using a masked attribute category encoding sequence sample. The method can quickly and effectively convert the input attribute category encoding into a vector by using the pre-established word vector table; and the converted output vector is low-dimensional and dense, which can greatly reduce vector operation and storage space.
Owner:GUANGZHOU HUADUO NETWORK TECH

Event detection method and device based on semantic network word representation and attention map

The application discloses an event detection method and device based on semantic network word representation and attention graph, comprising the following steps: splicing the word content vector, the word structure vector and the position feature vector of each word to generate the feature map of each sentence; combining the POS vector of each word to calculate the attention mechanism and generate the new feature map of each sentence; generating the sentence-level feature vector based on the new feature map; and obtaining the event detection result by using the splicing result of the sentence-level feature vector and the word content vector. The application comprehensively utilizes the external corpus, the semantic network, the part of speech and the attention graph, optimizes the features, more accurately extracts the trigger word, introduces more information, solves the polysemy problem, expresses the associated information between the synonymous words, and obtains more accurate event detection result.
Owner:NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT +1

Training data optimization method of multi-sag model and related device

ActiveCN121859001ASemantic analysisMeasure wordData space
The invention provides a training data optimization method for a multi-sag model and a related device, and the method comprises the steps: constructing a full quantity word list and a word attribute library corresponding to each sag field for the sag field; performing multi-dimensional value evaluation on each vocabulary in the full-quantity word list based on the word attribute library to obtain a comprehensive value of each vocabulary; based on the comprehensive value of each vocabulary, optimizing the full-quantity word list to obtain a high-value word list corresponding to the vertical class field; and carrying out combination and deduplication on the high-value word lists corresponding to each vertical class field to obtain a mixed word list. Thus, the comprehensive value of the vocabularies is quantified through the attribute information of the vocabularies, then the independent vocabularies corresponding to all the droop fields are optimized, the low-value vocabularies are screened out, the high-value vocabularies are reserved, finally, the high-value vocabularies of all the droop fields are subjected to merging and duplicate removal, duplicate data are removed, the mixed vocabularies are obtained, the vocabulary amount during model training is reduced, and the model training efficiency is improved. The training data space is fully compressed, and the training cost is reduced.
Owner:SHENZHEN XISHIMA DATA TECH CO LTD

Bidding text classification method and computer equipment

The invention discloses a bidding and tendering text classification method, which comprises the following steps of: obtaining each core word of a to-be-processed bidding and tendering text, constructing a graph structure by taking each core word as a node, and obtaining a feature matrix and a co-occurrence matrix; a first Transform network is adopted to process the feature matrix, and word-level coding features are obtained; determining each sentence vector, and processing each sentence vector, the word-level coding feature and the co-occurrence matrix by adopting a second Transform network to obtain a sentence-level coding feature; processing the sentence-level coding feature and the co-occurrence matrix by adopting a Transform variant to obtain a graph-level coding feature; fusing the word-level coding features, the sentence-level coding features and the graph-level coding features to obtain fused features; and processing the fusion features based on a multi-label text classification model, and outputting a multi-label classification result of the to-be-processed bidding and tendering text. According to the method, semantic information of different hierarchical features (words, sentences and graphs) of the bidding text can be captured, and the efficiency and accuracy of bidding text classification are improved.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD