Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

74 results about "Participle" patented technology

A participle (PTCP) is a form of a verb that is used in a sentence to modify a noun, noun phrase, verb, or verb phrase, and plays a role similar to an adjective or adverb. It is one of the types of nonfinite verb forms. Its name comes from the Latin participium, a calque of Greek μετοχή (metokhḗ) "partaking" or "sharing"; it is so named because the Ancient Greek and Latin participles "share" some of the categories of the adjective or noun (gender, number, case) and some of those of the verb (tense and voice).

Analysis method and device for credit granting approval process

The invention provides a credit granting approval process analysis method and device. Comprising the steps of performing word segmentation on text content in a credit granting approval document set, and combining the obtained segmented words to obtain combined phrases; calculating weighted word frequency-inverse document word frequency of the combined phrases; target phrases with weighted word frequency-inverse document word frequency larger than a word frequency threshold value are screened out from the combined phrases, and an amplified word list is generated according to the target phrases and the original word list; training a target word segmentation device based on the amplified word list; finely adjusting the basic vector model according to the training corpus to obtain a term vector model; analyzing the auditing opinion information based on a large language model to obtain an initial risk analysis rule containing a preset dimension; based on a similarity algorithm, a target word segmentation device and a term vector model, grouping and integrating the initial risk analysis rules to obtain risk analysis rules; and processing the information of each examination and approval stage by using a big language model and adopting a risk analysis rule to obtain examination and approval suggestion information.
Owner:MINSHENG BANKING CORP

Machine translation differential test method for multi-word expression

The invention provides a multi-word expression-oriented machine translation differential test method aiming at the problem of inaccurate multi-word expression semantic translation in a mainstream machine translation system. The method comprises the following steps that a word segmentation tool based on deep learning is adopted to divide words into vocabulary units, syntactic labels are distributed in combination with a pre-training sequence marking model, and a dependency analysis tool spaCy is utilized to mark the syntactic relation between the words; converting the tagged corpus into a standard CoNLL format, extracting a multi-word expression of a sentence through an automatic tool, and establishing a test data set of a sentence-level and phrase-level corresponding relationship; inputting the test set into a multi-translation system to generate a translation, and using an alignment tool AWESOME to accurately locate a corresponding relationship between a source language and a target language MWEs; the translation similarity is calculated based on BERTScore, mistranslation, translation omission and non-translation are recognized through an intra-group and inter-group dual check mechanism in combination with a dynamic threshold value, and evaluation of the translation accuracy of machine translation on multi-word expression is completed. According to the method provided by the invention, multi-word expression translation errors can be accurately recognized, and the accuracy of phrase-level semantic translation of a machine translation system is finely evaluated through a differential test method.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method, device and equipment for processing wrongly written characters

The embodiment of the invention discloses a wrongly written character processing method, device and equipment, and the method comprises the steps: receiving text data of a target domain input by a user, the text data comprising special words of the target domain; performing word segmentation processing on the text data through a preset word segmentation strategy aiming at a target field, and performing identification processing aiming at special words of the target field on segmented words corresponding to the obtained text data to obtain target special words in the segmented words corresponding to the text data; carrying out wrongly written character recognition processing on the target proprietary word through a proprietary word error recognition model for the target field to obtain a first recognition result, and carrying out wrongly written character recognition processing on segmented words, except the target proprietary word, in the segmented words corresponding to the text data through a general word error recognition model to obtain a second recognition result; and based on the first recognition result and the second recognition result, determining wrongly written character information contained in the text data through fusion processing of the two different recognition results.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

A pronunciation prediction method and related device

The present application discloses a pronunciation prediction method and related devices. First, the text to be synthesized is segmented to obtain a segmentation sequence. For the first category of words in the segmentation sequence, the pronunciation information of the first category of words is determined based on a preset corpus resource library. For the second category of words other than the first category of words in the segmentation sequence, their pronunciation categories are determined based on the part-of-speech information of each word in the segmentation sequence, and their pronunciation information is determined based on the pronunciation information determination method corresponding to their pronunciation category. In the present application, combined with the corpus resource library and the preset pronunciation information determination method corresponding to each pronunciation category, it is possible to cover the pronunciation information determination in various situations. Therefore, it is possible to accurately determine the pronunciation information of each word in the text to be synthesized, thereby improving the effect of speech synthesis.
Owner:合肥智能语音创新发展有限公司

A paraphrase sentence recognition method and system based on semantic primitive knowledge and abstract semantic representation

The application belongs to the field of natural language processing, and particularly relates to a method and system for paraphrase recognition based on semantic primitive knowledge and abstract semantic representation, which comprises the following steps: performing word segmentation on a sentence, and performing word-level vector representation and semantic primitive knowledge representation; performing mean value processing on the semantic primitive knowledge representation result, and extracting interactive attention feature information of the mean value processing result by using global semantic information to obtain global semantic primitive representation; performing abstract semantic analysis on a to-be-recognized paraphrase sentence from a sentence structure to obtain a single-root directed acyclic graph, and performing global semantic primitive representation and word-level vector representation; extracting global and local feature information in the order of the directed acyclic graph, and performing distance feature measurement on the information; inputting the distance feature measurement result into a neural network to obtain a recognition result; the application introduces external semantic primitive knowledge to perform semantic representation, the accuracy of the semantic primitive knowledge representation is assisted by global semantic information, and the abstract semantics of a Chinese paraphrase sentence is analyzed to obtain semantic relations, so that the accuracy of paraphrase recognition is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Large model-based work order integration method, apparatus and device, and storage medium

The invention discloses a work order integration method and device based on a large model, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining a target field corresponding to each initial work order, and dividing each initial work order into a plurality of work order clusters according to each target field; performing initial word segmentation processing on each initial work order based on a preset noun knowledge base to obtain a work order after word segmentation, performing vocabulary analysis on each work order after word segmentation, and constructing a target word segmentation table according to a corresponding analysis result; and performing word segmentation processing again on the work orders after word segmentation according to the target word segmentation list to obtain target work orders, refining the target work orders based on the number of the target work orders in the work order clusters and the work order refining large model to obtain corresponding refining results, and integrating the refining results to obtain a refined work order. And obtaining a question list corresponding to each question main body. And through two times of word segmentation processing, the reliability of a work order content extraction result is improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

A method and device for calculating word meaning similarity based on adjacent word features

ActiveCN116522949BThe calculation result is accurateSemantic analysisEnergy efficient computingSentence processingPart of speech
This invention provides a method for calculating semantic similarity based on adjacent word features, relating to the field of natural language processing. First, in the example word extraction module, example words are extracted from the example sentence and the sentence to be matched using the longest common substring algorithm. Second, in the example sentence processing module, word segmentation is performed on the example sentence excluding the example words, and part-of-speech tagging is applied to the example words. Then, in the feature extraction module, features are extracted from the words surrounding the example words in the example sentence. The correlation between the surrounding words and the example words and option words is calculated using a corpus. The calculated results are weighted and combined with their respective features to form the features of the example words and option words. Finally, in the option processing module, the part-of-speech of the option words is compared with that of the example words, and a similarity score is calculated based on the features of the option words to select the optimal option word. This invention determines the features of the current word by extracting features from surrounding words and calculating correlation; it uses a corpus to calculate the mutual information between two words to achieve the correlation between the two words.
Owner:KUNMING UNIV OF SCI & TECH

Determining semantic and grammatical correctness of user-expanded sentence using integrated programmatic and specialized guided and constrained artificial intelligence

A system and method guide an Artificial Intelligence engine to determine the semantic and grammatical correctness of a user-expanded sentence in real-time. The sentence validation process involves receiving input from the user, the input includes sentence fragment that the user wishes to expand and user-expanded sentence that the user constructs on the fragment provided. The inputs are broken down into tokens. The word-level tokenization algorithm is used, which identifies tokens by splitting the text into spaces, punctuation marks, and other delimiters. Further, a token comparison algorithm is used to assess the relationship between the sentence fragment and the user-expanded sentence to analyze order and placement. Once the token comparison is complete, a prompt is generated using prompt generator to evaluate grammatical and semantic evaluation of the user-expanded sentence. Real-time feedback is provided to the user based on grammatical and semantic evaluation.
Owner:2HR LEARNING INC

Foreign affair table intelligent filling error correction method and system based on natural language processing

The invention relates to the technical field of data processing, in particular to an intelligent foreign form filling and error correction method and system based on natural language processing. The method comprises the following steps: acquiring text images of a historical corpus and vocabularies; screening segmented words to be corrected; determining the ambiguity of the module based on the frequency proportion of the input vocabulary and the vocabulary type; determining a first possibility based on the co-occurrence probability of the adjacent vocabularies of the to-be-corrected segmented word and the similar segmented word and the similarity of the to-be-corrected segmented word and the similar segmented word; determining a second possibility through an error correction result corresponding to the similar word segmentation history and the image similarity of the word segmentation to be corrected and the similar word segmentation; and determining the total possibility and the necessity of word segmentation error correction through the first possibility, the second possibility and the ambiguity, and completing word segmentation error correction. The error correction accuracy is improved.
Owner:BEIJING DEXUN AVIATION SERVICE CO LTD

Medical automatic question and answer method and system based on common sense fusion

The application provides a medical automatic question and answer method and system based on common sense fusion, comprising: performing word segmentation on training sentences, querying a common sense database and a knowledge base to obtain entity relationship triples; fusing and encoding the triples and the training sentences; randomly selecting part of entities of the fused and encoded training sentences to mask, replacing the next sentence into other random sentences according to a fixed probability, and inputting the obtained training corpus into a multi-layer partition encoder for training; performing word segmentation and entity relationship query on question sentences in question and answer data, and fusing and encoding; taking common sense fusion encoding sequences of the questions as model inputs, taking answers as supervision labels, training a common sense fusion language model; and building a visual medical automatic question and answer system, inputting questions into the model through a front end, and displaying outputs of the model as answers to an interface. The application has no type limitation on user questions, and fuses common sense and medical knowledge into a language model, thereby ensuring that answers conform to grammatical rules and improving professionalism.
Owner:SHANGHAI JIAOTONG UNIV

Intelligent customer service interactive response method and system based on semantic analysis

The present application relates to the field of semantic analysis technology, and specifically to a method and system for intelligent customer service interactive response based on semantic analysis. The method comprises: obtaining historical data of intelligent customer service and voice data currently input by the user, and converting the voice data into text sentences; preprocessing the text sentences and performing word segmentation; obtaining the average phrase length of the sentence containing each phrase, and obtaining the phrase rationality of each phrase; obtaining the mean continuous length of phrases with the same part of speech in the sentence containing each phrase, and obtaining the unpopularity confidence of each phrase; obtaining the semantic rationality value of each phrase; obtaining the segmentation threshold of the semantic rationality value, and judging whether each phrase is interfered with; evaluating the current text sentence, obtaining an accurate text sentence; and completing the intelligent customer service interactive response. The present application improves the accuracy of semantic analysis by obtaining more accurate text sentences from the user.
Owner:BEIJING WEIHEGUANG DIGITAL TECH CO LTD

A large language model machine translation optimization method and system

This invention relates to a method and system for optimizing machine translation using a large language model, belonging to the field of machine translation technology. It includes: generating a corresponding first translation from a source sentence in a bilingual corpus; segmenting the source sentence into words, counting the frequency of easily misspelled words in each segment, and calculating the easily misspelled word score, obtaining an easily misspelled word set based on the score; calculating the semantic similarity between the sentence to be translated and multiple candidate examples, and calculating the quality scores of the multiple candidate examples; selecting the optimal k candidate examples from the multiple candidate examples based on semantic similarity and quality scores, and constructing them as prompt templates; obtaining training sentence pairs from the bilingual corpus based on the easily misspelled word scores to construct a training set; and using the training set and prompt templates to perform low-rank adaptive training on a large language model to obtain an optimized large language model. This invention not only improves translation quality but also enhances the interpretability of the translation process.
Owner:SUZHOU UNIV

An exaggeration representation word extraction method for Chinese irony text

The application discloses an exaggeration representation word extraction method for Chinese irony text, and belongs to the natural language processing technology, and comprises the following steps: step 1: after the irony data set is preprocessed, a bidirectional maximum matching method is used for word segmentation; step 2: the word frequency of the segmented text is calculated by using TF-IDF to construct a candidate word set; step 3: the chi-square statistic is used to measure the correlation degree between the irony text and the exaggeration representation, and the best threshold is set by the chi-square test method to select the strongly correlated exaggeration representation word, so as to construct an exaggeration representation seed word set; step 4: based on the WoBERT semantic similarity calculation framework, the dynamic word vector semantic similarity of the irony text and the seed word set is calculated, and the threshold is set to select the exaggeration representation word with high similarity, so as to construct an exaggeration representation word set. The application aims to extract the words containing the exaggerated expressions in the Chinese irony text to mine the characteristics of the irony sentences, so as to provide technical support for the Chinese irony text recognition task.
Owner:ANHUI UNIV OF SCI & TECH

Text generation content security filtering method and system based on semantic detection

The invention discloses a text generation content security filtering method and system based on semantic detection, and the method comprises the steps: obtaining and preprocessing an original content text generated based on the text to obtain a preprocessed content text, carrying out the sentence segmentation of the preprocessed content text to obtain a plurality of sentences, and determining the overall semantic features of the sentences in the context based on semantic detection; performing word segmentation on the sentence to obtain a plurality of vocabularies, and determining specific semantic features of each vocabulary in the sentence based on semantic detection; and evaluating the content security of each sentence based on the overall semantic feature and the specific semantic feature to obtain a comprehensive security evaluation value, determining the security filtering level of each sentence based on the comprehensive security evaluation value, and performing security filtering on each sentence in the original content text. According to the method, security detection is carried out on the text generation content through a multi-level and fine-grained semantic analysis and evaluation technology, complex harmful content which is difficult to detect through a traditional method can be effectively recognized, and accurate risk grading and intelligent flexible processing are achieved.
Owner:LUSHAN COLLEGE OF GUANGXI UNIV OF SCI & TECH

Text topic determination method and device and storage medium

The embodiment of the invention provides a text topic determination method and device and a storage medium, which can be applied to the technical field of natural language processing, in any segmented word in the method, according to the transmission weight of the segmented word and the underlying topic in a segmented word transmission matrix, a corresponding underlying topic vector when the segmented word vector of the segmented word is transmitted to the underlying topic is determined, and the underlying topic vector of the segmented word is transmitted to the underlying topic. A bottom-layer topic vector is used as an embedding anchor, the transmission direction of a plurality of segmented words is guided, key information is highlighted, and generation of a coherent fine-grained topic is facilitated; and for any (i-1) th layer of theme, according to the transmission weights of the (i-1) th layer of theme and the ith layer of theme in the theme transmission matrix, determining the corresponding ith layer of theme vector when the (i-1) th layer of theme vector of the segmented word is transmitted to the ith layer of theme, guiding the low-layer theme to be transmitted to the high-layer theme among different layers of themes, and realizing alignment of cross-layer themes, therefore, all layers of themes are kept consistent and hierarchical progressive semantically, and a hierarchical theme structure with interpretability and coverage is obtained.
Owner:WEBANK (CHINA) +1

Method, apparatus, electronic device, and medium for determining term relationships

The present disclosure provides a method, device, electronic equipment and medium for determining term relationship, relates to the field of data processing and artificial intelligence, and particularly relates to the field of natural language processing and knowledge graph. A computer-executed method for determining term relationship can include: obtaining a plurality of words, the plurality of words being obtained by performing word segmentation processing on a first term from a corpus; determining one or more dependency relationships between the plurality of words; based on at least one dependency relationship of the one or more dependency relationships, constructing a second term according to at least two words of the plurality of words; and in response to determining that the second term is in the corpus, determining the second term as a sub-concept of the first term.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

A text labeling method, system and storage medium

The embodiment of the application provides a text labeling method, system and storage medium, comprising: obtaining a syntax rule set; dividing syntax rules in the syntax rule set into a first rule subset and a second rule subset; obtaining a Chinese text to be labeled; performing sentence division processing on the Chinese text to be labeled to obtain a sentence set; performing word segmentation processing on each initial sentence in the sentence set to obtain a corresponding word set; traversing the syntax rule set, when the word set corresponding to a sentence matches a first verification condition of any syntax rule, generating a candidate syntax set; performing structure verification based on a second verification condition of a syntax candidate to obtain a screening set; when the syntax candidates in the screening set belong to the same rule subset and there is a preset mutual exclusion relationship between the syntax candidates, performing decision verification based on a decision condition corresponding to the rule subset to which the syntax candidate belongs to generate an actual syntax set; and performing syntax labeling on the Chinese text to be labeled based on the actual syntax set.
Owner:SHANGHAI YOULI NETWORK TECHNOLOGY CO LTD

English long sentence analysis method and system based on hierarchical perception

PendingCN122334237AQuestion analysisSentence analysis
This invention discloses a hierarchical perception-based method and system for analyzing complex English sentences, relating to the interdisciplinary fields of natural language processing and intelligent education. By defining structural delimiters, sentences are divided into syntactic levels according to the number of effective delimiters, first eliminating invalid parallel structural delimiters; then, a syntactic encoding module completes text segmentation, hierarchical tagging, word embedding, and feature fusion; a hierarchical perception module calculates feature weights and generates target syntactic features; finally, a syntactic decoding module outputs the sentence's main body and modifier decomposition results. The hierarchical framework constructed by this invention fully covers the core modifiers of English, is suitable for the scenario of complex sentences in the College Entrance Examination (Gaokao), and can achieve accurate and lightweight decomposition of all types of complex sentences. It effectively solves the problems of irregular decomposition, poor generalization, and poor adaptability of traditional methods, and the analysis process is clear and teachable.
Owner:SHENZHEN TAITAIGE TECHNOLOGY CO LTD

Semantic recognition method and device and storage medium

The invention relates to a semantic recognition method and device and a storage medium. The semantic recognition method comprises the steps of performing word segmentation processing on a to-be-recognized text to obtain N to-be-recognized words of the to-be-recognized text and a word order among the N to-be-recognized words, N being an integer greater than or equal to 1; character features of each to-be-recognized word in the N to-be-recognized words are recognized, N target words identical to the N to-be-recognized words are determined in the multiple preset words on the basis of the word order and the character features, the N to-be-recognized words correspond to the N target words in a one-to-one mode, and the target words correspond to preset semantic recognition results; and determining a preset semantic recognition result corresponding to the N target words based on the N target words, and taking the preset semantic recognition result as a semantic recognition result of the to-be-recognized text. Through the semantic recognition method disclosed by the embodiment of the invention, less processing resources can be occupied, and better universality can be achieved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Word segmentation method and device, equipment and storage medium

The invention relates to the technical field of natural language processing, and discloses a word segmentation method and device, equipment and a storage medium. The method comprises the following steps: for any sentence in a text, segmenting the sentence to obtain a plurality of word segmentation results corresponding to the sentence; for any word segmentation result in the plurality of word segmentation results, according to a word vector of any word segmentation in the word segmentation result, based on a frequency domain vector converted by the word vector and a font vector of the word segmentation, determining fusion features corresponding to the word segmentation; wherein the frequency domain vector of the segmented word is obtained by processing the word vector of the segmented word based on Fourier transform, and the font vector of the segmented word is obtained according to the stroke of the segmented word; and determining a target word segmentation result corresponding to the sentence from the plurality of word segmentation results according to the fusion feature corresponding to each word segmentation in each word segmentation result in the plurality of word segmentation results. Therefore, the accuracy of the word segmentation result can be improved.
Owner:WEBANK (CHINA)

A Chinese measure word correction method, system, computer device and storage medium

PendingCN122113908ANatural language data processingMeasure wordAlgorithm
The application relates to the technical field of natural language processing, and discloses a Chinese numeral quantifier correction method and system, computer equipment and a storage medium. The application constructs a dynamically expandable collocation word table and a numeral dictionary; performs sentence segmentation, word segmentation and part-of-speech tagging on input text, and identifies a three-element structure composed of numerals, quantifiers and nouns; excludes misjudgments on inherent expressions through a fixed phrase filtering mechanism; performs legality verification on the structure based on the collocation word table, and generates processing information for semantic correction if the verification fails; when collocation abnormalities are confirmed and the dictionary is insufficient, a large language model is triggered to replace the quantifier and adapt to the context based on the processing information, so that a final corrected sentence is generated. Through multi-layer cooperation of rule matching, word table verification and large model verification, the application effectively reduces the mis-correction rate while ensuring high-precision correction, and improves the accuracy and practicality of automatic correction of Chinese quantifiers.
Owner:山东齐鲁壹点传媒有限公司 +1

An entity extraction method based on MRC framework

The application discloses an entity extraction method based on an MRC framework, which comprises the following steps: firstly, obtaining a target sentence according to a device maintenance manual, generating a corresponding question according to the definition of an entity type, and splicing the target sentence and the question to obtain a corpus; then, performing word segmentation on the corpus by using a word segmentation tool, inputting the corpus into a BERT model after coding to obtain word embedding representation of the target sentence; secondly, obtaining sentence-level features of the target sentence by a sentence classification module; then, combining the sentence-level features and the word embedding representation of the target sentence to integrate into an entity extraction module; finally, combining the sentence classification module and the entity extraction module, training the two modules together, and completing entity extraction according to the two trained modules. The application can use the information at the sentence level in the entity extraction task, which helps to improve the precision of entity extraction and solves the problem of entity extraction in the device maintenance document.
Owner:ZHEJIANG UNIV

A Chinese word segmentation method, device and storage medium

The application discloses a Chinese word segmentation method, device and storage medium, and belongs to the technical field of natural language processing. The Chinese word segmentation method comprises the following steps: S1, a second language translation sentence of a to-be-detected sentence is acquired; S2, a Chinese Bert pre-training language model is used to code the to-be-detected sentence, so as to acquire vector representation of semantic information of the whole sentence and a sentence vector representation sequence; S3, a second language Bert pre-training language model is used to code the translation sentence, so as to acquire vector representation of semantic information of the whole sentence; S4, semantic features of the to-be-detected sentence and the translation sentence are fused, so as to acquire a predicted category of each word of the to-be-detected sentence; and S5, the to-be-detected sentence is segmented according to the predicted category, so as to acquire a word segmentation result. The method improves the accuracy of word segmentation, and has a good word segmentation effect on foreign words in particular.
Owner:ZHONGKE FANYU (WUHAN) TECH CO LTD

Word weight ranking method, device and equipment and storage medium

The application relates to the technical field of artificial intelligence, and discloses a word weight sorting method, device and equipment and a storage medium, the sorting method comprising the following steps: obtaining a first target sentence, performing word segmentation on the first target sentence to obtain a first vocabulary set contained in the first target sentence; removing each first vocabulary in the first vocabulary set from the first target sentence to obtain a second target sentence; forming a sentence pair by combining any second target sentence and the first target sentence, inputting the sentence pair into a pre-trained language model to generate a vector representation corresponding to the sentence pair; determining a target similarity between the second target sentence and the first target sentence in the sentence pair according to the vector representation corresponding to the sentence pair; sorting the target similarities of the second target sentences and the first target sentence; and determining a weight order of each first vocabulary according to the sorting of the target similarities. The application solves the problem that the word weight cannot be accurately sorted in the prior art.
Owner:PING AN TECH (SHENZHEN) CO LTD

A sentence scanning method and device and a storage medium

Embodiments of the present application disclose a sentence scanning method and device, and a storage medium. The method comprises: obtaining word elements of a teaching material file and an element information table establishing a mapping relationship with a single word element according to a current learning teaching material file of a user, the element information table comprising sentence information of all sentences containing the single word element in the teaching material file; obtaining a scanning field recognized by a scanning pen, performing word segmentation processing on the scanning field, and obtaining one or more key words contained in the scanning field; calling the element information table of the same word element as the key word, obtaining the sentence information of the key word; determining a sentence to which the scanning field belongs according to the sentence information of the key word, and playing or displaying the sentence. By using the above technical means, the problem that the existing scanning pen cannot play or translate a complete sentence according to a part field of the sentence is solved, and the user experience is improved.
Owner:DONGGUAN ELF EDUCATIONAL SOFTWARE CO LTD

Problem expansion method and device, electronic equipment and computer readable storage medium

The application relates to an artificial intelligence technology and discloses a question expansion method and device, equipment and a storage medium. The method comprises the following steps: extracting an inquiry mode word and an entity noun in a to-be-expanded question; extracting a standard question containing the inquiry mode word and a synonymous question with the same meaning as the standard question in a question and answer library, and marking the inquiry mode word and the inquiry mode word contained in the synonymous question as a keyword; performing word segmentation and part-of-speech tagging on the standard question and the synonymous question to obtain a sentence structure; extracting a front adjacent substantive word and a rear adjacent substantive word of the keyword in the standard question and the synonymous question according to the sentence structure; extracting the inquiry mode word in the keyword, and extracting the front adjacent substantive word and the rear adjacent substantive word; and according to the keyword, the front adjacent substantive word, the rear adjacent substantive word and the proper noun, composing a preset number of expansion questions in a preset grammar format. The application can automatically generate expansion questions according to an input question.
Owner:CHINA MERCHANTS FINANCE HLDG CO LTD

A text encoding method

The application discloses a text coding method, which comprises the following steps: data cleaning and preprocessing of a corpus, sentence division and word division of the corpus, extraction of relevant features, coding representation of the sentence, addition of three dimensions on the basis of coding, and construction of an upper model; compared with the prior art, the difference of the application lies in that the previous additional features are mostly intra-sentence features, such as part of speech, syntax dependency, relative position, absolute position and the like, while the application adds global statistical information, the overall operation is simple, the training cost is low, the prior knowledge injection based on global statistical information is carried out on the sentence level coding, and the accuracy of the upstream task is improved.
Owner:SSE INFORMATION NETWORK LTD

Pre-training language model training method, text sentiment classification method and device

The application provides a pre-training language model training method, a text sentiment classification method and device, and relates to the technical field of natural language processing. The pre-training language model training method comprises the following steps: obtaining a training sample and sequence text information corresponding to the training sample; performing word segmentation processing on the sequence text information to obtain a plurality of segmented words included in the training sample; obtaining a basic feature vector of each segmented word in the plurality of segmented words and a semantic element feature vector of each segmented word, and splicing the basic feature vector of each segmented word and the semantic element feature vector of each segmented word to obtain an input text vector; processing the input text vector based on a masking task and a sentence relationship task in the pre-training language model to obtain a loss function of the pre-training language model, and training the pre-training language model based on the loss function. By adopting the technical scheme, the semantic representation between sentences can be improved, and the classification of the text sentiment can be accurately predicted.
Owner:MASHANG CONSUMER FINANCE CO LTD

Chinese and English literature grammar processing method based on mixed grammar model

The invention discloses a Chinese and English literature grammar processing method based on a mixed grammar model, which is characterized by comprising the following steps of: (1) segmenting an original literature text into words or phrases by using a word segmentation and language recognition module, and simultaneously recognizing that the language attribute of each word or phrase is Chinese or English; (2) inputting the text segmented into words or phrases into a mixed grammar model module for primary analysis; the mixed grammar model module comprises a rule sub-model and a statistical sub-model; (3) the results of the primary analysis are fused, and then total confidence Ctotal scoring is carried out; if Ctotal is greater than or equal to 0.95, the analysis result is credible; and if the Ctotal is less than 0.95, repeating the step (2) to carry out secondary analysis, tertiary analysis and the like to carry out optimization until the Ctotal is greater than or equal to 0.95. The method solves the problem that in the prior art, the literature grammar analysis precision is reduced when Chinese and English mixed long and difficult sentences and terminologies are used.
Owner:CHINA TOBACCO YUNNAN IND