Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

131 results about "Participle" patented technology

A participle (PTCP) is a form of a verb that is used in a sentence to modify a noun, noun phrase, verb, or verb phrase, and plays a role similar to an adjective or adverb. It is one of the types of nonfinite verb forms. Its name comes from the Latin participium, a calque of Greek μετοχή (metokhḗ) "partaking" or "sharing"; it is so named because the Ancient Greek and Latin participles "share" some of the categories of the adjective or noun (gender, number, case) and some of those of the verb (tense and voice).

AI reading control method and system based on artificial intelligence

The invention relates to the technical field of natural language processing, in particular to an AI reading management and control method and system based on artificial intelligence, and the method comprises the following steps: carrying out text word segmentation processing on an input original text, segmenting the text into independent sentences, recognizing basic word units in each sentence, analyzing semantic adjacency relationships and syntactic structure features among vocabularies, and carrying out word segmentation processing on the word units; screening and extracting potential phrases representing paragraph meanings, and establishing a candidate semantic unit set; according to the method, through word segmentation, syntactic structure recognition and semantic adjacency analysis of the original text, potential phrases capable of representing paragraph significance are extracted, the candidate semantic unit set is constructed, and modeling of the semantic structure in the text is achieved. And associating the set with the reading fixation duration and the playback action of the user sentence by sentence to obtain a reading behavior response of each semantic fragment, and executing semantic weighting and reading behavior cross analysis according to the reading behavior response. Through the linkage mode, the semantic focus actually focused by the user at present can be recognized.
Owner:SHENZHEN JOYAR SMART MFG TECH LTD

Conversational AI multi-intention understanding and decision-making method for tourism scene

The invention is suitable for the technical field of semantic understanding, and provides a dialogue type AI multi-intention understanding and decision-making method for a tourism scene, and the method comprises the steps: obtaining a tourism scene dialogue text which is input by a user and is expressed by a dialect; dialect expression is converted into standard Chinese statements through a dialect translation engine, and translation semantic deviation is corrected through a semantic correction model; performing multi-intention modeling on the converted standard Chinese statements, adopting a pre-training language model to combine with dynamic word segmentation processing, aligning special words of dialects to a standard Chinese word list through a dialect word mapping table, and adjusting semantic representation to adapt to dialect features; generating word list probability distribution, and generating a dialect compatible reply containing a multi-intention decision result in combination with an autoregression model; a dialogue state tracking model is trained according to geographic information, user historical interaction data and a cross-dialect corpus, intention switching and composite requests in multiple rounds of dialogues are analyzed, and dialect compatibility and intention recognition accuracy are effectively improved.
Owner:GUANGXI LVFA TECH CO LTD

Knowledge base and knowledge graph combined text analysis method and system

The invention provides a knowledge base and knowledge graph combined text analysis method and system.The method comprises the steps that for each unregistered vocabulary, similar fields related to the unregistered vocabulary are searched for in at least two word segmentation fragments of a word segmentation sequence, the similar fields in the word segmentation sequence are replaced with the unregistered vocabulary, and the unregistered vocabulary is obtained; obtaining an optimized word segmentation sequence; the method comprises the steps of obtaining an optimized word segmentation sequence, converting the optimized word segmentation sequence into an optimized word segmentation vector, calculating a first similarity based on the optimized word segmentation vector, and generating a summary text and a viewpoint list according to the similarity. According to the method, not only can unregistered vocabularies be automatically recognized, but also the most possible meaning of the unregistered vocabularies can be determined according to the context environment of the unregistered vocabularies; and explanations of different types of texts need to be dynamically adjusted, so that the consistency and accuracy of the meanings of the texts are ensured.
Owner:TIANJIN HUIZHIYINGJIN TECHNOLOGY CO LTD

System and method for automatically generating access control strategy based on multi-task learning

The invention relates to an access control strategy automatic generation system and method based on multi-task learning, and the method comprises the steps: carrying out the word segmentation, cleaning and embedded vector conversion of an original access control text through a data preprocessing module, and constructing a normative input format; the feature sharing layer module is used for extracting deep semantic features of a text through multi-layer bidirectional coding and an attention mechanism and providing unified representation for downstream tasks; the access control statement identification module is used for judging whether each sentence in the text is an access control statement or not and realizing automatic identification of strategy related contents; and the attribute extraction and annotation module is used for annotating words in the access control statements and extracting subject, object and operation access control attributes. A word coding layer and a sentence coding layer are shared, local and global attention mechanisms are combined, key information of a text is extracted, the semantic understanding ability is enhanced, and cooperative training of statement recognition and attribute extraction is achieved; a conditional random field CRF structure is used for sequence labeling, and the structural rationality of attribute labels is ensured.
Owner:SUZHOU UNIV OF SCI & TECH +1

Translation quality evaluation method and device, electronic equipment and storage medium

The invention discloses a translation quality evaluation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a to-be-evaluated task; for each to-be-evaluated translated sentence, evaluating the target to-be-evaluated translated sentence based on a target reference translated sentence set associated with the target source text statement to obtain a sentence level evaluation attribute corresponding to the target to-be-evaluated translated sentence; evaluating each to-be-evaluated segmented word in the target to-be-evaluated translated sentence based on a reference translated word in the target reference translated sentence set to obtain a word level evaluation attribute corresponding to each to-be-evaluated segmented word; and according to the sentence level evaluation attribute of each to-be-evaluated translated sentence and the word level evaluation attribute of each to-be-evaluated segmented word, generating a translated text evaluation report corresponding to the to-be-evaluated task. The effect of more objectively and accurately evaluating the overall translation quality of the translated text and the translation quality of the proprietary entity is achieved.
Owner:AGRICULTURAL BANK OF CHINA

Risk portrait generation method based on knowledge graph

The invention relates to the technical field of risk portrait generation based on a knowledge graph, and discloses a risk portrait generation method based on the knowledge graph. According to the scheme, the document is segmented through the sentence end symbols, semantic coherence is ensured, and the accuracy and stability of candidate word extraction are improved. Continuous Chinese character sequence segmentation is adopted, dependence of a word segmentation tool is avoided, the occurrence frequency and position of candidate words are recorded, and a detailed statistical basis is provided for risk factor screening. A minimum occurrence number constant is set, noise interference and information coverage are balanced, the screening flexibility is improved, and candidate word screening is more representative. And counting the co-occurrence relationship of candidate words in the sentence, quantifying the co-occurrence frequency through an indicator function, enhancing word association expression, and providing key data for initial graph construction. Finally, the scheme constructs a knowledge graph based on the candidate words and the co-occurrence relation thereof, ensures clear expression of risk information, and improves the accuracy of the system in risk identification and portrait generation.
Owner:SHANGHAI JIEJI INFORMATION TECHNOLOGY CO LTD

Malicious content detection method and device, electronic equipment and storage medium

The invention discloses a malicious content detection method and device, electronic equipment and a storage medium, which are used for solving the problem of low detection accuracy of the existing malicious content detection method, and the method comprises the following steps: obtaining a to-be-detected text; performing word segmentation processing on the to-be-detected text to obtain a word sequence corresponding to the to-be-detected text; the word sequence is input into a malicious content detection model, sensitive words in the word sequence and context information of the sensitive words are output, the malicious content detection model is obtained through training based on feature vectors, part-of-speech features and dependency syntax features of words in a training sample text, the part-of-speech features represent the types of the words, and the dependency syntax features represent the types of the words; the dependency syntactic features represent a syntactic association relationship between the word and other words in the sentence to which the word belongs. According to the malicious content detection model, on the basis of the features of the words of the training sample text, the part-of-speech features and dependency syntax features of the words are combined for training, key features of the text are highlighted, the context perception ability of the model to the malicious content is enhanced, and the detection accuracy of the malicious content is improved.
Owner:NSFOCUS INFORMATION TECHNOLOGY CO LTD +1

Analysis method and device for credit granting approval process

The invention provides a credit granting approval process analysis method and device. Comprising the steps of performing word segmentation on text content in a credit granting approval document set, and combining the obtained segmented words to obtain combined phrases; calculating weighted word frequency-inverse document word frequency of the combined phrases; target phrases with weighted word frequency-inverse document word frequency larger than a word frequency threshold value are screened out from the combined phrases, and an amplified word list is generated according to the target phrases and the original word list; training a target word segmentation device based on the amplified word list; finely adjusting the basic vector model according to the training corpus to obtain a term vector model; analyzing the auditing opinion information based on a large language model to obtain an initial risk analysis rule containing a preset dimension; based on a similarity algorithm, a target word segmentation device and a term vector model, grouping and integrating the initial risk analysis rules to obtain risk analysis rules; and processing the information of each examination and approval stage by using a big language model and adopting a risk analysis rule to obtain examination and approval suggestion information.
Owner:MINSHENG BANKING CORP

Machine translation differential test method for multi-word expression

The invention provides a multi-word expression-oriented machine translation differential test method aiming at the problem of inaccurate multi-word expression semantic translation in a mainstream machine translation system. The method comprises the following steps that a word segmentation tool based on deep learning is adopted to divide words into vocabulary units, syntactic labels are distributed in combination with a pre-training sequence marking model, and a dependency analysis tool spaCy is utilized to mark the syntactic relation between the words; converting the tagged corpus into a standard CoNLL format, extracting a multi-word expression of a sentence through an automatic tool, and establishing a test data set of a sentence-level and phrase-level corresponding relationship; inputting the test set into a multi-translation system to generate a translation, and using an alignment tool AWESOME to accurately locate a corresponding relationship between a source language and a target language MWEs; the translation similarity is calculated based on BERTScore, mistranslation, translation omission and non-translation are recognized through an intra-group and inter-group dual check mechanism in combination with a dynamic threshold value, and evaluation of the translation accuracy of machine translation on multi-word expression is completed. According to the method provided by the invention, multi-word expression translation errors can be accurately recognized, and the accuracy of phrase-level semantic translation of a machine translation system is finely evaluated through a differential test method.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

International trade contract intelligent auditing system based on natural language processing

The invention relates to the technical field of contract intelligent auditing, in particular to an international trade contract intelligent auditing system based on natural language processing, which comprises a data acquisition preprocessing module, a key vocabulary extraction module, a key vocabulary analysis screening module and a text similarity association module, the data acquisition preprocessing module carries out word segmentation and segmentation on a contract text and a regulation standard text, then the key word extraction module extracts a key word set in each text segment through a word frequency statistical method, and in order to prevent the situation that the key word set extracted by the key word extraction module is inaccurate, the key word set is extracted by the key word extraction module. The key vocabulary analyzing and screening module analyzes the paraphrases of the key vocabularies in the corresponding words, removes the key vocabularies which are not nouns in the key vocabulary set, provides a paraphrasing basis for the vocabulary same paraphrasing mapping unit, unifies same paraphrasing word vectors, and transmits the same paraphrasing word vectors to the key vocabulary same paraphrasing mapping unit; and the contract and the regulation text are deeply associated in terms of vocabulary semantics.
Owner:NANJING LUNAIZUN TECHNOLOGY CO LTD

Method, device and equipment for processing wrongly written characters

The embodiment of the invention discloses a wrongly written character processing method, device and equipment, and the method comprises the steps: receiving text data of a target domain input by a user, the text data comprising special words of the target domain; performing word segmentation processing on the text data through a preset word segmentation strategy aiming at a target field, and performing identification processing aiming at special words of the target field on segmented words corresponding to the obtained text data to obtain target special words in the segmented words corresponding to the text data; carrying out wrongly written character recognition processing on the target proprietary word through a proprietary word error recognition model for the target field to obtain a first recognition result, and carrying out wrongly written character recognition processing on segmented words, except the target proprietary word, in the segmented words corresponding to the text data through a general word error recognition model to obtain a second recognition result; and based on the first recognition result and the second recognition result, determining wrongly written character information contained in the text data through fusion processing of the two different recognition results.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

A pronunciation prediction method and related device

The present application discloses a pronunciation prediction method and related devices. First, the text to be synthesized is segmented to obtain a segmentation sequence. For the first category of words in the segmentation sequence, the pronunciation information of the first category of words is determined based on a preset corpus resource library. For the second category of words other than the first category of words in the segmentation sequence, their pronunciation categories are determined based on the part-of-speech information of each word in the segmentation sequence, and their pronunciation information is determined based on the pronunciation information determination method corresponding to their pronunciation category. In the present application, combined with the corpus resource library and the preset pronunciation information determination method corresponding to each pronunciation category, it is possible to cover the pronunciation information determination in various situations. Therefore, it is possible to accurately determine the pronunciation information of each word in the text to be synthesized, thereby improving the effect of speech synthesis.
Owner:合肥智能语音创新发展有限公司

A paraphrase sentence recognition method and system based on semantic primitive knowledge and abstract semantic representation

The application belongs to the field of natural language processing, and particularly relates to a method and system for paraphrase recognition based on semantic primitive knowledge and abstract semantic representation, which comprises the following steps: performing word segmentation on a sentence, and performing word-level vector representation and semantic primitive knowledge representation; performing mean value processing on the semantic primitive knowledge representation result, and extracting interactive attention feature information of the mean value processing result by using global semantic information to obtain global semantic primitive representation; performing abstract semantic analysis on a to-be-recognized paraphrase sentence from a sentence structure to obtain a single-root directed acyclic graph, and performing global semantic primitive representation and word-level vector representation; extracting global and local feature information in the order of the directed acyclic graph, and performing distance feature measurement on the information; inputting the distance feature measurement result into a neural network to obtain a recognition result; the application introduces external semantic primitive knowledge to perform semantic representation, the accuracy of the semantic primitive knowledge representation is assisted by global semantic information, and the abstract semantics of a Chinese paraphrase sentence is analyzed to obtain semantic relations, so that the accuracy of paraphrase recognition is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Large model-based work order integration method, apparatus and device, and storage medium

The invention discloses a work order integration method and device based on a large model, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining a target field corresponding to each initial work order, and dividing each initial work order into a plurality of work order clusters according to each target field; performing initial word segmentation processing on each initial work order based on a preset noun knowledge base to obtain a work order after word segmentation, performing vocabulary analysis on each work order after word segmentation, and constructing a target word segmentation table according to a corresponding analysis result; and performing word segmentation processing again on the work orders after word segmentation according to the target word segmentation list to obtain target work orders, refining the target work orders based on the number of the target work orders in the work order clusters and the work order refining large model to obtain corresponding refining results, and integrating the refining results to obtain a refined work order. And obtaining a question list corresponding to each question main body. And through two times of word segmentation processing, the reliability of a work order content extraction result is improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

A method and device for calculating word meaning similarity based on adjacent word features

ActiveCN116522949BThe calculation result is accurateSemantic analysisEnergy efficient computingSentence processingPart of speech
This invention provides a method for calculating semantic similarity based on adjacent word features, relating to the field of natural language processing. First, in the example word extraction module, example words are extracted from the example sentence and the sentence to be matched using the longest common substring algorithm. Second, in the example sentence processing module, word segmentation is performed on the example sentence excluding the example words, and part-of-speech tagging is applied to the example words. Then, in the feature extraction module, features are extracted from the words surrounding the example words in the example sentence. The correlation between the surrounding words and the example words and option words is calculated using a corpus. The calculated results are weighted and combined with their respective features to form the features of the example words and option words. Finally, in the option processing module, the part-of-speech of the option words is compared with that of the example words, and a similarity score is calculated based on the features of the option words to select the optimal option word. This invention determines the features of the current word by extracting features from surrounding words and calculating correlation; it uses a corpus to calculate the mutual information between two words to achieve the correlation between the two words.
Owner:KUNMING UNIV OF SCI & TECH

Input method word frequency adjustment method and device

The present application discloses an input method word frequency adjustment method and device, which are used to solve the technical problem of poor word frequency adjustment effect of input method phrases. An input method word frequency adjustment method includes the following steps: obtaining corpus data; segmenting the corpus data through a word segmentation model to generate a number of word segmentation units; annotating the word segmentation units through a phonetic recognition model to generate word segmentation unit syllables; saving word segmentation units with the same syllables to the same syllable vocabulary; counting the occurrence probability of the first word segmentation unit in the same syllable vocabulary; comparing the occurrence probability of the first word segmentation unit with a preset threshold to obtain a comparison result; adjusting the word frequency of the first word segmentation unit according to the comparison result; arranging the word segmentation unit order of the syllable vocabulary where the first word segmentation unit is located in a preset order according to the adjusted word frequency of the first word segmentation unit, and updating the syllable vocabulary. By dynamically adjusting the word frequency of phrases in the same syllable vocabulary, the accuracy of input is improved.
Owner:BEIJING THUNISOFT INFORMATION TECH

Determining semantic and grammatical correctness of user-expanded sentence using integrated programmatic and specialized guided and constrained artificial intelligence

A system and method guide an Artificial Intelligence engine to determine the semantic and grammatical correctness of a user-expanded sentence in real-time. The sentence validation process involves receiving input from the user, the input includes sentence fragment that the user wishes to expand and user-expanded sentence that the user constructs on the fragment provided. The inputs are broken down into tokens. The word-level tokenization algorithm is used, which identifies tokens by splitting the text into spaces, punctuation marks, and other delimiters. Further, a token comparison algorithm is used to assess the relationship between the sentence fragment and the user-expanded sentence to analyze order and placement. Once the token comparison is complete, a prompt is generated using prompt generator to evaluate grammatical and semantic evaluation of the user-expanded sentence. Real-time feedback is provided to the user based on grammatical and semantic evaluation.
Owner:2HR LEARNING INC

Foreign affair table intelligent filling error correction method and system based on natural language processing

The invention relates to the technical field of data processing, in particular to an intelligent foreign form filling and error correction method and system based on natural language processing. The method comprises the following steps: acquiring text images of a historical corpus and vocabularies; screening segmented words to be corrected; determining the ambiguity of the module based on the frequency proportion of the input vocabulary and the vocabulary type; determining a first possibility based on the co-occurrence probability of the adjacent vocabularies of the to-be-corrected segmented word and the similar segmented word and the similarity of the to-be-corrected segmented word and the similar segmented word; determining a second possibility through an error correction result corresponding to the similar word segmentation history and the image similarity of the word segmentation to be corrected and the similar word segmentation; and determining the total possibility and the necessity of word segmentation error correction through the first possibility, the second possibility and the ambiguity, and completing word segmentation error correction. The error correction accuracy is improved.
Owner:BEIJING DEXUN AVIATION SERVICE CO LTD

Improved key phrase extraction method and system based on key word screening

The invention relates to an improved key phrase extraction method based on key word screening, the improved key phrase extraction method is suitable for key phrase extraction of a single English patent text, the method is based on an existing KeyBERT technology, a key word screening step is introduced, and the accuracy of key phrase extraction is improved. Performing word segmentation on the patent text, removing stop words and punctuation marks, and extracting key words of the patent text; secondly, generating a key phrase candidate list by utilizing a CountVectorizer function in the KeyBERT, and generating a key phrase candidate list according to the key phrase candidate list; then, screening the candidate key phrases based on strong correlation key words in the field to which the patent text belongs; finally, final key phrases are determined through cosine similarity calculation, and the key word screening step is introduced, so that the number of irrelevant candidate phrases can be effectively reduced, the key phrase extraction precision and efficiency are improved, and the method is particularly suitable for patent text analysis with high technicality and high specialty.
Owner:FUDAN UNIVERSITY

Processing method and electronic equipment

The invention provides a processing method and electronic equipment, and the method comprises the steps: obtaining a to-be-processed text, carrying out the word segmentation of the to-be-processed text, and obtaining a plurality of ordered lexical elements; the lexical elements comprise semantic words, and the semantic words are vocabularies representing text semantics in the to-be-processed text; for any semantic word of the plurality of semantic words, sequentially connecting the next semantic word from any semantic word until the text content of the next semantic word is the same as that of any semantic word, and obtaining a plurality of first word groups; combining any two first phrases with the same text content of the final lexical elements in the plurality of first phrases to obtain a plurality of second phrases; screening at least part of the second phrases from the plurality of second phrases according to a predetermined rule, and splicing the at least part of the second phrases to obtain a plurality of text features; and generating core information of the to-be-processed text according to the plurality of text features.
Owner:LENOVO (BEIJING) LTD

Medical automatic question and answer method and system based on common sense fusion

The application provides a medical automatic question and answer method and system based on common sense fusion, comprising: performing word segmentation on training sentences, querying a common sense database and a knowledge base to obtain entity relationship triples; fusing and encoding the triples and the training sentences; randomly selecting part of entities of the fused and encoded training sentences to mask, replacing the next sentence into other random sentences according to a fixed probability, and inputting the obtained training corpus into a multi-layer partition encoder for training; performing word segmentation and entity relationship query on question sentences in question and answer data, and fusing and encoding; taking common sense fusion encoding sequences of the questions as model inputs, taking answers as supervision labels, training a common sense fusion language model; and building a visual medical automatic question and answer system, inputting questions into the model through a front end, and displaying outputs of the model as answers to an interface. The application has no type limitation on user questions, and fuses common sense and medical knowledge into a language model, thereby ensuring that answers conform to grammatical rules and improving professionalism.
Owner:SHANGHAI JIAOTONG UNIV

A bid evaluation method, device and medium based on electronic trading platform

The present invention discloses a bid evaluation method, device and medium based on an electronic transaction platform, which belongs to the field of electronic bid evaluation technology; monitors the project progress in real time, determines whether the bidding project has reached the current progress, and automatically obtains the current project information if the project progress reaches the current node, enters the bidding document preparation and bid evaluation method setting; otherwise, the system continues to monitor the project progress; by introducing a conflict detection mechanism, it effectively solves the conflict problem that may exist in the information block processing process. First, the NLP technology is used to segment the information block, remove stop words and punctuation marks, obtain the text after segmentation, and obtain sentence vectors through the BERT model, extract named entities, and use syntactic analysis technology to build a grammatical tree of the sentence, thereby realizing semantic analysis of the information block. By calculating the semantic similarity between information blocks and setting a threshold to determine whether there is repeated or similar content, the conflict between information blocks can be effectively identified and resolved.
Owner:ANHUI TENDERING GRP INC

A multilingual place name translation method combining syllable segmentation and adaptive learning

The present application relates to the field of place name translation technology and discloses a place name translation method that combines multilingual syllable segmentation with adaptive learning. The method performs word segmentation and common name translation annotation on the source language place name address based on a general dictionary, converts the unannotated proper noun portion into an IPA phoneme sequence, and then performs syllable segmentation and translation matching on the IPA phoneme sequence of each word segment. Multiple possible candidate transliteration results are matched for each word segment. The candidate transliteration results of the proper noun segment are then combined with the common name translation result to construct a candidate set of place name and address target language translation results. The candidate set is then adaptively learned and aggregated using a deep learning algorithm to comprehensively consider the semantic information and contextual coherence of multiple transliteration versions to generate the final place name and address target language translation. The present application can solve the problems of inaccurate, stiff, and ambiguous transliterations in traditional place name and address translation methods, thereby improving translation efficiency and quality.
Owner:SHAANXI TIRAIN TECH CO LTD

Intelligent customer service interactive response method and system based on semantic analysis

The present application relates to the field of semantic analysis technology, and specifically to a method and system for intelligent customer service interactive response based on semantic analysis. The method comprises: obtaining historical data of intelligent customer service and voice data currently input by the user, and converting the voice data into text sentences; preprocessing the text sentences and performing word segmentation; obtaining the average phrase length of the sentence containing each phrase, and obtaining the phrase rationality of each phrase; obtaining the mean continuous length of phrases with the same part of speech in the sentence containing each phrase, and obtaining the unpopularity confidence of each phrase; obtaining the semantic rationality value of each phrase; obtaining the segmentation threshold of the semantic rationality value, and judging whether each phrase is interfered with; evaluating the current text sentence, obtaining an accurate text sentence; and completing the intelligent customer service interactive response. The present application improves the accuracy of semantic analysis by obtaining more accurate text sentences from the user.
Owner:BEIJING WEIHEGUANG DIGITAL TECH CO LTD

A large language model machine translation optimization method and system

This invention relates to a method and system for optimizing machine translation using a large language model, belonging to the field of machine translation technology. It includes: generating a corresponding first translation from a source sentence in a bilingual corpus; segmenting the source sentence into words, counting the frequency of easily misspelled words in each segment, and calculating the easily misspelled word score, obtaining an easily misspelled word set based on the score; calculating the semantic similarity between the sentence to be translated and multiple candidate examples, and calculating the quality scores of the multiple candidate examples; selecting the optimal k candidate examples from the multiple candidate examples based on semantic similarity and quality scores, and constructing them as prompt templates; obtaining training sentence pairs from the bilingual corpus based on the easily misspelled word scores to construct a training set; and using the training set and prompt templates to perform low-rank adaptive training on a large language model to obtain an optimized large language model. This invention not only improves translation quality but also enhances the interpretability of the translation process.
Owner:SUZHOU UNIV

Language Model Training, Video Subtitle Verification Method, Device, Equipment and Medium

The present invention relates to the field of artificial intelligence technology, and provides a method, device, equipment and medium for language model training and video subtitle verification. In this language model training method, sample sentences containing only Chinese characters in a text sample set are input into an initial Chinese character splitting pre-training model with initial parameters, and the sample sentences are successively subjected to word segmentation, radical splitting, granularity splitting and decoding recognition to obtain sample decoded sentences; according to the sample decoded sentences and the sample sentences containing only Chinese characters, a text loss value is determined; when the text loss value does not reach a preset convergence condition, the initial parameters are updated iteratively until the text loss value reaches the preset convergence condition, and the initial Chinese character splitting pre-training model after convergence is recorded as a Chinese pre-training language model based on character splitting. The present invention also relates to blockchain technology, and the Chinese pre-training language model based on character splitting is stored in the blockchain. The present invention can improve the accuracy of preprocessing of characters or texts.
Owner:SHENZHEN PING AN SMART HEALTHCARE TECH CO LTD

An exaggeration representation word extraction method for Chinese irony text

The application discloses an exaggeration representation word extraction method for Chinese irony text, and belongs to the natural language processing technology, and comprises the following steps: step 1: after the irony data set is preprocessed, a bidirectional maximum matching method is used for word segmentation; step 2: the word frequency of the segmented text is calculated by using TF-IDF to construct a candidate word set; step 3: the chi-square statistic is used to measure the correlation degree between the irony text and the exaggeration representation, and the best threshold is set by the chi-square test method to select the strongly correlated exaggeration representation word, so as to construct an exaggeration representation seed word set; step 4: based on the WoBERT semantic similarity calculation framework, the dynamic word vector semantic similarity of the irony text and the seed word set is calculated, and the threshold is set to select the exaggeration representation word with high similarity, so as to construct an exaggeration representation word set. The application aims to extract the words containing the exaggerated expressions in the Chinese irony text to mine the characteristics of the irony sentences, so as to provide technical support for the Chinese irony text recognition task.
Owner:ANHUI UNIV OF SCI & TECH

Text generation content security filtering method and system based on semantic detection

The invention discloses a text generation content security filtering method and system based on semantic detection, and the method comprises the steps: obtaining and preprocessing an original content text generated based on the text to obtain a preprocessed content text, carrying out the sentence segmentation of the preprocessed content text to obtain a plurality of sentences, and determining the overall semantic features of the sentences in the context based on semantic detection; performing word segmentation on the sentence to obtain a plurality of vocabularies, and determining specific semantic features of each vocabulary in the sentence based on semantic detection; and evaluating the content security of each sentence based on the overall semantic feature and the specific semantic feature to obtain a comprehensive security evaluation value, determining the security filtering level of each sentence based on the comprehensive security evaluation value, and performing security filtering on each sentence in the original content text. According to the method, security detection is carried out on the text generation content through a multi-level and fine-grained semantic analysis and evaluation technology, complex harmful content which is difficult to detect through a traditional method can be effectively recognized, and accurate risk grading and intelligent flexible processing are achieved.
Owner:LUSHAN COLLEGE OF GUANGXI UNIV OF SCI & TECH