Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

48 results about "Syntactic structure" patented technology

Syntactic Structures is a major work in linguistics by American linguist Noam Chomsky. It was first published in 1957. It introduced the idea of transformational generative grammar.This approach to syntax (the study of sentence structures) was fully formal (based on symbols and rules). At its base, this method uses phrase structure rules. These rules break down sentences into smaller parts ...

Aspect-level sentiment analysis optimization method based on multivariate external knowledge fusion

The invention discloses an aspect-level sentiment analysis optimization method based on multivariate external knowledge fusion, which relates to the technical field of sentiment analysis optimization, and comprises the following steps: constructing a multivariate external knowledge source comprising a Chinese sentiment dictionary, a domain knowledge graph and a user comment prior mode library; the sentiment module is used for providing vocabulary-level sentiment polarity, entity attribute relations and high-frequency evaluation semantic modes; semantic coding is performed on the input text and the specified aspect words to generate context semantic representation, and global semantic features and local position features are extracted in combination with aspect word position information; based on a semantic coding result, converting the multivariate external knowledge sources into structured knowledge representations, and dynamically adjusting contribution weights of various types of knowledge through a gating fusion mechanism to generate fused knowledge representations; performing dependency syntactic analysis on the input text, constructing an original syntactic structure, and calculating the correlation strength of each grammatical component and aspect words in combination with context semantic representation; and pruning the original syntactic structure according to the correlation intensity.
Owner:HUANENG JINCHANG PHOTOVOLTAIC POWER GENERATION CO LTD

Fine-grained intention recognition method for large-model complex instruction

The invention relates to the technical field of natural language processing and artificial intelligence, and discloses a large-model complex instruction-oriented fine-grained intention recognition method, which comprises the following steps of: constructing a candidate analysis forest of an input instruction, and executing entity alignment of each node and a domain knowledge graph; performing semantic compatibility verification on the candidate dependency trees by utilizing entity relationships in the atlas, and screening an optimal analytic tree in combination with syntactic probability and knowledge consistency scores; if the analytic tree meeting the threshold value does not exist, a semantic conflict edge is positioned, and candidate mounting points are searched by using a map neighborhood relation to reconstruct a dependency structure; and finally, generating standardized data containing a structured analytic path based on the optimal analytic tree, and performing fine adjustment on the large model. According to the method, logic constraint and dynamic repair are carried out on the syntactic structure by introducing the knowledge graph, analysis errors caused by multiple modification or structural ambiguity in a complex instruction are effectively solved, and the accuracy and robustness of intention recognition of a large model in the vertical field are improved.
Owner:BEIJING ZHONGWEI SHENGDING TECH CO LTD

Textual encoding and analysis with a large graphical language model

The techniques discussed herein enhance the operation of content generation and analysis systems. Namely, textual content applications such as technical documentation, creative writing, and content moderation. This is accomplished through generating a graphical representation of a body of text (e.g., a document). The graphical representation can comprise a plurality of nodes representing the words of the document and a plurality of lines that join the nodes representing a level of association between individual words. As such, the graphical representation can capture the semantic and syntactical structure of the associated document while omitting the original textual content. The graphical representation can be subsequently evaluated for complexity based on the density of nodes and lines. Accordingly, the disclosed system can assign a score to a document based on the evaluation of the graphical representation. In addition, various documents can be ranked based on such scores.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Ancient book proofreading method and device and storage medium

The embodiment of the invention provides an ancient book proofreading method and device and a storage medium. In the method, each ancient book page of an ancient book file is divided into a plurality of ancient book segments, a text segment corresponding to each ancient book segment is generated, and a first ancient book character of each ancient book segment corresponds to a second ancient book character in the corresponding text segment; according to the first ancient book character in any ancient book segment, performing letter proofreading on the second ancient book character in the text segment corresponding to any ancient book segment to obtain first proofreading information; according to the syntactic structure in any ancient book page and the first proofreading information in the text page, performing word and sentence proofreading on the text page to obtain second proofreading information; and according to the first and / or second proofreading information in the text file corresponding to the ancient book file, carrying out grammar proofreading on text contents with impaired meanings in the text file to obtain a corrected text file. In this way, the ancient book recognition text can be corrected from multiple dimensions of characters, words, sentences and semantics, and the accuracy of the recognition text is improved.
Owner:HANGZHOU MOULING TECHNOLOGY CO LTD

A machine translation software defect detection method based on combined semantics

The application discloses a machine translation software defect detection method based on combined semantics, comprising the following steps: S1, obtaining different Chinese translation sentences of the same English source sentence from different translation software; S2, using a word alignment model to correspond English words in the source sentence with segmented Chinese words in the Chinese translation sentences; S3, using a sentence compression model and a syntactic structure analysis method to obtain a main part and each additional part of the English source sentence respectively and form a sub-sentence set; S4, aligning each part of the English source sentence obtained in the step S3 with the corresponding translation part to obtain aligned translation of the main part and the additional part of the source sentence; and S5, detecting errors, including semantic similarity calculation and synonym checking.
Owner:TIANJIN UNIV

Knowledge graph relation extraction method fusing syntactic structure and domain rule

PendingCN121787547AEnrich supervision signalsAccurately monitor signalsBiological modelsNatural language data processingRelation classificationAlgorithm
The invention relates to the technical field of natural language processing and mapping knowledge domain, and particularly discloses a mapping knowledge domain relation extraction method fusing a syntactic structure and a domain rule. The method aims at solving the problems of semantic understanding deviation and insufficient domain knowledge utilization in vertical domain relation extraction. The method comprises the steps that dependency syntactic analysis is conducted on a text, a syntactic dependency adjacency matrix is constructed, and meanwhile a rule adjacency matrix is generated based on domain rule base matching; carrying out weighted fusion and filtering on the two types of matrixes to obtain an enhanced adjacent matrix; inputting the matrix and a text vector into a graph convolutional network fused with a dependency type attention mechanism, and learning to obtain a node enhancement representation; and finally constructing an entity pair feature vector to finish relationship classification. Through explicit fusion of interpretable syntactic rules and domain priori, the accuracy and robustness of relation extraction in professional fields such as a power distribution network are improved. Experiments show that the model relation extraction F1 value reaches 86.07% and is improved by 1.31% compared with a baseline model with the optimal performance.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Low-resource event argument extraction method based on language consistency example

The invention relates to a low-resource event argument extraction method based on a language consistency example, and belongs to the technical field of event argument extraction. According to the method, by introducing an example selection mechanism based on language consistency, the accuracy of event argument extraction under the low-resource condition is effectively improved. Compared with an existing method for conducting example matching only according to sentence-level semantic similarity, the method has the advantages that the granularity of semantic modeling is refined, the consistency of a syntactic structure is further combined, it is ensured that the selected example is highly consistent with a target event in the aspects of event trigger words, argument distribution, syntactic dependency and the like, and the accuracy of example matching is improved. And the method has relatively high precision when low-resource event argument extraction is carried out.
Owner:BEIJING INST OF COMP TECH & APPL

Intelligent analysis system for English long and difficult sentence structure in combination with context characteristics

The invention, which relates to the technical field of English parsing, discloses an intelligent parsing system for a long and difficult English sentence structure in combination with context features, comprising a context feature hierarchical extraction module, a dynamic coupling association module, an adaptive weight adjustment module and a syntactic structure parsing module. According to the method, a three-layer context feature layered extraction model of a syntactic structure layer, a chapter semantic layer and a sentence pattern function layer is constructed, the multi-dimensional context features of English long and difficult sentences are dynamically coupled and associated, and the feature analysis weight of each layer is optimized in real time according to the context feature distribution through an adaptive weight adjustment module; synchronous linkage of syntactic structure analysis and context semantics is achieved, the problems that in the prior art, due to the fact that multi-dimensional context association is ignored, nested subordinate sentence splitting errors, logic confusion of master and slave sentences, core component positioning deviation and the like are caused are solved, and the accuracy of English long-difficult sentence structure analysis is improved; and a technical support is provided for application scenes depending on long and difficult sentence analysis, such as machine translation, academic literature research and reading, language teaching and the like.
Owner:吕冰

Configurable natural language output

A system is provided for determining a natural language output, responsive to a user input, using different speech personality profiles. The system may determine to user a particular language generation profile based at least in part on data relating to the user input and data corresponding to the response to the user input. The language generation profile may include different attributes that are used to determine the natural language output, such as, prosody, replacement words, injected words, sentence structure, etc.
Owner:AMAZON TECH INC

Real-time tone feedback in video conferencing

A computer-implemented process is programmed to programmatically receive, using a first computer system, electronic digital data representing input time-correlated speech data and video data, determine a first text sequence corresponding to the input time-correlated speech data, the first text sequence comprising unstructured natural language text, determining syntactic structure data associated with the first text sequence, inputting the time-correlated video data and the syntactic structure data associated with the first text sequence into one or more machine learning models, the machine learning models producing an output of one or more scores for at least a portion of the time-correlated video data and first text sequence, transforming the output of one or more scores to yield and output set of summary points and suggestions, and transmitting a graphical element of the output set of summary points and suggestions for display.
Owner:SUPERHUMAN PLATFORM INC

Triage methods and systems for neurological diseases

PendingCN122314446ATriageNeurology department
This invention relates to the field of natural language processing technology, specifically to a triage method and system for neurological diseases, comprising the following steps: segmenting the syntactic structure of the patient's chief complaint text sequence and identifying nodes; calculating and quantifying the nesting degree of the patient's linguistic organization logic; and obtaining a syntactic potential depth value. This invention can filter out noise interference from complex clinical expressions and extract neurological syndrome features by comparing them with a preset internal capsule lesion pattern. By jointly analyzing feature tensors pointing to limb paralysis with syntactic potential depth values ​​below the normal threshold, dual cross-validation of somatic motor disorders and central cognitive impairment is achieved. Furthermore, by comparing with feature thresholds of peripheral nerve or spinal cord lesions, the specificity of identifying central acute and severe illnesses such as ischemic stroke is improved, optimizing the accuracy of emergency triage.
Owner:THE SECOND HOSPITAL OF HEBEI MEDICAL UNIV

Neural machine translation selective knowledge distillation method based on dependency constraint self-attention

The invention relates to a neural machine translation selective knowledge distillation method based on dependency constraint self-attention, and belongs to the technical field of machine translation. An existing knowledge distillation method has the problems that only vocabulary-level probability distribution is transmitted, syntactic structure constraints are ignored, and the capacity of a student model is reduced, so that the complex syntactic modeling capacity is insufficient. Therefore, according to the method, a syntactic matrix converted through a source language dependency syntactic tree is provided, linear combination is adopted to dynamically correct self-attention weight distribution of an encoder, and explicit syntactic constraints are synchronously injected into a teacher-student model. Through a selective distillation strategy of syntax perception, deep knowledge effective for training in a teacher model is screened, a cross-language syntax corresponding relation is obtained in combination with structure alignment distillation, and model compression and translation performance enhancement is realized through attention optimization and a selective knowledge transmission mechanism guided by a dependency syntax.
Owner:KUNMING UNIV OF SCI & TECH

Ancient Chinese machine translation method for syntactic perception and knowledge enhancement

The invention discloses an ancient language machine translation method for syntactic perception and knowledge enhancement, and belongs to the technical field of natural language processing. Aiming at the difficulties of complex sentence patterns and scarcity of historical knowledge of ancient Chinese, three innovation modules are designed: firstly, a dynamic dependency syntactic analysis module explicitly models an ancient Chinese syntactic structure by utilizing probability analysis and a graph network; secondly, a retrieval enhancement generation module accurately extracts related knowledge fragments from the high-quality ancient language corpus, and semantic comprehension is enhanced; and finally, the stream grammar constraint decoder fuses the source language syntax and the target language generation process through a double-stream mechanism and a stream mechanism, and loyalty and smooth modern text output is realized. The three modules are deeply fused, so that the accuracy and interpretability of ancient language translation are effectively improved, and a reliable technical path is provided for ancient book digitization.
Owner:ZHONGBEI UNIV

A fault scenario search method, system, device, storage medium and product

This invention discloses a method, system, device, storage medium, and product for searching fault solutions. It involves using a search engine to perform preliminary matching of fault descriptions to obtain a set of candidate fault solutions; extracting semantic vectors from both the fault descriptions and candidate fault solutions; constructing a syntactic dependency tree for the fault descriptions and candidate fault solutions to obtain a graph structure; inputting the graph structure and the corresponding semantic vectors into a graph convolutional neural network to obtain syntactic feature vectors for the fault descriptions and the candidate fault solutions; bidirectionally fusing and concatenating the syntactic feature vectors to obtain a fused feature for each candidate fault solution; and ranking the candidate fault solutions based on the fused features to obtain an ordered fault solution matching the fault description. This invention enables deep capture of semantic and syntactic structures to improve search relevance and provide more accurate results for fault solution searching.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

An AIGC text detection method and system based on multi-dimensional feature fusion

This invention discloses an AIGC text detection method and system based on multi-dimensional feature fusion, belonging to the field of natural language processing. It addresses the performance bottlenecks and interpretability issues of existing technologies in text rewriting / polishing and multi-model hybrid content detection. Closed-loop detection is achieved through S1 multi-dimensional feature extraction (four orthogonal dimensions: statistical linguistics, semantics and fluency, syntactic structure, and information theory), S2 attention-weighted feature fusion (Z-score normalization + dynamic weighting by an attention network), S3 adaptive threshold determination and confidence evaluation (dynamically generating thresholds based on text metadata), and S4 interpretability report generation. The system includes feature extraction, fusion, discrimination, and report generation modules, as well as offline construction and online inference units. It balances detection accuracy, anti-circumvention robustness, cross-scenario generalization ability, and result interpretability, with low inference costs and easy scalability.
Owner:HANGZHOU ZHULONG ZHIYUAN TECHNOLOGY CO LTD

A Chinese sentence classification method that integrates dependency syntax and global co-occurrence information

PendingCN122309738APart of speechAlgorithm
This invention discloses a Chinese sentence classification method that integrates dependency syntax and global co-occurrence information. First, text data is collected for preprocessing and syntactic analysis. Then, a pre-trained language model aggregates sub-word-level features into word-level contextual semantic features. Based on the syntactic analysis results, a dependency graph and a word-part-of-speech co-occurrence graph are constructed. Word-level features are used as node features and input into a graph attention network and a graph convolutional network respectively to complete dual-graph modeling, obtaining syntactic enhancement features and global co-occurrence enhancement features. Subsequently, an adaptive weighted fusion of the two types of enhancement features is performed through a gating mechanism. The fused features are concatenated with the original word-level features and input into a Transformer block to complete global self-attention interaction. Finally, the interacted features are pooled and linearly mapped, and the label with the highest probability is used to achieve Chinese sentence classification. This invention solves the problem that existing technologies mainly rely on pre-trained language model context representations, making it difficult to fully model syntactic structural information and global statistical knowledge.
Owner:XIAN UNIV OF TECH

Hybrid reasoning BIM (Building Information Modeling) design achievement auditing method for complex standard specification provisions

The invention discloses a hybrid reasoning BIM (Building Information Modeling) design achievement auditing method for complex standard specification provisions. The method comprises the following steps: S1, defining a grammar structure and an entity class in a standard specification ontology; s2, constructing a standard specification knowledge graph based on the natural language large model; s3, constructing a mapping table; s4, constructing a design result file and text description information; s5, generating an executable rule statement; and S6, performing hybrid reasoning based on a rule reasoning engine and a natural language large model cue word technology. The method comprises the following steps: disassembling a standard text by using a natural language large model cue word technology, and constructing a knowledge graph; the natural language standard specification text is converted into an executable rule, and compliance auditing of the BIM design result is supported; and a rule inference engine and a large language model are combined to perform hybrid inference, so that the BIM model auditing precision is improved.
Owner:CHINA RAILWAY DESIGN GRP CO LTD +1

Sensitive regions around register transfer level (RTL) code objects on visual display

Disclosed subject matter relates to verification and debugging tool and method for providing sensitive regions around and in Register Transfer Level (RTL) code objects. The verification and debugging tool includes input interface configured to receive RTL code in Hardware Description Language (HDL). Further, verification and debugging tool includes RTL parser configured to parse RTL code, and generate abstract syntax tree representing syntactic structure of RTL code. Thereafter, verification and debugging tool includes graph converter configured to convert abstract syntax tree into RTL directed graph. Furthermore, verification and debugging tool includes visual display configured to display code blocks, each exhibiting sensitive regions. Finally, verification and debugging tool includes objects extractor configured to detect user actions related to sensitive regions. The sensitive regions are provided around and in each of RTL code objects to facilitate extraction and display of relevant information regarding dependencies and operations of RTL code upon user interactions.
Owner:SINGULARITY DYNAMICS PTE LTD

A word unit consistency framework recommendation method for framework semantic knowledge base construction

The present application belongs to the field of natural language processing, and particularly relates to a word unit consistency framework recommendation method for framework semantic knowledge base construction. The method combines the application scenarios faced by framework recommendation in the knowledge base construction process, designs and provides a framework recommendation method based on word unit consistency for framework creation, word unit expansion and framework semantic data set construction. From the perspective of framework recommendation, the method fuses a Chinese framework recommendation method with syntactic structure information, provides related auxiliary analysis tools for improving the coverage of frameworks and word units in the knowledge base, and provides a more robust representation technology for the same word unit in different contexts. In addition, the method takes into account research and practicality, and from the engineering perspective, can be integrated into the existing framework semantic labeling platform to improve the usability of the platform and accelerate the construction of the knowledge base.
Owner:SHANXI UNIV

Named entity recognition enhancement method and system based on large language model

The invention discloses a named entity recognition enhancement method and system based on a large language model, and belongs to the technical field of natural language processing, and the method comprises the steps: firstly, training the large language model through direct preference alignment, and enabling a generated rewritten text to be more adaptive to a recognition mode of a named entity recognition model in a semantic or syntactic structure; secondly, performing context rewriting on the input text for multiple times by utilizing a preference alignment large language model, and converting a difficult sample which causes recognition failure of the named entity recognition model or low-confidence output into a high-confidence recognizable candidate text; and finally, integrating a plurality of rewritten statement recognition results through a majority voting mechanism, and outputting a final entity recognition result in combination with an original sentence entity recognition result. According to the method, the recognition accuracy and robustness of the named entity recognition model in long-tail entity, unregistered word and complex context scenes are effectively improved.
Owner:WUHAN UNIV

A multilingual syntax tree generation method based on multi-source syntax guidance and large model collaborative optimization

This invention relates to a multilingual syntax tree generation method based on multi-source syntax guidance and large-scale model collaborative optimization, belonging to the field of natural language processing. The method first generates a dependency syntax tree from the source language text using a traditional syntax analyzer, and performs quality assessment based on the reliability of the root node and dependency relations. For structures with low scores, the Qwen3-8B model is introduced for preliminary correction to improve the quality of the source tree. Subsequently, the segmented target language text and the quality-aligned source language syntax structure are used together as input prompts to guide multiple large-scale models to collaboratively generate the target language syntax tree, achieving cross-language syntax structure transfer. Finally, through consistency integration, rule compliance verification, and format standardization steps, high-quality, uniformly formatted pseudo-syntax tree data is generated. This invention can significantly improve the accuracy of low-resource language syntax analysis and the effectiveness and reliability of data generation standardization in low-resource dependency analysis tasks.
Owner:KUNMING UNIV OF SCI & TECH

Intention recognition method, device and related equipment

This disclosure provides an intent recognition method, apparatus, and related devices, relating to the fields of computer and internet technology. The method includes: performing word segmentation and word vector encoding on target text to obtain multiple node features; constructing a basic hypergraph using at least one of the following methods, and predicting the intent in the target text based on the basic hypergraph: clustering multiple node features and associating semantically similar word nodes with semantic concept hyperedges in the basic hypergraph; parsing the target text using syntactic analysis tools and associating word nodes corresponding to core grammatical components with syntactic structure hyperedges in the basic hypergraph; matching the target text using predefined regular expressions and associating word nodes that match the same matching rule with pattern template hyperedges in the basic hypergraph; and extracting predefined types of named entities from the target text and associating entity word nodes of the same type with entity association hyperedges in the basic hypergraph. Embodiments of this disclosure can improve the accuracy of intent recognition.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

A structured dynamic semantic reconstruction method, device and storage medium

The embodiment of the application discloses a structured dynamic semantic reconstruction method, device and medium. The method comprises: in the process of dynamic dialogue reconstruction, obtaining to-be-processed data; inputting the to-be-processed data into a pre-trained dynamic semantic reconstruction model, and outputting a result of semantic reconstruction of speaking information of a current last person. The application enhances the understanding ability of the model to the sentence with complex syntax structure by performing syntax dependency analysis on the sentence and fusing the syntax dependency analysis into the word coding, position coding and other modules of the model, so that the sentence can be correctly and smoothly reconstructed, and the semantic understanding deviation problem of the algorithm model caused by the existence of reference and omission in human-computer dynamic interaction is solved. Meanwhile, the application does not need a large amount of manual annotation work, and can be trained only by using dialogue data, thereby reducing the dependence on high-quality manual annotation data in the training process of the algorithm model, and reducing the dependence on hardware materials required for deployment of the algorithm model.
Owner:CHONGQING JUEXIAO TECH CO LTD

Language structure analysis method using language expression

The invention relates to a method for analyzing a language structure by using a language expression, which is suitable for analyzing structures of more than 7,000 languages on the earth. The method comprises the following steps of: analyzing a morphologic element and a syntactic structure of a target sentence by a morphologic element and syntactic analysis unit; and based on a pre-training language expression learning algorithm provided by a language expression learning algorithm unit, a language expression calculation unit generates a language expression for the morphologies and syntactic structures of the sentences. Wherein the language expression calculation unit is used for dividing a sentence into an upper left area, a lower left area, an upper right area and a lower right area by taking a sentence center symbol as a benchmark, a subject symbol is distributed in the upper left area, a core verb symbol is distributed in the lower left area, and a verb type symbol is distributed in the lower right area; non-predicate symbols representing remaining phrases other than subjects and verbs are assigned to the right upper region. Furthermore, symbolization specific information including marks (1-4) indicating a person and a character number mark (.) indicating the number of word characters such as nouns and adjectives is added to at least one of the left upper mark, the left lower mark, the right upper mark and the right lower mark of each symbol. By means of the method, the sentences in the specific language can be accurately analyzed into structured syntactic expressions, and therefore the efficiency of language processing such as sentence recognition, semantic understanding, translation and retrieval is improved.
Owner:崔朝植

A method and apparatus for quantifying continuous emotion intensity based on multidimensional feature decoupling

PendingCN122311193AMathematical OperatorsFeature Dimension
This invention provides a method and apparatus for quantifying continuous sentiment intensity based on multi-dimensional feature decoupling. The method identifies the attribution relationship between semantic carriers and subjects in the text to be analyzed to obtain a first feature basis vector; extracts absolute sentiment values ​​or feature vector information of words to obtain a second feature basis vector; and identifies syntactic structure level, organizational rules, or context topology layer to obtain a third feature basis vector. A high-dimensional decoupled feature space with a total dimension of is constructed, and the above three are mapped as a subset of baseline dimensions into it. This space supports dynamic insertion and modular expansion of any external variables across scenarios and topics as incremental feature dimensions, and adaptively adjusts the input matrix of the combination operator. The aggregation engine uses mathematical operators to perform spatial mapping and dimensionality reduction aggregation on the multi-dimensional feature tensor to obtain the absolute value of continuous sentiment intensity corresponding to each subject. This invention can capture structural changes and abnormal moments in sentiment intensity, demonstrating the evolution of public sentiment and opinions.
Owner:陈轶伦 +2

Unsupervised Chinese contrast learning method based on prompt learning and mutual information

The invention belongs to the technical field of semantic text similarity, and particularly discloses an unsupervised Chinese comparative learning method based on prompt learning and mutual information. According to the method, prompt learning is introduced to serve as a data enhancement method, on the basis, a method for obtaining sentence embedding based on manual design prompt is used, a sentence embedding task is reexpressed as a mask language task, the gap between a pre-training model and a training target is reduced, an original BERT or RoBERTa layer can be more effectively used, the performance of the model is improved, and the user experience is improved. Besides, the positive sample pairs are constructed by the same sentences through data enhancement, and the mode of constructing the positive sample pairs can influence the alignment of the structures in the positive sample pair enhancement view, so that mutual information maximization is adopted on the attention tensor of the positive sample pairs, and the accuracy of the positive sample pairs is improved. The alignment of view structures is maximized by enhancing the distribution similarity of attention values in enhanced views of positive sample pairs, so that the distance between the positive samples is better shortened, and the learning ability of a model to a syntactic structure is improved.
Owner:GUANGDONG UNIV OF TECH

Dependency analysis model and chinese joint event extraction method based on dependency analysis

The application discloses a Chinese joint event extraction method based on dependency analysis, first introduces dependency analysis to construct a syntactic structure and strengthens the depth interaction of information; secondly, three types of edge representations are designed to calculate graph convolution features in order to bridge the inconsistency of words; finally, the cascade error propagation problem of the traditional pipeline method is relieved through joint learning of the event trigger word classification task and the event argument classification task, and the effect of extracting event triggers and arguments from documents is improved. The Chinese joint event extraction model based on dependency analysis integrates syntactic structure information while encoding semantics, enhances the information flow between words, and designs different types of edge representations for the construction of an undirected graph according to the characteristics of Chinese word segmentation. The application enriches semantic feature representation by integrating the syntactic structure knowledge contained in the Chinese text, and effectively improves the effect of sentence-level event extraction by using the joint learning method.
Owner:MYRON INTELLIGENT TECH (SHANGHAI) CO LTD

Improved TextRank text summarization method based on key phrase selection

PendingCN121880551Aavoid semantic fragmentationincrease information densityNatural language data processingEnergy efficient computingFeature vectorSemantic network
The invention relates to the technical field of text summarization, and discloses an improved TextRank text summarization method based on key phrase selection, and the method comprises the steps: firstly, carrying out the preprocessing of an original text, and converting the original text into an independent sentence set; obtaining a syntactic structure tree, selecting non-terminal clause nodes, and constructing a key phrase set; performing feature vectorization on the key phrases by using a SimBERT pre-training model, calculating semantic similarity and constructing a text semantic network graph; meanwhile, auxiliary weights such as sentence positions, title similarity and keywords are calculated, and a final sentence weight matrix is generated through linear weighted fusion; and finally, obtaining a global score, carrying out iterative screening and redundancy elimination processing on the candidate phrases, and outputting an abstract after reordering according to the original text logic. According to the method, syntactic structure fine-grained screening, deep semantic interactive calculation and a dynamic redundancy elimination mechanism are combined, the problems of insufficient semantic mining and result redundancy of a traditional abstract method are solved, and the abstract quality is remarkably improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Corpus quality evaluation method and device

The invention provides a corpus quality evaluation method and device. The method comprises the following steps: acquiring a corpus text; quality evaluation results of the corpus text are determined, and the quality evaluation results comprise a first quality evaluation result of sentences in the corpus text, a second quality evaluation result of paragraphs and a third quality evaluation result of the corpus text; wherein the first quality evaluation result is obtained by analyzing sentences based on a pre-trained syntactic structure model, the paragraph is obtained by analyzing a semantic relationship among a plurality of sentences based on a pre-trained semantic fragmentation model, and the second quality evaluation result is obtained by analyzing the paragraph through a pre-trained logic association model; the third quality evaluation result is obtained by analyzing the corpus text through a pre-trained information density model. In this way, the model training speed and the model performance are improved through the corpus training model with higher quality.
Owner:CHENGDU HUAWEI TECH CO LTD