Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

85 results about "Named entity" patented technology

In information extraction, a named entity is a real-world object, such as persons, locations, organizations, products, etc., that can be denoted with a proper name. It can be abstract or have a physical existence. Examples of named entities include Barack Obama, New York City, Volkswagen Golf, or anything else that can be named. Named entities can simply be viewed as entity instances (e.g., New York City is an instance of a city).

Techniques for classifying data using large language models

A system and method for classification. A method includes identifying candidate entities among text data by applying at least one entity identification rule to the text data. Inputs are constructed based on the identified candidate entities, where each input includes a first portion of text indicating a candidate entity and at least one second portion of text and where the at least one second portion of text of each input is adjacent to the first portion of text of the input. Multiple language models are applied to the inputs, where each language model is trained to identify a respective set of entities and where outputs of the language models include at least one portion of entity-indicating text for each input. Based on the outputs of the language models, at least one named entity in the text data is determined.
Owner:CYERA LTD

Entity enhancement and context-aware paragraph retrieval method for RAG system

The invention provides an entity enhancement and context-aware paragraph retrieval method for an RAG system for solving the problems of fuzzy query intention and insufficient paragraph context modeling in RAG retrieval. The method comprises the following steps: firstly, identifying and extracting a key entity by using a named entity, carrying out weighted fusion on the key entity and question representation obtained by a pre-training model to form an enhanced query vector, and accurately describing a core semantic intention of the question; secondly, performing semantic modeling on paragraphs in a document library, mining a potential semantic association relationship between the paragraphs, and constructing a context interaction model between the paragraphs based on a graph neural network and a gating loop unit mechanism, so as to obtain paragraph vector representation with complete semantics and clear hierarchy; and finally, calculating the similarity between the enhanced query and the paragraph vector, and completing high-precision paragraph-level retrieval. According to the method, the correlation and the recall rate are remarkably improved, insufficient entity utilization and weak context modeling are relieved, and good expansibility and cross-domain applicability are achieved.
Owner:SOUTHEAST UNIV +1

Network security domain knowledge graph construction method, system and device, processor and computer readable storage medium thereof

The invention relates to a network security domain knowledge graph construction method. The method comprises the steps of (1) performing named entity extraction for a network security domain based on multi-model cooperative verification, and training a lightweight model; (2) segmenting a long text based on an entity perception multi-dimensional scoring dynamic sliding window; (3) performing named entity and relation extraction and lightweight entity relation identification model construction based on multi-model collaborative network security; and (4) based on the extracted and disambiguated entities and relationships, designing a knowledge graph mode to construct a network security knowledge graph. The invention also relates to a corresponding system, device, processor and computer readable storage medium. By adopting the network security domain knowledge graph construction method, system and device, the processor and the computer readable storage medium, the computing power demand of a large model during element extraction is effectively reduced, and the accuracy and recognition types of network security entities and relationships thereof when the large model processes a long text are improved.
Owner:THE THIRD RES INST OF MIN OF PUBLIC SECURITY

Causal knowledge graph construction and question-answering system based on retrieval enhancement generation and large language model

The invention discloses a causal knowledge graph construction and question-answering system based on retrieval enhancement generation and a large language model. The system comprises a multi-source heterogeneous knowledge base construction module, a retrieval enhancement generation module, a named entity recognition and causal triple extraction module and a knowledge fusion reasoning module, according to the method, a high-quality named entity annotation data set and a causal data set are constructed, and normalization and integrity of input knowledge are guaranteed; a mixed retrieval strategy (keyword + vector + sparse embedding) is provided, and the evidence coverage rate and recall precision are improved. Under RAG driving, LLM is combined, two rounds of causal triple extraction are achieved, and the causal relationship coverage degree and the direction judgment confidence degree are improved; a conflict detection and atlas fusion mechanism is designed to ensure the unity and consistency of new and old causal knowledge; a question answering system based on a causal atlas is also established, multi-hop causal reasoning is supported, and answers with controllable credibility and explainable are output.
Owner:CHONGQING UNIV

Determining labels of inheritance datasets using simulated data instances

Disclosed is a method for determining inheritance labels of users based on inheritance datasets of the users. The method includes generating a plurality of reference panels for a plurality of data-inheritance origins, each reference panel corresponding to a data-inheritance origin and comprising reference-panel datasets representative of the data-inheritance origin. The method constructs a plurality of simulated data trees that are built using the reference-panel datasets that are selected from the plurality of reference panels. The method generates a plurality of simulated inheritance datasets representing a plurality of simulated named entities, each representing a descendant named entity in one of the simulated data trees. The method trains a machine learning model to determine inheritance labels of an inheritance dataset.
Owner:ANCESTRY COM DNA LLC

A network security entity identification method and system based on multi-layer channel attention

The application discloses a network security entity identification method and system based on a multi-layer channel attention, and belongs to the technical field of network security entity identification. The specific method comprises the following steps: performing data preprocessing on a real network security entity text data set; constructing a named entity identification model based on a multi-layer channel attention; inputting the preprocessed network security entity text data set into the named entity identification model for model training, so as to obtain a trained named entity identification model; inputting an actual network security entity text data set into the trained named entity identification model, and outputting a label sequence; identifying the entity category of each word in the actual network security entity text data set according to the output label sequence, and realizing network security entity identification based on a multi-layer channel attention. The application can accurately identify various entity types in network security text, and improve the monitoring and analysis capability of network security events.
Owner:STATE GRID LIAONING ELECTRIC POWER CO LTD +1

Customs declaration information processing method and device and electronic equipment

Embodiments of the present application disclose a customs declaration information processing method and device and electronic equipment. The method comprises: determining a customs declaration material file associated with a to-be-generated customs declaration form; identifying text information content of the customs declaration material file to determine at least one named entity included therein, the named entity comprising a continuous character fragment in the text information content; and during an information input operation for a target field in the customs declaration form, providing recommended information about to-be-input information in the target field according to the named entity identified from the customs declaration material file. Through the embodiments of the present application, the efficiency of generating a customs declaration form can be improved, and the probability of input errors and the like caused by a manual input process can be reduced.
Owner:ALIBABA SINGAPORE HLDG PTE LTD

Real-time context sensing dynamic cue word generation method based on large model translation

The invention discloses a real-time context sensing dynamic cue word generation method based on large model translation. The method comprises the steps of initialization, context sensing and core concept extraction, dynamic cue word generation, new concept recognition and list updating. According to the method, a traditional static translation process is converted into a dynamic process with memory and state maintenance capabilities, a global core concept list is constructed and dynamically maintained in real time, and dynamic cue words containing determined standard translations are automatically generated in the translation process; the global consistency of the translations on core concepts such as term named entities is remarkably improved, the problem that core concept translations are inconsistent in long document translation is fundamentally solved, and the method is suitable for professional fields such as technical documents, academic papers and commercial reports with extremely high term consistency requirements.
Owner:IOL WUHAN INFORMATION TECH CO LTD

NLP-driven district-level BIM data space conflict automatic checking method and system

The invention relates to the technical field of BIM conflict checking, and discloses an NLP-driven district-level BIM data space conflict automatic checking method and system, and the method comprises the steps: receiving a natural language checking instruction of a user, inputting a pre-training multi-modal field adaptive model through an NLP processing module, and carrying out the NLP-driven district-level BIM data space conflict automatic checking. Completing intention recognition and named entity extraction to obtain key semantic elements and generate unified semantic representation; a dynamic task graph is constructed, tasks are simplified, after a target model and a collision rule are called, geometric intersection detection is accelerated by a GPU through a collaborative architecture scheduling engine cluster, and a CPU is responsible for topology analysis; acquiring original collision data, determining collision points by combining engineering knowledge graph clustering and grading, and generating an intelligent interpretation report; conflicts are visualized on a CIM platform, a BIM correction script is generated in combination with user operation and a grading result, and an updating model and a collision rule are synchronously collected and fed back. According to the BIM data processing method and the BIM data processing device, automation and high efficiency of BIM data processing can be considered while natural language instruction driven checking is carried out.
Owner:URBAN PLANNING & DESIGN INST OF SHENZHEN UPDIS

An autoregressive model-based cross-domain named entity recognition method

The application discloses a cross-domain named entity recognition method based on an autoregressive model, and comprises the following steps: S1, encoding an input sequence; S2, encoding a label through a label encoder; S3, obtaining label background information; S4, obtaining label context information; S5, connecting the label background information to the input sequence and connecting the label context information to the predicted named entity label as final label perception information z i , and finally obtaining a final sequence representation u; the application provides a cross-domain named entity recognition method based on an autoregressive model, which improves the relationship between a source text and its named entity label, improves the portability of label information, and helps the model to promote domain adaptation.
Owner:BEIJING INST OF TECH

Named entity bias detection and mitigation techniques for sentence sentiment analysis

Techniques for named entity bias detection and mitigation for sentence sentiment analysis. In one particular aspect, a method is provided that includes obtaining a training set of labeled examples for training a machine learning model to classify sentiment, preparing a list of named entities using one or more data sources, for each example in the training set of labeled examples with a named entity, replacing the named entity with a corresponding entity type tag to generate a labeled template data set, executing a sampling process for each entity type t within the labeled template data set to generate a augmented invariance data set comprising one or more invariance groups having labeled examples for each entity type t, and training the machine learning model using labeled examples from the augmented invariance data set.
Owner:ORACLE INT CORP

A small sample named entity recognition model training method and recognition method

ActiveCN115759103BNatural language data processingNamed-entity recognitionNearest neighbor classifier
The application provides a small sample named entity recognition model training method, comprising the following steps: S1, obtaining a training set, a training set type description set, a support set and a support set type description set; S2, mining clue words in each sample on the training set and the support set respectively and performing clue word labeling to obtain the training set and the support set containing named entity labels and clue word labels respectively; S3, performing multi-round iterative training on a basic named entity recognition model until convergence by using the training set and the training set type description set processed in step S2; and S4, performing migration training on the basic named entity recognition model trained in step S3 until convergence by using the support set and the support set type description set processed in step S2, to obtain a small sample named entity recognition model composed of an encoder and a nearest neighbor classifier.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A method and system for Chinese-oriented named entity recognition

PendingCN122509177AStrong semantic expression abilityavoid relational ambiguityNamed-entity recognitionMulti-label classification
The application provides a Chinese-oriented named entity recognition method and system, and belongs to the technical field of natural language processing; comprising: an encoding stage step, context encoding of input Chinese text; a feature extraction stage, a double-branch structure of global semantic branch and local boundary branch in parallel is adopted, wherein the local boundary branch carries out boundary feature enhancement through multi-granularity convolution combined with a learnable Gaussian Laplacian operator; a feature fusion and prediction stage step, global and local features are fused through a gating mechanism, and multi-label classification prediction is carried out based on a preset six-tuple grid label system; a decoding stage step, named entities are obtained through bidirectional search strategy decoding according to predicted relationship probability; through boundary perception feature extraction and enhanced label system, the recognition accuracy and recall rate of nested entities and discontinuous entities in Chinese text are significantly improved.
Owner:SOUTHWEST UNIVERSITY FOR NATIONALITIES

Data processing method and computer readable storage medium

The present disclosure provides a data processing method and a computer readable storage medium, the data processing method is used for training a named entity recognition model, comprising: obtaining a labeled training sample pair and an unlabeled training sample pair; for each training sample pair, obtaining the corresponding latent representation features of the corresponding training sample pair and fusing, and then obtaining a first prediction result of the labeled training sample pair and a second prediction result of the unlabeled training sample pair by performing named entity prediction; obtaining the reconstruction features of each training sample pair according to the latent representation features of each training sample pair; determining three loss functions based on the first prediction result, the second prediction result, each sample pair and the reconstruction features of each sample pair; and training the named entity recognition model according to the three loss functions. The present disclosure adopts semi-supervised training, which can reduce the training cost while ensuring the accuracy of the model.
Owner:MASHANG CONSUMER FINANCE CO LTD +1

Joint extraction method for judicial text entity relationship based on regional vertex labeling

The invention discloses a judicial text entity relationship joint extraction method based on regional vertex labeling. The method comprises the following steps: coding judicial text cleaning data through a pre-training language model BERT to obtain corresponding vector representation and embedded representation; a head entity and a tail entity in a triple corresponding to the original judicial text data form a rectangular area in the target representation set, and four vertexes of the rectangular area are identified to identify the triple; calculating probability scores that target representations in the target representation set belong to the four vertexes; and extracting the triple of the original judicial text data by combining the loss function with the character pairs with the distributed labels. According to the method, the key information in the original judicial text data is converted into the formatted triple, the core relationship between the named entity and the positioning entity in the judicial text is accurately identified, the structured representation of the irregular text is realized, and judicial workers are further assisted in understanding the case.
Owner:DALIAN UNIV OF TECH

Systems and techniques for handling long text for pre-trained language models

In some aspects, a computing device can receive, at a data processing system, a set of utterances for training a named entity recognizer or for inference with a named entity recognizer to assign a label to each token segment in the set of utterances. The computing device can determine a length of each utterance in the set and, when the length of the utterance exceeds a predetermined threshold of token segments: divide the utterance into a plurality of overlapping token segment chunks; assign a label and a confidence score to each token segment in the token chunk; determine a final label and an associated confidence score for each token segment chunk by merging two confidence scores; determine a final annotated label for the utterance based at least on merging the two confidence scores; and store the final annotated label in a memory.
Owner:ORACLE INT CORP

Value-directed parsing for data extraction

Described are methods and systems for parsing unstructured or semi-structured text to extract named entities, data types defined to include semantic fields. Fields are constrained to sets of potential field values. These sets can overlap, leading to ambiguous parses. For example, the text string “3-4-2023” parsed as a date can yield Mar. 4, 2023 or Apr. 3, 2023. Potential field values in alternative parses are scored and the scores used to select and disambiguate the resultant value parses.
Owner:ZOHO OFFICE SUITE

Chinese named entity recognition retrieval enhancement framework based on uncertain components

The application provides a Chinese named entity recognition retrieval enhancement framework based on uncertain components, which can sample uncertain components of an input Chinese sequence by including a named entity recognition model, and can retrieve uncertain components based on an entity set obtained by sampling, so compared with a traditional method relying on a dictionary, the required knowledge sequence can be effectively retrieved without the need to spend high costs to construct and dynamically maintain a high-quality dictionary, thereby saving a large amount of computing power, and further, since a knowledge fusion model is included, the Chinese named entity can be recognized and predicted based on the retrieved knowledge sequence, so the ambiguity in the recognition process can be eliminated by the knowledge sequence, and a more accurate prediction result can be obtained, and compared with a traditional method of detecting by using a traversal strategy, the efficiency is higher. In summary, by using the enhancement framework, accurate Chinese named entity recognition results can be efficiently obtained, and a large amount of computing power can be saved.
Owner:FUDAN UNIVERSITY

Named entity sampling method and device based on diffusion model

The invention provides a named entity sampling method and device based on a diffusion model, and the method comprises the steps: carrying out the marking and coding of named entities in training text data, and obtaining the boundary information corresponding to each named entity in the training text data; gradually adding noise in the boundary information to obtain target training text data, and training a preset diffusion model to obtain a target diffusion model; in response to the obtained target text data, determining the sampling time step length of the target diffusion model in the next time step based on sampling state information predicted by the target diffusion model for the target text data in the current time step and the historical time step; and performing iterative sampling on the named entities in the target text data according to the sampling time step length by using the target diffusion model to obtain named entity information corresponding to the target text data output by the target diffusion model. Through the method, the accuracy, flexibility and efficiency of named entity recognition are improved.
Owner:CHINA ELECTRONICS CORP 6TH RES INST +1

Colored lamp culture field named entity identification method

The invention discloses a colored lamp culture field named entity recognition method, which belongs to the technical field of natural language processing, and comprises the following steps: acquiring colored lamp culture text data, and extracting context feature representation through a pre-training language model; according to the context feature representation, applying adversarial disturbance in an embedding space to obtain a feature representation with optimal robustness; according to the feature representation with the optimal robustness, deep semantic coding is carried out through a residual bidirectional long-short-term memory network, and coding features are obtained; according to the coding features, sequence labeling decoding is carried out through a conditional random field, and predicted entity fragments are obtained; according to the predicted entity segment, decoding and filtering after type sensing are carried out, and a final named entity recognition result is obtained. According to the method, accurate identification of high-density process terms and heterogeneous data in the colored lamp culture field is realized, and high-quality entity extraction capability is provided for intangible cultural heritage digitization and knowledge graph construction.
Owner:SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Chinese named entity recognition method, electronic device, and storage medium

The application provides a Chinese named entity recognition method based on an attention mechanism, an electronic device and a storage medium, which comprises the following steps: inputting a text to be recognized into an embedding layer to obtain a word vector; using a Transformer encoder to extract features of the word vector to obtain a first context feature; using a Bi-LSTM model to extract features of the word vector to obtain a second context feature; fusing the first context feature and the second context feature to obtain a fused feature; and decoding the fused feature to obtain Chinese named entities corresponding to the text to be recognized. The Chinese named entity recognition method based on the attention mechanism realizes deep fusion of global semantic information and directional information. In order to obtain more context information and solve the problem of polysemy, a RoBERTa-wwm pre-training model is used as a character-level embedding, so that the model recognition effect is improved.
Owner:JIANGXI UNIV OF SCI & TECH

Emerging risk event detection and evaluation

PCT designated stageWO2026015082A1FinanceForecastingEmerging riskData mining
Automatically collating information from a corpus of publications regarding effects of emerging risks on organizations includes collecting digital resources relevant to an emerging risk event type, analyzing, using natural language classifier(s), the digital resources to identify a set of named-entity values, and, using the set of named-entity values, clustering subsets of the digital resources as belonging to a same event of a set of emerging risk events For each cluster subset, the systems and methods may include determining counts of named-entity values within each of the digital resources, classifying a depth of information of each digital resource based on the counts, comparing the digital resources according to semantic similarity to define groups of similar digital resources, and, based on the groups and the depth of information of each of the digital resources, selecting a representative set of digital resources.
Owner:AON GLOBAL OPERATIONS LTD (SINGAPORE BRANCH)

Determining cross-document rhetorical connections based on parsing and identifying named entities

An extended discourse tree and a method for navigating text using the extended discourse tree are provided. [Solution] A discourse navigation application creates a first discourse tree for a first paragraph of a first document and a second discourse tree for a second paragraph of a second document, determines entities and corresponding first basic discourse units from the first discourse tree, determines second basic discourse units in the second discourse tree that match the first basic discourse units, determines rhetorical connections between the two basic discourse units, and creates navigable links between the two discourse trees.
Owner:ORACLE INT CORP

A text information extraction method and system based on multi-agent large language model

The application relates to a text information extraction method and system based on a multi-agent large language model. The method extracts a global named entity list in target text information based on a main-agent large language model, and the global named entity list at least includes a sample name list and an experimental method name list; constraint features are constructed according to the global named entity list, and prompt texts of at least two sub-agent large language models are constructed based on the constraint features; all sub-agent large language models are run in parallel, structured data fields are extracted from the target text information according to the constraint features in the corresponding prompt texts of each sub-agent large language model, and names in the structured data fields are unified based on the global named entity list. The application improves the accuracy and data integrity of long document multi-task extraction by constructing a two-stage collaborative architecture combining global coordination of a main agent and constraint of sub-agents in parallel and an explicit term constraint mechanism.
Owner:SHANGHAI INST OF CERAMIC CHEM & TECH CHINESE ACAD OF SCI

Named entity recognition method and device, electronic equipment and storage medium

Embodiments of the present application provide a named entity recognition method and device, electronic equipment and a storage medium. The named entity recognition method comprises: inputting a to-be-recognized text into a pre-trained named entity recognition model; and performing the following operations in the named entity recognition model: for each entity type, obtaining a relative position vector, a query vector and a key vector of a segmented word in the to-be-recognized text under the entity type, and obtaining global pointer information of the to-be-recognized text under the entity type based on the relative position vector, the query vector and the key vector, the global pointer information indicating a probability that the segmented word belongs to the entity type and a probability that the segmented word belongs to an entity start-end position pair; and identifying an entity type and a start-end position pair of a named entity in the to-be-recognized text based on the global pointer information. The embodiments of the present application can improve the accuracy of named entity recognition.
Owner:CHINA TELECOM CORP LTD

Entity question and answer generation method and device, equipment and storage medium

The embodiment of the invention provides an entity question and answer generation method and device, equipment and a storage medium. The method comprises the following steps: analyzing a to-be-processed knowledge file to obtain a processable text; according to the processable text, performing semantic segmentation by adopting a bidirectional language representation model to obtain one or more to-be-extracted text blocks; adopting a trained entity question and answer generation model to obtain a named entity corresponding to each to-be-extracted text block, and obtaining an entity question and answer pair corresponding to each named entity; the method comprises the following steps: acquiring a training named entity and a training question-answer pair for training an entity question-answer generation model through a cue word project and a language large model; and according to the question statement, screening out an entity question and answer pair matched with the question statement, and outputting an answer statement in the entity question and answer pair. The training named entities and the training question and answer pairs are generated through the cue word engineering and the language large model and are used for training the entity question and answer generation model, and the richness and the retrieval recall rate of a knowledge base are improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

A method and device for training a named entity and relationship joint recognition model

The application relates to a training method and device of a named entity and relationship joint recognition model, and belongs to the technical field of natural language processing; the method solves the problem that in the prior art, named entity and relationship recognition needs to be performed in two independent tasks, namely, entity recognition and classification, and the time consumption is long, and the resource use efficiency is reduced; the model training method comprises the following steps: obtaining a named entity and relationship joint annotation data set D ALL , adding task descriptions related to named entity recognition and relationship recognition to original text data in the data set D ALL , constructing target outputs according to the task descriptions, further constructing a training sample set, training data in the training sample set, and obtaining a named entity and relationship joint recognition model.
Owner:BEIJING ZHONGKE ZHIJIA TECH CO LTD

Named entity recognition system and named entity recognition method

Provided is a named entity recognition system, including: an input module configured to recognize a speech input of a user and convert the speech input into text; a preprocessing module configured to separate the text in units of syllables and perform transformation; and a learning module configured to perform multi-task learning for recognizing a named entity and identifying a boundary of spacing with respect to the transformed text, and output a result of recognizing the named entity and a result of identifying the boundary of spacing, based on recognizing the named entity and identifying the boundary of spacing.
Owner:HYUNDAI MOTOR CO LTD +1

System and method for fact verification using blockchain and machine learning technologies

A method for performing fact verification includes receiving a document for verification; identifying and extracting, on the computer network, at least two named entities from the document for verification and associated information as metadata; identifying, off-chain, identifiers of relevant documents corresponding to each of the at least two named entities, the relevant documents being authenticated documents including one or more of the at least two named entities; identifying a predetermined number of relevant documents among the identified relevant documents based on a number of the at least two named entities present and a number of occurrences for each of the named entities; and determining whether the document to be verified is supported by the predetermined number of relevant documents.
Owner:JPMORGAN CHASE BANK NA

Data query method and device, equipment, storage medium and program product

The embodiment of the invention provides a data query method and device, equipment, a storage medium and a program product, and relates to the technical field of big data, the technical field of artificial intelligence and application of a large model in the field of financial science and technology. The method comprises the following steps: receiving a query request input by a user; the query request is a text which is described based on a natural language and requests query data; based on a preset filtering rule corresponding to at least one regularized named entity, identifying the regularized named entity and a semantic type thereof from the query request; based on an entity recognition model, recognizing an irregular named entity and a semantic type thereof from the query request; generating a query statement template of the query request according to the query request; according to the semantic type of the named entity, filling the named entity into a query statement template to obtain a query statement corresponding to the query request; and executing the query statement to obtain a query result corresponding to the query request. According to the method, the accuracy of data query is improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA