Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Noun phrase" patented technology

A noun phrase or nominal (phrase) is a phrase that has a noun (or indefinite pronoun) as its head or performs the same grammatical function as such a phrase. Noun phrases are very common cross-linguistically, and they may be the most frequently occurring phrase type.

Literature semantic search method and system based on elastic search

The invention discloses a literature semantic search method and system based on elastic search, and relates to data retrieval. The literature semantic search method comprises the steps that vectorization processing is conducted on a noun phrase list by means of a text2vec-based-multilingual model trained based on a CoSENT method; according to the semantic vector, performing approximate nearest neighbor search in a second retrieval module to obtain first candidate data; inputting the query text data into a first retrieval module, and performing keyword matching through a BM25 algorithm to obtain second candidate data; fusing the first candidate data and the second candidate data to obtain third candidate data; a Sequence Matcher algorithm is adopted to calculate character string similarity between expansion words in the third candidate data, a similarity threshold value is set based on the length of the longest common subsequence, duplicate removal is carried out, and fourth candidate data is obtained; and performing weight distribution based on positions and similarity scores on the fourth candidate data, and enhancing the distinction degree of the extension words by expanding a score interval to obtain extension word recommendation list data. According to the method, the accuracy of document retrieval is remarkably improved.
Owner:CHINA EDUCATIONAL PUBLICATIONS IMPORT & EXPORT CORP LTD

System and Method for Accurate Responses from Chatbots and LLMs

Systems and methods are described for obtaining accurate responses from large language models (LLMs) and chatbots, including for question and answering, exposition, and summarization. These systems and methods accomplish these objectives via use of noun phrase avoiding processes such as a noun phrase collision detection process, a query splitting process, and a topical splitting process as well as by use of formatted facts, formatted fact model correction interfaces (FF MCIs), bounded-scope deterministic (BSD) neural networks, processes and methods, and intelligent storage and retrieval (ISAR) systems and methods. These systems and methods avoid and bypass noun phrase collisions and correct for errors caused by noun phrase collisions so that hallucinations are eliminated from LLM responses.
Owner:ACURAI INC

Document intelligent writing and analysis system based on knowledge graph

The invention relates to the technical field of document processing, in particular to a knowledge graph-based document intelligent writing and analysis system, which comprises a graph dynamic updating module, an entity recognition and mapping module, a context connection analysis module, a semantic structure rearrangement module and a semantic coherence verification module. According to the method, the real-time data source is adopted, the knowledge graph is dynamically updated and expanded, synchronization of document content and the current information trend is ensured, noun phrases and entity distribution in a text are accurately analyzed through intelligent entity recognition, the accuracy of information extraction is improved, the coincidence degree of keywords between paragraphs is calculated, and synonymous entities are intelligently inserted, so that the information extraction efficiency is improved. According to the method, the context connection quality of the document is improved, the document structure is automatically rearranged according to the correlation of the content, the logic presentation of the information is optimized, the reading experience is enhanced, and the overall semantic consistency of the document is ensured and the interpretation ambiguity is reduced through complex semantic coherence verification.
Owner:WUDU INTERNET (XIAMEN) INFORMATION TECHNOLOGY CO LTD

System and Method for Accurate Responses from Chatbots and LLMs

Systems and methods are described for obtaining accurate responses from large language models (LLMs) and chatbots, including for question and answering, exposition, and summarization. These systems and methods accomplish these objectives via use of noun phrase avoiding processes such as a noun phrase collision detection process, a query splitting process, and a topical splitting process as well as by use of formatted facts, formatted fact model correction interfaces (FF MCIs), bounded-scope deterministic (BSD) neural networks, processes and methods, and intelligent storage and retrieval (ISAR) systems and methods. These systems and methods avoid and bypass noun phrase collisions and correct for errors caused by noun phrase collisions so that hallucinations are eliminated from LLM responses.
Owner:ACURAI INC

Government affair application platform based on artificial intelligence model

The invention relates to the technical field of government affair intelligence, in particular to a government affair application platform based on an artificial intelligence model, and the platform comprises a semantic recognition module, a tag affiliation module, a rule screening module, a path generation module and a path verification module. According to the method, the subject-called structure for recognizing noun phrases and approval verbs in the government affair approval text is adopted, semantic boundaries are defined, the structure recognition accuracy is enhanced, keyword and field item comparison adopts a field-level consistency mode, the tag affiliation precision is improved, field value comparison logic is introduced in rule node screening, and the accuracy of tag affiliation is improved. In path generation, a path structure with clear logic is established through sorting index and field dependence identification, field conflicts and path chaos are avoided, field value consistency verification ensures that path content is highly matched with an approval text, and it is guaranteed that the path is effective and available; automatic processing from text analysis and field affiliation to path verification is achieved, and the efficiency of the government affair approval process is improved.
Owner:SHENZHEN YUNHENG INTELLIGENT CO LTD

Multi-document key phrase extraction method based on graph structure node influence

The invention provides a multi-document key phrase extraction method based on graph structure node influence, and relates to the technical field of natural language processing and text mining. Firstly, a candidate phrase set is generated through noun phrase extraction and standardization; secondly, a semantic relation between phrases is captured through local subgraph construction and a sliding window mechanism, and the semantic relation is integrated into a global phrase co-occurrence graph; thirdly, dynamically dividing theme communities based on two-dimensional structure entropy minimization and a potential game model, and identifying phrase groups with high semantic aggregation; then, cross-topic nodes are processed through a structure entropy heuristic function, and flexibility of topic division is enhanced; and finally, in combination with node influence sorting, extracting key phrases with theme representativeness and propagation capability. The method does not need to label data, is suitable for multiple fields of academic literatures, news texts and the like, has high efficiency, accuracy and universality, and provides an innovative solution for multi-document key phrase extraction.
Owner:YUNNAN POWER GRID CO LTD +1

Open domain text information extraction method based on knowledge injection and graph neural network

The invention relates to the technical field of natural language processing, and discloses a knowledge injection and graph neural network-based open domain text information extraction method, which comprises the following steps of: extracting all noun phrases from input text data to construct a candidate entity set; combining the candidate entities in pairs, and constructing a self-attention incidence matrix of each entity pair; performing sequence sampling on the self-attention incidence matrix to generate a candidate triple sequence set; calculating semantic similarity between the candidate triple sequence and the input text data, and outputting the first k high-correlation triple sequences as initial information extraction results of the input text data; and performing dependency structure analysis on the initial information extraction result based on a graph neural network, and generating a triple sequence through redundant sequence labeling as a final information extraction result. According to the method, the recognition rate of the complex syntactic structure triad in the open domain information extraction task is remarkably improved, and meanwhile, the redundancy of the extraction result is effectively reduced.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Chapter-level relation extraction method and device, electronic equipment and storage medium

ActiveCN115618846BSemantic analysisInference methodsAlgorithmNoun phrase
The application provides a passage-level relation extraction method and device, electronic equipment and a storage medium, wherein the passage-level relation extraction method comprises: obtaining a target passage text, the target passage text being a passage text identified with pronouns and noun phrase references aligned with the pronouns; based on the target passage text and a preset relation extraction model, obtaining semantic relations between different entity pairs in the target passage text and relation categories of the semantic relations. The application can effectively improve the accuracy and reliability of passage-level relation extraction.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Unified cascade panoramic narrative detection and segmentation method

A unified cascaded approach for panoramic narrative detection and segmentation involves 1) multimodal encoding; 2) multimodal interaction; 3) coordinate-guided aggregation (CGA); 4) centroid-driven localization (BDL); and 5) training loss. A unified framework based on dynamic kernels is constructed, where a learnable kernel is built for each noun phrase to predict its corresponding mask and bounding box. To address the prediction conflict issue, two new cascaded modules are proposed to continuously process segmentation and detection to achieve cross-task alignment: a coordinate-guided aggregation (CGA) module and a centroid-driven localization (BDL) module. Using the centroid of the segmentation mask as an anchor, segmentation and detection are connected in series, naturally aligning the two tasks. The combination of the two modules mutually promotes their respective performance: the position information of the mask drives the BDL module to produce accurate boxes, while the backward guidance of the BDL module promotes the CGA module's ability to distinguish different instances during training.
Owner:XIAMEN UNIV

A complete knowledge graph construction method based on multi-source vulnerability data

The present invention discloses a method for constructing a complete knowledge graph based on multi-source vulnerability data, which mainly solves the problems of complex vulnerability data sources and missing relationships in the vulnerability database. The method includes: 1) collecting multi-source vulnerability data, where the data comes from unstructured vulnerability data recorded in CVE, CWE, CAPEC, and the security community; 2) normalizing sentences describing unstructured vulnerability information, performing sentence boundary detection and noun phrase normalization; 3) labeling the processed sentences with semantic roles, extracting sentence role construction data triples, and generating a vulnerability knowledge graph; 4) using node2vec to represent the nodes and relationships in the vulnerability knowledge graph into a low-dimensional dense vector space, and performing similarity calculation on the graph embedding results; 6) completing missing relationships through link prediction to obtain a complete knowledge graph. The present invention can establish a complete vulnerability database and effectively improve vulnerability retrieval efficiency.
Owner:XIDIAN UNIV +1

Systems and methods of incident resolution using chunking, vector embedding, and clustering techniques

A method for finding historically similar incidents includes receiving a plurality of historical data objects corresponding to a plurality of previous incidents, each of the plurality of historical data objects indicating an occurrence of a previous incident and including a historical resolution text description; generating a historical embedding of each of the plurality of historical data objects; extracting noun phrases from each of the historical resolution text descriptions; applying topic modeling to the extracted noun phrases; receiving a current data object indicating an occurrence of a current incident associated with a configurable item, the current data object including an incident description; generating a current embedding of the current data object; extracting noun phrases from the current embedding; applying topic modeling to the extracted noun phrase of the current data object; and identifying a set of historically similar incidents by applying a Euclidean distance formula to the current embedding.
Owner:FIDELITY INFORMATION SERVICES LLC

A patent tree and a method of constructing the same

The application relates to the technical field of patent data management, in particular to a patent tree and a construction method thereof, which comprises the following steps: based on patent text data, the technical features in the patent text are analyzed, the technical features of each patent are classified and integrated, and a patent technical feature set is obtained. In the application, the accurate positioning of patent technical elements is realized through intelligent information extraction and semantic analysis, the deep semantic analysis of verbs and noun phrases is combined, the technical target and implementation mode of the patent are extracted in detail, and the functions are decomposed and integrated, so that the core content of the patent is presented in a clear and structured manner. Further, the technical features are segmented and integrated according to time nodes and application fields, and combined with time sequence fluctuation characteristic analysis, historical dependent features are identified, and the development state of the patent technology can be effectively predicted. Through feature matching and similarity calculation, the accurate grouping of the patent and the analysis of cross-field feature combination can be realized.
Owner:GUANGDONG ZHIDELI NETWORK TECH CO LTD

Systems and methods of incident resolution using chunking, vector embedding, and clustering techniques

A method for finding historically similar incidents includes receiving a plurality of historical data objects corresponding to a plurality of previous incidents, each of the plurality of historical data objects indicating an occurrence of a previous incident and including a historical resolution text description; generating a historical embedding of each of the plurality of historical data objects; extracting noun phrases from each of the historical resolution text descriptions; applying topic modeling to the extracted noun phrases; receiving a current data object indicating an occurrence of a current incident associated with a configurable item, the current data object including an incident description; generating a current embedding of the current data object; extracting noun phrases from the current embedding; applying topic modeling to the extracted noun phrase of the current data object; and identifying a set of historically similar incidents by applying a Euclidean distance formula to the current embedding.
Owner:FIDELITY INFORMATION SERVICES LLC

A multi-scale visual positioning method and system based on semantic consistency guidance

The application relates to the technical field of visual language fusion positioning, and discloses a multi-scale visual positioning method and system based on semantic consistency guidance, which comprises the following steps: extracting a noun phrase based on a Stanza syntax analysis model, generating text semantic features, phrase semantic features and visual features respectively based on a pre-trained BEiT-3 image-text encoding model; calculating a text consistency constraint loss through the attention interaction of the noun phrase features and the visual features, and simultaneously generating a concept-level semantic heat map; based on the semantic heat map, applying weights corresponding to semantic responses to multi-scale visual features generated by a feature pyramid network; and according to the multi-scale visual features generated after weighting, carrying out adaptive sampling based on an offset amount on a candidate region through a deformable attention module and completing multi-scale candidate frame generation; the application solves the problem that the candidate region generation of an existing model lacks semantic directionality, and improves the problems of feature redundancy and unstable semantic alignment.
Owner:HUNAN NORMAL UNIVERSITY

A method and system for entity relation extraction from open-domain text

This invention proposes a method and system for entity relation extraction from open-domain text, comprising: acquiring labeled text as training data; extracting all named entities and noun phrases from the training data using entity recognition and performing data augmentation; training a neural network model using the augmented data as input to obtain an entity relation classification model; statistically analyzing the word frequencies of each named entity and noun phrase in the augmented data, and marking named entities and noun phrases with word frequencies greater than a preset value as filter words; acquiring the open-domain text and its corresponding head entities; extracting named entities and noun phrases from the open-domain text excluding filter words and inputting them into the entity relation classification model to obtain the entity relations of the open-domain text. Through effective data augmentation strategies, without incurring additional costs, this method effectively solves the problem of poor performance in practical applications of entity relation recognition due to noise caused by candidate tail entities.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

System and method for accurate responses from chatbots and llms

PCT designated stageWO2025193562A1Semantic analysisKnowledge representationCollision detectionNoun phrase
Systems and methods are described for obtaining accurate responses from large language models (LLMs) and chatbots, including for question and answering, exposition, and summarization. These systems and methods accomplish these objectives via use of noun phrase avoiding processes such as a noun phrase collision detection process, a query splitting process, and a topical splitting process as well as by use of formatted facts, formatted fact model correction interfaces (FF MCIs), bounded-scope deterministic (BSD) neural networks, processes and methods, and intelligent storage and retrieval (ISAR) systems and methods. These systems and methods avoid and bypass noun phrase collisions and correct for errors caused by noun phrase collisions so that hallucinations are eliminated from LLM responses.
Owner:ACURAI INC

Automatically generated product recommendations based on questions and answers

An automated technique for enriching presented answers by highlighting related shopping recommendations is disclosed. Shopping recommendations can be highlighted within the answers themselves, and can also be used as an auxiliary list of suggestions. A model is described for selecting phrases from answer text, referred to as a continuous sequence of lexical items of noun phrases, relating to potential products that may represent relevant shopping recommendations in the context of question and answer pairs. And sorting the noun phrases according to an importance sequence. The noun phrases ranked first are used for searching for products displayed in association with the noun phrases. Clicking or tapping the highlighted noun phrases initiates a shopping-related process, such as presenting gadgets with product recommendations or running a search in a search engine.
Owner:AMAZON TECH INC

Multimodal entity and relationship extraction method and system based on cross-modal alignment and fusion

The present invention discloses a multimodal entity and relationship extraction method and system based on cross-modal alignment and fusion, comprising: processing and encoding input text and images to obtain a variety of image and text features; using the semantic representation of the image as an anchor point, aligning fine-grained and coarse-grained text features with pixel-level image representations, respectively, and mapping the image and text features to the same semantic space; performing multi-granularity feature fusion through text-guided dynamic gating aggregation, visual prefix cross-modal fusion, and cross-modal image-text matching, thereby increasing feature complementarity and modeling the association between noun phrases in the text and image objects to obtain a multi-granularity multimodal feature representation; fusing multi-granularity multimodal features through entity-guided attention gating, aggregating visual information related to text entities, and obtaining a final multimodal fusion representation; and performing task predictions for multimodal named entity recognition and multimodal relationship extraction based on the multimodal fusion representation.
Owner:YANBIAN UNIV

Text-to-person retrieval method based on bounding box extraction and semantic consistency constraints

This paper discloses a cross-modal text-to-person retrieval method based on bounding box extraction and semantic consistency constraints. The method comprises the following steps: extracting fine-grained bounding boxes from images; extracting fine-grained noun phrases from text; generating a training set; constructing a fine-grained aggregation network; training the fine-grained aggregation network; and searching for people using text. This paper constructs a text-to-person retrieval model based on bounding box extraction and semantic consistency constraints. This method leverages the visual language knowledge from existing large-scale pre-trained models (GLIP and CLIP) and uses textual cues and GLIP to accurately extract key local features for identifying pedestrians, improving the accuracy of pedestrian retrieval. CLIP is used to extract visual and linguistic features to obtain a more comprehensive semantic representation. Finally, a constraint method is designed to maintain semantic consistency of features, reducing noise interference and improving the stability of pedestrian retrieval.
Owner:XIDIAN UNIV

Method for generating training data for training machine learning models and apparatus for generating training data using the same

This invention provides a method and apparatus for generating training data for training machine learning models. [Solution] The method involves acquiring the original image S100, generating an image caption for the original image S211, extracting noun phrases S212, generating a first pseudo-label including a first category name and a first bounding box S214, extracting proposals corresponding to objects from the original image S221, generating a region description S222, generating a second pseudo-label with a second category name and a second bounding box S223, filtering the first pseudo-label and the second pseudo-label according to pre-set filtering conditions to generate an integrated pseudo-label S300, and annotating the original image to generate training data S400.
Owner:SUPERB AI CO LTD

Non-transitory computer-readable recording medium, text generation method, and text generation device

A non-transitory computer-readable recording medium stores therein a text generation program that causes a computer to execute a process including acquiring a first text serving as a norm and a second text related to a case example, first generating graph data of the second text including noun phrases included in the second text and information about a relation between the noun phrases in the second text, based on the second text, and first inputting a prompt including the graph data of the second text generated, and the first text, to a large-scale language model to generate a third text satisfying a requirement defined in the first text.
Owner:FUJITSU LTD

A method for mapping noun phrases to description logic concepts based on externalization

The method for mapping a noun phrase to a description logic concept based on epitaxy firstly exhaustively lists all text segments of the noun phrase, generates a mapping table of the text segments to resources in a knowledge base; then generates an analysis sequence according to the word segmentation, part-of-speech tagging and syntax tree of the noun phrase; and finally, according to the analysis sequence, continuously refines basic concepts generated by the indexed resources from the concept of EL++, until all words are analyzed, to obtain the description logic concept to which the noun phrase is mapped. The application can automatically process complex noun phrases containing implicit relations and generate high-quality description logic concepts through analysis of the syntax tree.
Owner:NANJING UNIV

An English tweet named entity extraction method and device based on subjective and objective word lists

The present invention relates to a method and device for extracting named entities from English tweets based on subjective and objective word lists, belonging to the technical field of data processing; it solves the problem that a large number of subjective words in English tweets affect the performance of subsequent named entity recognition; the named entity extraction method of the present invention includes the following steps: obtaining English texts in multiple fields to construct a corpus; performing word segmentation and word frequency statistics on the texts in the corpus, and constructing a subjective word list through screening; preprocessing the English tweet to be recognized to obtain a standard tweet text; using a syntactic dependency analysis model to extract all noun phrases in the standard tweet text, and preprocessing the noun phrases based on the subjective word list to construct a noun phrase set NP p ; based on the noun phrases in the noun phrase set NP p construct a tree-shaped parent-child structure and perform named entity extraction to obtain the named entity recognition result of the English tweet.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

Generating recommendations by using communicative discourse trees of conversations

Techniques are disclosed for improved autonomous agents that can provide a recommendation in a non-intrusive, conversational manner. In an aspect, a method determines a first sentiment score for a first utterance and a second sentiment score for a second utterance, each sentiment score indicating an emotion indicated by the respective utterance. The method further identifies that a difference between the first sentiment score and the second sentiment score is greater than a threshold. The method further extracts a noun phrase from the second utterance. The method identifies a text fragment that includes an entity that corresponds to the noun phrase. The method identifies that the text fragment addresses a claim of the second utterance. The method forms a third utterance that includes the a recommendation related to the second utterance and adds the third utterance to the sequence of utterances after the second utterance.
Owner:ORACLE INT CORP

Multi-scale visual positioning method and system based on semantic consistency guidance

The invention relates to the technical field of visual language fusion positioning, and discloses a multi-scale visual positioning method and system based on semantic consistency guidance, and the method comprises the steps: extracting noun phrases based on a Stanza syntactic analysis model, and respectively generating text semantic features, phrase semantic features and visual features based on a pre-trained BEiT-3 image-text coding model; text consistency constraint loss is calculated through attention interaction of noun phrase features and visual features, and meanwhile a concept-level semantic thermodynamic diagram is generated; based on the semantic thermodynamic diagram, applying a weight corresponding to the semantic response to a multi-scale visual feature generated by the feature pyramid network; according to the multi-scale visual features generated after weighting, adaptive sampling based on offset is carried out on the candidate area through a deformable attention module, and multi-scale candidate box generation is completed; according to the method, the problem that existing model candidate region generation lacks semantic directivity is solved, and the problems of feature redundancy and unstable semantic alignment are improved.
Owner:HUNAN NORMAL UNIVERSITY

Intelligent generation system for resume file talent portrait report based on AI large model

The invention relates to the technical field of named entity recognition, in particular to an intelligent resume file talent portrait report generation system based on an AI large model, and the system comprises a resume file analysis module which converts a Chinese resume file format; the resume information analysis module performs public word replacement on each Chinese resume text and analyzes the word group nesting probability of word groups corresponding to nouns in public words; obtaining a supplementary resume set of each Chinese resume; based on the word group nesting probability of the word group corresponding to each noun in the supplementary resume set, determining an entity nesting confidence coefficient; a sparse resume sample is determined, the word group nesting probability of the word group corresponding to each noun in the webpage text is counted, and a search auxiliary discriminant value is determined; determining a nested entity screening value, and obtaining a nested long entity; a resume information extraction module obtains a Chinese resume named entity recognition result; and the talent portrait report generation module is used for generating a talent portrait report. The invention aims to improve the resume entity recognition accuracy.
Owner:JINAN KEJIN INFORMATION TECH CO LTD

Method for generating training data to be used for training machine learning model and training data generating device using the same

There is provided a method for generating training data for training a machine learning model by a training data generating device, comprising: (a) in response to acquiring an original image, performing sub-processes of: generating an image caption for the original image, extracting at least one noun phrase from the image caption, and performing an open vocabulary object detection on the original image, to thereby generate at least one first pseudo label and performing sub-processes of: extracting at least one proposal from the original image, generating at least one region description, and generating at least one second pseudo label; and (b) filtering the at least one first pseudo label and the at least one second pseudo label according to a preset filtering condition, to thereby generate at least one integrated pseudo label, and generating the training data by annotating the original image with the at least one integrated pseudo label.
Owner:SUPERB AI CO LTD

A method and system for literature semantic search based on elastic search

The application discloses a literature semantic search method and system based on elastic search, and relates to data retrieval. A text2vec-base-multilingual model trained based on a CoSENT method is used to perform vectorization processing on a list of noun phrases. According to semantic vectors, approximate nearest neighbor search is performed in a second retrieval module to obtain first candidate data. Query text data is input into the first retrieval module, and keyword matching is performed through a BM25 algorithm to obtain second candidate data. The first candidate data and the second candidate data are fused to obtain third candidate data. A Sequence Matcher algorithm is used to calculate the string similarity between extended words in the third candidate data, a similarity threshold is set based on the length of the longest common subsequence, and deduplication is performed to obtain fourth candidate data. Weight distribution based on position and similarity score is performed on the fourth candidate data, the score interval is expanded to enhance the distinguishability of the extended words, and an extended word recommendation list data is obtained. The application significantly improves the accuracy of literature retrieval.
Owner:CHINA EDUCATIONAL PUBLICATIONS IMPORT & EXPORT CORP LTD

An english news element extraction method based on core word diffusion

The application provides a WHO element extraction method in an English news scene, which comprises the following steps: cleaning and preprocessing data of network news, combining an article title, a lead and an article theme into one text data, and analyzing a tree structure of the text data; extracting all nouns in all noun phrases from a subtree of the text data tree structure, screening core words, recalling a core word set, calculating a characteristic value of each core word, and extracting a first core word in a sorting result as a final core word; expanding the core word into a complete element text, traversing leaf nodes corresponding to the core word, and finding simple sentence nodes in a parent node direction, connecting all texts from the simple sentence nodes to the leaf nodes and then to intermediate links, and finally generating the element text; and the application realizes extraction of WHO elements on multiple news samples, has high accuracy, can improve a recall rate of core words, and can expand a recall set.
Owner:XIAMEN UNIV