Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

13 results about "Contextual similarity" patented technology

Contextual Word Similarity is nothing but identifying different types of similarities between words. It is one of the goals of Natural Language Processing. Statistical approaches are used for computing the degree of similarity between words.

A method and device for constructing an incomplete speech rewriting model

The present application relates to a kind of incomplete speech rewriting model construction method and device, method includes: based on span dependency and insertion dependency dependency modeling and using node link resolution mode obtains the dependency graph of incomplete speech rewriting text editing operation;Using GPT model, the context similarity feature and / or rewriting consistency characteristic of current incomplete speech sentence are calculated, the context similarity feature and / or rewriting consistency characteristic are used to enhance the interactive inference of incomplete speech rewriting;The dependency graph score feature is fused with the context similarity feature and / or rewriting consistency characteristic, and the final feature after the feature fusion is pushed to rewrite result based on. More abundant semantic features can be provided for parsing model, and the speech rewriting effect is improved.
Owner:WUHAN UNIV

Code similarity detection method and device

The invention relates to the technical field of software analysis, and provides a code similarity detection method and device. The method comprises the following steps: determining context information of a query function according to a binary dependency graph of the query function; determining context information of each candidate function according to the binary dependency graph of each candidate function; determining context similarity between the query function and each candidate function according to the context information of the query function and the context information of each candidate function; according to the content of the query function and the content of each candidate function, determining the content similarity between the query function and each candidate function; selecting an initial similar function from the candidate functions according to the context similarity and the content similarity; and inputting the initial similarity function and the query function into the large language model to obtain a code similarity detection result output by the large language model. According to the code similarity detection method and device provided by the invention, the code similarity detection efficiency and accuracy are improved.
Owner:INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Text2SQL semantic caching method based on context and mode matching

The invention provides a Text2SQL semantic caching method based on context and pattern matching, and relates to the field of natural language processing and database interaction.The method comprises the steps that user query is input into a semantic compression module, semantic significant keywords are extracted to obtain a candidate set, a vocabulary set is extracted and constructed through a large language model, and a database is obtained; obtaining a compressed query through a syntactic organizer; inputting the historical dialogue into a context encoder, and obtaining global context representation through encoding and two-stage attention; performing coarse-grained filtering on the compressed query to form a newest candidate set; performing fine-grained context matching on the global context representation and the newest candidate set to obtain context similarity; and comparing the context similarity with a preset threshold value, judging whether the cache is hit or not, if the cache is hit, interacting the hit SQL statement with a database, and returning a query result. According to the method and the device, the problems of low SQL multiplexing accuracy and high response delay are solved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Geographic entity information processing method and device, storage medium and program product

The embodiment of the invention provides a geographic entity information processing method and device, a storage medium and a program product. In the embodiment of the invention, for at least two pieces of geographic entity description information, a geographic semantic feature encoder, a distance feature encoder and a geographic context feature encoder are respectively utilized to encode names and address semantics, spatial positions and surrounding geographic environments; obtaining a geographic semantic similarity, a spatial distance feature and a geographic context similarity; and inputting the geographic semantic similarity, the spatial distance feature and the geographic context similarity into a classification head network, and outputting a classification result of whether the at least two pieces of geographic entity description information point to the same geographic entity or not. Therefore, by introducing the model to learn the incidence relation between different features, full utilization of multi-dimensional features such as geographic semantics, spatial positions and surrounding geographic environments is realized, feature fusion and accurate judgment are realized in the classification head network, and the accuracy and robustness of geographic entity matching are improved.
Owner:TAOBAO CHINA SOFTWARE

A method for protecting cue words in large models based on differential privacy

This application provides a method for protecting prompt words in large models based on differential privacy. By constructing a perturbation mapping table, differential privacy perturbation is applied only to the first token encountered in a multi-turn dialogue, and the mapping relationship is stored. Subsequent repeated tokens directly reuse the historical perturbation results, thus avoiding the accumulation of privacy budget with the increase of dialogue turns. Simultaneously, an attention-based context-aware utility function is combined to integrate the static embedding similarity of tokens with dynamic context similarity to maintain semantic consistency across turns. Finally, a two-stage bucketing index mechanism is adopted, first bucketing candidate tokens according to utility, and then fine-sampling within the selected buckets, effectively mitigating the long-tail effect under large vocabulary and improving the stability of perturbation selection. This application achieves a balance between privacy protection strength and semantic usability in multi-turn dialogues, and only requires perturbing the input prompt words without accessing the model's internal architecture, making it suitable for black-box large language model inference scenarios.
Owner:XIDIAN UNIV

Area perception open set identification method based on causal class activation mapping

The invention discloses a causal class activation mapping-based region perception open set identification method, which comprises the following steps of: performing closed set training through a convolutional neural network by taking a training image of a known class as input, and generating class activation mapping reflecting causal contribution of a local region to a classification result based on region ablation analysis; the method comprises the following steps: in a training stage, extracting a region perception feature vector of a training image by utilizing a saliency map, and constructing a region prototype corresponding to a known category; in a test stage, region perception feature vectors are extracted from a test image, a region consistency score is calculated with the region prototype constructed in the training stage, and fusion judgment is carried out in combination with a global uncertainty score based on an energy model, so that known category identification and unknown category rejection are realized. According to the method, causal class activation mapping is introduced into an open set identification judgment process, so that the capability of distinguishing context similar unknown classes can be enhanced, and a regional interpretable basis is provided.
Owner:XIAN TECH UNIV

Content descriptor

An apparatus, method, system and computer-readable medium are provided for generating one or more descriptors that may potentially be associated with content, such as video or a segment of video. In some embodiments, a teaser for the content may be identified based on contextual similarity between words and / or phrases in the segment and one or more other segments, such as a previous segment. Text / characters may serve as a candidate descriptor(s). In some embodiments, one or more strings of characters or words may be compared with (pre-assigned) tags associated with the content, and if it is determined that the one or more strings or words match the tags within a threshold, the one or more strings or words may serve as a candidate descriptor(s). One or more candidate descriptor identification techniques may be combined.
Owner:COMCAST CABLE COMM LLC

System and method for artificial intelligence-based matchmaking

Disclosed is an artificial intelligence-based matchmaking system (100) for suggesting compatible user profiles on a graphical user interface (GUI). The system (100) includes a user device (102) comprising an input unit adapted to receive one or more input data from a user and an output unit communicatively coupled with a server (104) through a communication network (106). The server (104) includes a processing unit configured to receive and store profile inputs comprising textual, categorical, or behavioral data such as occupation, interests, and intent parameters, process the profile inputs to generate a prioritized list of candidate profiles based on semantic correlation, contextual similarity, and inferred intent, display the candidate profiles sequentially on the GUI, receive gesture inputs indicating interest or skip actions and update a dynamic parameter-weighted preference model to adaptively reorder subsequent profiles in real-time. The present disclosure also relates to a method (200) for artificial intelligence-based matchmaking system (100).
Owner:YARASI MUNUSWAMY RAGAVENDRA SWAMY +1

A knowledge graph pipeline multi-strategy entity disambiguation incremental alignment method and system

PendingCN122432352AEngineeringKnowledge graph
The application discloses a kind of knowledge graph pipeline multi-strategy entity disambiguation incremental alignment method and system.The method normalizes candidate entity and generates fingerprint, to block key in bucket generation candidate entity pair, comprehensive similarity is obtained by fusing character, semantic embedding and structural context similarity, when comprehensive similarity is not lower than class correlation threshold, similar graph is constructed and representative entity is determined, only incremental alignment is executed to change fingerprint set and historical fingerprint index, and output incremental change set and conflict details.Compared with full amount pairwise comparison, the application reduces the scale of candidate pair and improves the processing efficiency in continuous updating scenario.
Owner:HANGZHOU BUSINESS ENTERPRISE HUITONG NETWORK TECHNOLOGY CO LTD

A text2sql semantic caching method based on context and pattern matching

The application provides a Text2SQL semantic caching method based on context and pattern matching, relates to the field of natural language processing and database interaction, and comprises the following steps: inputting a user query into a semantic compression module, extracting semantic significant keywords to obtain a candidate set, extracting and constructing a vocabulary set through a large language model, and obtaining a compressed query through a syntax organizer; inputting historical dialogues into a context encoder, obtaining a global context representation through coding and two-stage attention; performing coarse-grained filtering on the compressed query to form a latest candidate set; performing fine-grained context matching on the global context representation and the latest candidate set to obtain a context similarity; comparing the context similarity with a preset threshold to determine whether the cache hits, and if the cache hits, interacting with a database with the hit SQL statement and returning a query result. The application solves the problems of low SQL reuse accuracy and high response delay.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Transform optimization method and system based on token graph query propagation mechanism

This invention relates to the field of deep learning technology, particularly to a Transformer optimization method and system based on a token graph-based query propagation mechanism. The method comprises an input processing layer, a dynamic graph construction layer, a propagation control layer, a computational optimization layer, a feedback adjustment layer, and an output integration layer. This invention captures the contextual similarity or relationship between tokens by constructing a token-graph (dynamic or static structure) and introduces a stochastic propagation of Q vectors mechanism. Only the seed token generates a complete Q vector, while the remaining tokens generate approximate Q vectors through graph propagation. Furthermore, K and V are projected only onto the central node, and a local error control feedback module is included. This solves the problems of computational resource constraints, memory limitations, and accuracy stability in long sequence processing of Transformers. It provides a practical and feasible technical path for its application in complex tasks such as long document understanding, multi-turn dialogue, and code generation.
Owner:HANGZHOU DIANZI UNIVERSTIY INFORMATION ENG SCHOOL

A multi-system data unified logistics asset internet of things monitoring processing system

The present application relates to the technical field of logistics asset internet of things monitoring and data unification, in particular to a multi-system data unified logistics asset internet of things monitoring processing system, comprising a graph construction unit, used for real-time acquisition of multi-source heterogeneous data streams, and for aggregated generation of dynamic time series graph snapshots; a heuristic quantification unit, used for taking scores as pseudo labels; a time series graph neural network unit, used for generation of high-dimensional state vectors for independent nodes; a model training unit, used for making the similarity of high-dimensional state vectors in vector space reproduce heuristic contextual similarity; a unified asset aggregation unit, used for solving high-dimensional state vectors and aggregating them into a unified asset view; and a hidden event inference unit, used for generation of hidden event alarms, and for determination of normal state when an abnormal inference score does not exceed an alarm threshold. The present application solves the technical problem that business system identification and internet of things identification need to be manually annotated to be aligned, and realizes automatic association of heterogeneous data and construction of a unified asset view.
Owner:ANWOOD LOGISTICS SYSTEMS (SUZHOU) CO LTD