Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

34 results about "Pairwise similarity" patented technology

Pairwise similarity score provides a relevant measure of similarity between protein sequences. This similarity incorporates biological knowledge about proteins and it is extremely powerful when combined with support vector machine to predict PPI.

Pair-wise graph querying, merging, and computing for account linking

There are provided systems and methods for pairwise graph querying, merging, and computing for account linking. A service provider may provide an account graph system to identify pairwise similarities between different accounts based on shared data that may be identified through one or more linking characteristics. When providing pairwise graph similarities, a service provider may receive a query identifying two or more accounts and / or an account with a parameter for graph exploration and querying. The service provider may utilize connection, link, or relationship graphs, queried and generated using a graph database, to determine pairwise similarities between the designated seed account and one or more selected accounts. The graph may include vertices for different queried data points and edges connecting such queries, where directionality of the edges or other vectors may be used to identify links or hops between accounts for data querying and exploration.
Owner:PAYPAL INC

Course map fusion processing method and device, electronic equipment and storage medium

The invention provides a course map fusion processing method and device, equipment and a medium, and the method comprises the steps: obtaining a plurality of knowledge points of a first course file and a plurality of knowledge points of a second course file, and enabling the knowledge points to be configured with index information; the index information comprises course names and chapter levels; dividing a plurality of knowledge point groups according to the pairwise similarity of each knowledge point; fusing the knowledge points in each knowledge point group to obtain a fusion result of each knowledge point group; and according to the fusion result of each knowledge point group and the index information, counting the fusion number of knowledge points of the first course file and the second course file under different chapter levels, and generating a thermodynamic diagram according to the fusion number. According to the method, high-level generation of fused new knowledge points with teaching value is realized, and the knowledge graph structure of the course is optimized, so that the high-level system structure of the course is better reflected, the system structure and association of the course are mined, and personalized teaching suggestions are given.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Human-ai collaborative prompt engineering

Embodiments herein describe techniques for optimizing prompts describing tasks for a large language model (LLM), to enable effective and efficient operations to optimize prompt generation and selection for various tasks in LLMs through human-AI collaboration. In an embodiment, a computing system applies a set of candidate prompts to the LLM, evaluates responses to the candidate prompts received from the LLM and calculates a pairwise similarity matrix of responses based on the received responses to the candidate prompts. The computing system evaluates the pairwise similarity matrix to determine whether or not to present a prompt to receive a user feedback input. The described techniques can enable enhanced processing speed, reducing an overall computer system time typically required for implementing optimized prompt generation, and enhancing performance of the computing system executing the LLM.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Classifier-guided dataset compression using distribution-aware selection

An example operation may include at least one of determining, by a transformer encoder trained on annotated image-text data, first latents for a first dataset stored in a memory, and second latents for a second dataset stored in the memory, generating a similarity matrix based on comparisons between the first latents and the second latents, constructing a graph comprising nodes corresponding to the first latents and edges based on pairwise similarity exceeding a threshold, identifying connected components in the graph and selecting, from each component, at least one latent having a highest score from a classifier trained to approximate divergence between the first dataset and the second dataset, forming a reduced dataset comprising the at least one latent, providing the reduced dataset to a model training module, and training an image classifier using the reduced dataset and the second dataset.
Owner:THE TORONTO DOMINION BANK

Similarity sensitive diversity

Similarity sensitive diversity is utilized to measure variation in a distribution of item listings along one or more categories. A similarity between category vectors of each category pair in a set of categories is determined and utilized to generate a pairwise similarity matrix. The pairwise similarity matrix may be pruned to remove category pairs below a threshold. Utilizing the pairwise similarity matrix, similarity sensitive diversity between one or more items of a plurality of items may be determined. In various aspects, the similarity sensitive diversity may be utilized to: generate a list of relevant items in an appropriate distribution, suggest refinements of a search query; generate navigation modules; categorize or recategorize the plurality of items; or generate autosuggestions.
Owner:EBAY INC

System for carrying out snoRNA and disease association prediction based on multi-graph SAGE network

PendingCN120544677ABiostatisticsSequence analysisFeature extractionTanimoto coefficient
The invention discloses a system for carrying out snoRNA and disease association prediction based on a multi-graph SAGE network, and relates to the technical field of biological information, the system comprises the following components: a data acquisition and preprocessing module, which obtains data containing snoRNA and disease association information from MNDRv3.1, carries out feature extraction on a snoRNA sequence, calculates snoRNA pairwise similarity by using a Tanimoto coefficient, and sends the snoRNA pairwise similarity to the MNDRv3.1; meanwhile, the disease pairwise similarity is obtained by adopting a DAG-based semantic similarity method; according to the method, the multi-graph heterogeneous network is constructed by integrating the paired similarity data of the snoRNA and the disease and the associated network of the snoRNA and the disease, the feature information of the snoRNA and the disease is comprehensively captured by using the GraphSAGE architecture with an attention mechanism, the complex relationship between the snoRNA and the disease is effectively captured, the prediction precision of snoRNA-disease association is remarkably improved, and the prediction efficiency of the snoRNA-disease association is improved. According to the method, higher AUC, AUPRC and F1 scores are obtained on a test set, higher classification capability and prediction reliability are displayed, and a more accurate prediction result is provided for biomedical researchers.
Owner:SOUTHWEST MEDICAL UNIV

Method for monitoring a work system and system

A method for monitoring a work system including steps of receiving a plurality of scan data values acquired by a sensor of a sensor device and order indicators associated with the scan data values, determining pairwise similarities between the scan data values, determining a number of clusters for grouping the scan data values based on the determined pairwise similarities, grouping the scan data values into N clusters, wherein N is the number of clusters determined based on the determined pairwise similarities, assigning the scan data values to the clusters obtaining a grouped dataset of the scan data values, and outputting a signal based on the grouped dataset. Further, a system and a computer program are shown.
Owner:WORKAROUND GMBH

Media retrieval with semi-supervised contrastive learning

Methods and apparatus for media retrieval. In some examples a method of media retrieval includes receiving an input media dataset of a first modality and obtaining, with a first-modality encoder, a first embedding representing the input media dataset in a latent space. The method also includes selecting a subset of second embeddings based on a similarity metric configured to quantify pairwise similarity between different embeddings in the latent space. Each of the second embeddings represents, in the latent space, a respective candidate media dataset of a different second modality. The second embeddings are obtained using a second-modality encoder. The method also includes identifying one or more output media datasets of the different second modality among the respective candidate media datasets based on the selected subset of the second embeddings. The first- and second-modality encoders are jointly trained using semi-supervised contrastive learning.
Owner:DOLBY LABORATORIES LICENSING CORP

Multimodal AI model protection using embeddings

Techniques for assessing multi-modal inputs to a machine learning model involve receiving a multimodal input containing an image, producing several transformed versions of that image, and generating embeddings for both the original and transformed images. A pairwise similarity analysis among all embeddings is conducted to determine distance values. Two dissimilarity metrics can then be calculated: one reflecting the differences among the transformed images, and another comparing the original image to its transformed versions. If the dissimilarity among the transformed images is greater than that between the original and transformed images plus a threshold, the system triggers a remediation action. This action either blocks the input from being processed by the machine learning model or prevents the model's output from being returned to the requester, thereby enhancing the reliability and security of the model.
Owner:HIDDENLAYER INC

Method and system for ballistic specimen clustering

There are described a method and a system for generating clusters of ballistic specimens. Using an image acquisition tool, topographic data for at least three ballistic specimens of at least one region of interest is acquired. From the topographic data, at least one parameter characterizing optimal toolmark alignment is determined for every distinct pair of the at least three ballistic specimens, each ballistic specimen having a plurality of toolmarks formed thereon. From the topographic data, at least one pairwise similarity score associated with optimal toolmark alignment is determined for every distinct pair of the at least three ballistic specimens. At least one triplet-wise consistency measure indicative of a consistency of optimal toolmark alignment is determined for every distinct triplet of the at least three ballistic specimens. A cluster analysis is conducted based on the at least one similarity score and the at least one consistency measure to generate the clusters.
Owner:FORENSIC TECH (CANADA) INC +1

Retrieval enhancement generation system member leakage risk assessment method based on few queries

The invention discloses a member leakage risk assessment method and system for a retrieval enhancement generation system based on a small amount of queries. Under the black box condition, according to the target member document, generating a plurality of queries with equivalent semantics and diverse syntaxes by the auxiliary large language model, submitting the queries to a target retrieval enhancement generation system, and collecting responses; the response is encoded as a sentence vector, an average pairwise similarity is calculated and a rejection response is identified to apply a penalty, and a final risk assessment score is synthesized and compared with a threshold to assess the likelihood of the target document as a knowledge base member. The method is implemented under the complete black box condition, does not depend on the internal architecture or parameter information of the system, is low in query overhead, is easy to calibrate, and is suitable for compliance auditing, red team evaluation and security detection before online, so as to quantify the member privacy leakage risk and guide a protection strategy.
Owner:HUAZHONG NORMAL UNIV

Method and device for constructing quantifiable index system based on semantic extraction and alignment

This invention discloses a method and apparatus for constructing a quantifiable indicator system based on semantic extraction and alignment. The method includes: extracting explicit and implicit indicators from the target text to form an initial indicator set; converting the explicit and implicit indicators into semantic vectors and determining the pairwise similarity between all vectors; clustering the explicit and implicit indicators based on vector similarity to select core indicators to form a core indicator set; matching each core indicator in the core indicator set with the target dataset to obtain a target dataset that meets preset matching conditions; and constructing the indicator system based on the target dataset and the corresponding core indicators. This invention relies on large language models and natural language processing technology to achieve fully automatic construction, redundancy removal, and localization adaptation of the indicator system, improving the efficiency of indicator construction, reducing subjective intervention, and enhancing the objectivity and practicality of the system.
Owner:SHENZHEN RES INST OF BIG DATA +1

Super-network representation learning method based on multi-head attention mechanism applied to breeding

The application discloses a multi-head attention mechanism-based super network representation learning method applied to breeding, relates to the technical field of crop breeding, and comprises the following steps: constructing a crop heterogeneous super network, and constructing an initial embedding matrix corresponding to all nodes; obtaining a normalized node representation vector corresponding to each tuple and a similarity score of each tuple; calculating the pairwise similarity between each two nodes; constructing a joint loss function, and optimizing the parameters of a target model; generating a final representation vector of a node in the crop heterogeneous super network by using the target model after parameter optimization; and predicting a crop variety according to the final representation vector. The application avoids complex graph conversion and explicit high-order combination calculation, improves the calculation efficiency and scalability, can more comprehensively and accurately encode semantic and structural information of phenotypic traits, and thus provides reliable and high-quality data basis for crop variety prediction based on the final representation vector.
Owner:QINGHAI UNIVERSITY

Discovering Novel Artificial Neural Network Architectures

Methods, systems and apparatuses for discovering novel artificial neural network architectures (ANN) architecture are disclosed. One method includes calculating ANN architecture fingerprints including an ANN architecture fingerprint of each of a plurality of existing ANN architectures, creating a plurality of next-generation candidate ANN architectures, calculating a plurality of next-generation candidate ANN architecture fingerprints including an ANN architecture fingerprint of each of the plurality of next-generation candidate ANN architectures, calculating ANN architecture pairwise similarities between each of the plurality of existing ANN architectures and each of the plurality of next-generation candidate ANN architectures using the plurality of existing ANN architecture fingerprints and the plurality of next-generation candidate ANN architecture fingerprints, retraining each of the plurality of next-generation candidate ANN architectures on the training dataset, obtaining a performance score of each of the next-generation candidate ANN architectures, and calculating a fitness score for each of the next-generation candidate ANN architectures.
Owner:BLAIZE INC

A data feature screening method, device, equipment and medium

PendingCN122087403Aimprove accuracyExclude highly similar situationsScreening methodEngineering
This application discloses a method, apparatus, device, and medium for screening data features, relating to the field of data mining and analysis. The method includes: sequentially determining multiple information indicators for each candidate data feature based on a preset set of test threshold parameters; wherein the set of test threshold parameters is configured based on out-of-sample window data; screening the multiple candidate data features to obtain multiple screened data features based on the multiple information indicators and the set of test threshold parameters, combined with reference data features corresponding to each candidate data feature; wherein the reference data features are constructed based on the corresponding candidate data features; grouping the multiple screened data features according to the pairwise similarity between them to obtain multiple similar feature groups, and selecting a representative data feature for each similar feature group based on the multiple information indicators of the screened data features. By implementing this application, the accuracy of data feature screening can be improved.
Owner:E FUND MANAGEMENT CO LTD

System and method for cross-modal interaction based on pre-trained model

A method is provided for data processing performed by a processing system. The method comprises determining a set of first tokens for first data and a set of second token for second data, each token comprising information associated with a segment of the respective data, determining pair-wise similarities between the set of first tokens and the set of second tokens, each pair comprising a first token in the set of first tokens and a second token in the set of second tokens, determining, for each first token in the set of first tokens, a maximum similarity based on the determined pair-wise similarities between the respective first token and the second tokens in the set of second tokens, and determining a first similarity between the first data and the second data by aggregating the maximum similarities corresponding to the first tokens in the set of first set of tokens.
Owner:HUAWEI TECH CO LTD

Similarity sensitive diversity

Similarity sensitive diversity is utilized to measure variation in a distribution of item listings along one or more categories. A cosine similarity between category vectors of each category pair in a set of categories is determined and utilized to generate a pairwise similarity matrix. The pairwise similarity matrix may be pruned to remove category pairs below a threshold. Utilizing the pairwise similarity matrix, similarity sensitive diversity between one or more items of a plurality of items may be determined. In various aspects, the similarity sensitive diversity may be utilized to: generate a list of relevant items in an appropriate distribution, suggest refinements of a search query; generate navigation modules; categorize or recategorize the plurality of items; or generate autosuggestions.
Owner:EBAY INC

A method, system and rail vehicle for detecting axle box bearing faults

PendingCN122364956ABogiePairwise similarity
This invention discloses a method, system, and rail vehicle for detecting axle box bearing faults, relating to the field of rail transport safety technology. It extracts frequency domain characteristic waveforms from the axle box bearing's state signals, performs pairwise similarity comparisons of the frequency domain characteristic waveforms of different components at the same axle position on the same bogie, and compares the similarity of the frequency domain characteristic waveforms of the same components at the same axle position on different bogies. By combining the comparison results of different components within the same axle position and the comparison results of the same components between different bogies, fault determination is made. This enables the detection of lateral anomalies in multiple components under the same operating conditions. Compared to traditional methods that use single-component, fixed-threshold determination, this method not only adapts to different operating conditions but also eliminates the need to pre-establish health baselines for different components. It effectively improves the accuracy of axle box bearing fault detection in rail vehicles while simplifying deployment and increasing detection sensitivity.
Owner:CRRC QINGDAO SIFANG CO LTD

Similarity self-adaption-based semi-supervised open world radiation source individual identification method

PendingCN120561673AData setFeature extraction
The invention discloses a semi-supervised open world radiation source individual identification method based on similarity self-adaption, and aims to solve the problems that unmarked data cannot be effectively utilized and unknown classes cannot be subdivided in an open world scene in existing open set and semi-supervised radiation source individual identification. The method comprises the following steps: acquiring a data set containing marked samples and unmarked samples, and extracting sample features by using a feature extractor; clustering is converted into sample pair similarity prediction through pairwise similarity loss, and a new category is identified; identifying a known class by using self-adaptive cross entropy loss, and balancing the learning speeds of the new class and the known class; introducing an entropy regularization item to prevent model overfitting; and through minimizing an overall objective function training model formed by pairwise similarity loss, adaptive cross entropy loss and entropy regularization, identification of known and new radiation sources is realized. According to the method, finite marked data and a large amount of unmarked data are effectively utilized, known classes and new classes can be accurately identified, and the adaptability to new class radiation sources is enhanced.
Owner:ARMY ENG UNIV OF PLA

Apparatus, system, and method for grouping data records

The present application relates to apparatuses, systems, and methods for grouping data records based on entities referenced by the data records. The disclosed grouping mechanism can include determining pairwise similarities among a large number of data records, and clustering subsets of data records based on their pairwise similarities.
Owner:FACTUAL

A similar protein structure retrieval method based on hash learning

This invention discloses a method for retrieving similar protein structures based on hash learning: First, a protein structure dataset is acquired, and the pairwise similarity information is calculated. For the query sample, positive and negative samples are sampled. The protein structures are modeled as a graph, and node and edge features are extracted, the feature representations of nodes and edges are updated, quantization loss and similarity loss are defined, and the final loss function is obtained. The model is then trained. Based on the trained model, each collected protein structure is represented as a binary vector, resulting in a binary vector database. During retrieval, the trained model represents the new protein structure to be retrieved as a binary vector. Similar structures are retrieved by directly calculating the Hamming distance between binary vectors or by constructing an inverted index. Alternatively, the top-ranked protein structures returned based on binary vector retrieval can be reordered using other real-valued vectors or more complex algorithms. This invention reduces storage overhead and improves retrieval speed.
Owner:NANJING UNIV

Apparatus, systems, and methods for grouping data records

The present application relates to apparatus, systems, and methods for grouping data records based on entities referenced by the data records. The disclosed grouping mechanism can include determining a pair-wise similarity between a large number of data records, and clustering a subset of the data records based on their pair-wise similarity.
Owner:FOURSQUARE LABS INC

Graph based metadata structuring algorithm to enable machine learning

ActiveUS12682613B2Medical recordText string
In the disclosed systems and methods for categorizing medical data, a computer system obtains, in electronic form, a plurality of medical records. Each medical record includes corresponding medical data from a respective medical evaluation and corresponding metadata comprising a plurality of attributes about the respective medical evaluation. Each respective attribute comprises a corresponding string of text. The computer system determines, for each respective pair of medical records consisting of a first medical record and a second medical record, a corresponding pairwise similarity between, for each respective attribute in a set of attributes, the corresponding string of text for the first medical record and the corresponding string of text for the second medical record. The computer system identifies a first subset of the plurality of medical records. Each respective medical record in the first subset is connected to each other through pairwise similarities that each satisfies a similarity threshold.
Owner:TEMPUS AI INC

An intelligent vehicle networking stream data privacy protection method and system based on federated learning

The application provides a kind of intelligent vehicle networking flow data privacy protection method and system based on federal learning, by in the intelligent vehicle networking framework based on federal learning, the replacement of new and old training samples is carried out to the flow data generated in the framework using pair similarity strategy, and the total size of local training data is kept unchanged using dynamic data storage strategy;Among them, the pair similarity strategy is to replace the data points in the newly generated flow data to the training sample as the data points that need to be replaced, which satisfy the preset similarity condition;Dynamic data storage strategy is to store the newly generated flow data and the replaced old data points in the buffer area.The application uses pair similarity strategy and dynamic data storage strategy to replace similar samples in the original data set using samples in the flow data and keep the number of training samples unchanged, solves the problem of catastrophic forgetting by keeping the data distribution unchanged, and also ensures the effective convergence of the federal learning training process.
Owner:ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1

Quantifiable index system construction method and device based on semantic extraction and alignment

The invention discloses a quantifiable index system construction method and device based on semantic extraction and alignment, and the method comprises the steps: extracting dominant indexes and recessive indexes based on a target text, and forming an initial index set; respectively converting the dominant indexes and the implicit indexes into semantic vectors, and determining pairwise similarities among all the vectors; clustering the dominant indexes and the recessive indexes based on the vector similarity, and selecting core indexes to form a core index set; matching each core index in the core index set with a target data set to obtain a target data set meeting a preset matching condition; and constructing an index system based on the target data set and the corresponding core index. According to the method, full-automatic construction, redundancy elimination and localization adaptation of the index system are realized by relying on a large language model and a natural language processing technology, the index construction efficiency is improved, subjective intervention is reduced, and the objectivity and the landing performance of the system are enhanced.
Owner:SHENZHEN RES INST OF BIG DATA +1

Method for monitoring deformation of fan blade in non-stop state

The invention provides a method for monitoring fan blade deformation in a non-stop state, and belongs to the technical field of wind driven generator blade state monitoring, and the method comprises the following steps: using a wide-angle camera to shoot a to-be-monitored fan rotating blade in an upward shooting manner, extracting blade features frame by frame, and generating arc length and chord length time sequence data; the pairwise similarity of the time sequence data of the three leaves is calculated, and a similarity-time curve is generated; constructing a curvature-position-time three-dimensional space-time atlas; a wavelet decomposition-Kalman filter combined noise reduction algorithm is adopted for the abnormal feature sequence; and when the filtering energy value continuously exceeds a working condition self-adaptive threshold value, or the curvature of a specific position in the three-dimensional space-time atlas deviates from similar blades, judging that the corresponding blade is abnormal in deformation and outputting a positioning result. According to the method, under the condition that a fan is not stopped, automatic, high-precision and spatial positioning of blade deformation monitoring is achieved, and early diagnosis of blade deformation abnormity and damage spatial positioning are achieved in a complex operation environment.
Owner:WUHAN DITE ENERGY TECH CO LTD

Robot motion control method and system based on multi-task reinforcement learning

The invention relates to the technical field of deep reinforcement learning, and provides a robot motion control method and system based on multi-task reinforcement learning, and the method comprises the steps: dividing robot control tasks into corresponding task groups, each task group having an independent strategy network and an independent value network; the current robot state is obtained, a control instruction is obtained through the strategy network of the corresponding task group, the robot is driven to execute actions, environment feedback is obtained, the value network evaluates the value of the executed actions, the value is fed back to the strategy network, and strategy optimization is conducted; wherein the training of the strategy network and the value network comprises the steps of sampling tracks from a plurality of control tasks, calculating the similarity of every two control tasks, constructing a similarity matrix, grouping the control tasks through clustering, obtaining a plurality of task groups, and performing joint training of the strategy network and the value network by sharing parameters in the task groups. The stability and the overall performance of the reinforcement learning model are improved, and the robot control precision is improved.
Owner:SHANDONG UNIV

SYSTEMS AND METHODS FOR QUALITY ASSESSMENT FOR LARGE LANGUAGE MODELS (LLMs) BASED ON CONSISTENCY QUANTIFICATION

Systems and methods for LLM assessment are disclosed herein. Embodiments may provide a quality assessment of an LLM that is based on the consistency of that LLM. Utilizing response sampling, pairwise similarity, and an uncertainty score, embodiments may measure response consistency, which can be correlated with output quality for an LLM.
Owner:Q2 SOFTWARE

Methods of analyzing similarity of at least two samples of a plurality of samples comprising genomic DNA

The present application relates to a method for analyzing the similarity of at least two samples of a plurality of samples comprising genomic DNA. The method comprises the following steps: a) providing a plurality of samples comprising genomic DNA; b) performing a deterministic restriction site whole genome amplification (DRS-WGA) of the genomic DNA separately for each sample; c) preparing a massively parallel sequencing library from each product of DRS-WGA using a no-fragmentation, sequencing adapter / WGA fusion primer PCR reaction; d) performing low-pass whole genome sequencing of the massively parallel sequencing library at an average coverage depth of less than 1x; e) aligning the reads of each sample obtained in step d) to a reference genome; f) extracting the allele content at a plurality of polymorphic loci for each sample; g) calculating a pairwise similarity score locus for at least two samples based on the measured allele content at the plurality of loci; h) determining the similarity of at least two samples based on the similarity score, the method being used for non-invasive prenatal testing or diagnosis.
Owner:MENARINI SILICON BIOSYSTEMS SPA

Multi-modal Knowledge Graph Entity Alignment Method and System Based on Interaction between Modes

The present disclosure provides a multi-modal knowledge graph entity alignment method and system based on inter-modal interaction, which relates to the technical field of multi-modal knowledge graphs, and includes obtaining entity data of different modalities to be aligned in the knowledge graph; extracting the structural information features, visual information features, relationship information features, and attribute information features of the entity data; using a low-rank multi-modal fusion method to model the interaction between modalities for the obtained individual modal features, and then using a cross-modal attention mechanism to enable the individual modal features to learn the interaction between modalities from the low-rank fusion modality in parallel, so as to generate an overall entity feature representation; performing similarity contrast learning using different individual modal features and the overall entity feature representation, and updating the overall entity feature representation; calculating the pairwise similarity through the updated overall entity feature representation, and selecting the two entities with the highest similarity for entity alignment. The present disclosure can capture the interaction between multi-modal information of entities.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)