Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Levenshtein distance" patented technology

In information theory, linguistics and computer science, the Levenshtein distance is a string metric for measuring the difference between two sequences. Informally, the Levenshtein distance between two words is the minimum number of single-character edits (insertions, deletions or substitutions) required to change one word into the other. It is named after the Soviet mathematician Vladimir Levenshtein, who considered this distance in 1965.

Data asset classification and dynamic authority management system

The invention relates to the technical field of data security, in particular to a data asset classification and grading and authority dynamic management system, which comprises an asset parameter acquisition module, a classification feature judgment module, a grading weight calculation module and an authority dynamic adaptation module. According to the method, metadata and path parameters are automatically collected through a distributed crawler in combination with a regular expression, naming parameters and content feature parameters are separated to generate a structured parameter set, field naming similarity is quantized through a Levenshtein distance, content character distribution features are verified through chi-square verification, and a naming similarity and distribution feature dual verification mechanism is constructed. According to the method, sensitivity, access frequency and data volume weight are dynamically distributed through an entropy weight method, grading parameters are generated through linear superposition, an RBAC model is combined with a Dijkstra algorithm to verify access path topology legality, unauthorized access is blocked, and the heterogeneous data classification grading and authority control dynamic adaptive capacity is improved.
Owner:国义招标股份有限公司

Conversational Artificial Intelligence Regression Ensemble

A method of improving chatbot accuracy includes receiving an input test file containing chatbot interactions data including utterances, expected intents, and expected responses. For each interaction, the method may generate predicted intents and responses using an AI-based chatbot trained on previous interaction data. The method may determine similarity between predicted and expected responses / intents by calculating Levenshtein distances and comparing to predetermined thresholds. Failed interaction keywords may be extracted using natural language processing to identify failing topics. Machine learning decision tree classification may analyze patterns in failed interactions to generate resolution suggestions. The method may generate an interactive dashboard displaying historical performance trends, domain-specific accuracy metrics, visualizations of frequently keywords, and confidence distributions for failed interactions. This automated regression testing approach enables efficient identification and resolution of chatbot performance issues while maintaining response quality.
Owner:ELEVANCE HEALTH INC

Data exporting method based on Excel template

The invention discloses a data exporting method based on an Excel template. The data exporting method comprises the following steps: S1, configuring a user template; s2, dynamic data source connection; s3, extracting and mapping data; s4, intelligent field matching domain; s5, data cleaning and conversion; s6, file generation and distribution; through dynamic template analysis, intelligent field matching, multi-source secure connection and a multi-level verification alarm mechanism, three core problems of low efficiency, poor adaptability and weak security in a traditional data export technology are solved; the method comprises the following steps: by integrating an Apache POI library and an OpenCSV tool, covering a mainstream Excel file type; the template analysis engine ensures that the format of the exported file is completely consistent with the preset template; field similarity calculation based on the Levenshtein distance is combined with a historical use frequency weight, so that the matching accuracy is improved; and the operation and maintenance response efficiency is improved by a multi-dimensional verification and intelligent alarm system.
Owner:WUXI RONGZHI TECH CO LTD +1

Index question answering method and electronic equipment

The invention discloses an index question answering method and electronic equipment, and the method comprises the steps: determining an initial feature extraction entity corresponding to a to-be-answered question when the to-be-answered question is received, and the to-be-answered question is an index question; target similarities between the initial feature extraction entity and candidate entities and candidate entity relationships in the knowledge graph are determined, and the target similarities comprise a Lelwstein distance and a cosine similarity; determining a target entity and a target entity relationship from the knowledge graph according to the target similarity; and according to the target entity and the target entity relationship, generating a target answer corresponding to the to-be-answered question through a large language model. According to the technical scheme provided by the embodiment of the invention, the technical problems of low answering efficiency and low accuracy of index questions in the prior art are solved, and the purposes of improving the question answering efficiency and the accuracy of the target answer are achieved.
Owner:AGRICULTURAL BANK OF CHINA

Conversational artificial intelligence regression ensemble

PendingCN122663581AData packRegression testing
A method for improving chatbot accuracy includes receiving an input test file containing chatbot interaction data, the data including utterances, intended intents, and intended responses. For each interaction, a predicted intent and response can be generated using an AI chatbot trained on prior interaction data. The method determines the similarity between the predicted and intended responses / intents by calculating the Levenshtein distance and comparing it to a predetermined threshold. Keywords from failed interactions are extracted through natural language processing to identify failure themes. Machine learning decision tree classification can analyze patterns in failed interactions to generate solution recommendations. The method can generate an interactive dashboard displaying historical performance trends, accuracy metrics for specific domains, visualizations of high-frequency keywords, and confidence distributions for failed interactions. This automated regression testing method can effectively identify and resolve chatbot performance issues while maintaining response quality.
Owner:ELEVANCE HEALTH INC

IP network attack group classification method based on behavior sequence

The application discloses a kind of IP network attack group classification method based on behavior sequence, its steps are as follows: S101, obtain the malicious IP behavior data set of operator security protection platform;S102, data cleaning, and dictionary coding is carried out to behavior sequence, generate behavior sequence data;S103, to behavior sequence data, behavior frequent item calculation is carried out;S104, Levenshtein distance similarity calculation is carried out to behavior sequence, similarity range is divided into four intervals, for the IP value in the corresponding range in each interval, four intervals are used as characteristic field;S105, the appearance frequency of each instruction in all IP behavior sequences is counted as characteristic field;S106, all characteristic fields are handled by feature engineering, and are arranged into model input format;S107, using cluster analysis, obtain IP network attack group classification.The application realizes IP network attack group classification, and then solves the problem that existing technology cannot be applied to IP group classification scene without communication or associated attribute.
Owner:JIANGSU HONGXIN SYST INTEGRATION

Fuzzy address matching method based on parallel processing in electric power scene

The invention discloses a fuzzy address matching method based on parallel processing in an electric power scene. The method comprises the following steps: data preprocessing: preprocessing a marketing address and an intranet address in distributed photovoltaic grid-connected information; address grouping and region division; a parallel computing framework is introduced to decompose a matching task of a network address and a marketing address; calculating the similarity between the character strings by using a Levenshtein distance algorithm to obtain a similarity score between the intranet address and the marketing address; high-matching-degree result screening: according to the calculated similarity score and a preset high-score threshold value, screening out a high-score matching item as a preliminary result; carrying out leak checking and vacancy filling on a low-matching-degree result; and final result output: combining and outputting the preliminary result and a result obtained after leak checking and vacancy filling. According to the method, parallel tasks can be efficiently distributed and processed, the time complexity of data processing is reduced, the processing capacity of the system is enhanced, and address information with slightly different expressions but substantially identical expressions is effectively identified.
Owner:HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER +1

A data asset classification and dynamic management system of grading and permission

The application relates to the technical field of data security, in particular to a data asset classification and grading and dynamic management system of authority, which comprises an asset parameter acquisition module, a classification feature discrimination module, a grading weight calculation module and a dynamic adaptation module of authority. In the application, distributed crawlers are combined with regular expressions to automatically collect metadata and path parameters, naming parameters and content feature parameters are separated to generate a structured parameter set, Levenshtein distance is used to quantify the naming similarity, chi-square test is used to verify the content character distribution feature, a double-checking mechanism of naming similarity and distribution feature is constructed, the entropy weight method is used to dynamically distribute the sensitivity, access frequency and data volume weight, the grading parameters are generated through linear superposition, the RBAC model is combined with the Dijkstra algorithm to verify the legality of the access path topology, the over-authorization access is blocked, and the dynamic adaptation capability of heterogeneous data classification and grading and authority control is improved.
Owner:国义招标股份有限公司

Iterative chart code generation method based on chart and code correction large model

This invention discloses an iterative chart code generation method and apparatus based on a chart and a large-scale code correction model, belonging to the field of chart code generation technology. The method includes: based on a code error classification system, performing data mutation using a multimodal large-scale model based on the original chart code pairs, and constructing an original dataset; based on the code error classification system and the row-level Levenshtein distance algorithm, adding correction prompts to the original dataset, and constructing a training dataset; based on the LLaMA-Factory fine-tuning framework and the GRPO training strategy, performing two-stage training on the large-scale code correction model based on the training dataset to obtain an optimized large-scale code correction model; and based on the optimized large-scale code correction model and the large-scale chart code generation model, iteratively generating code based on a target reference chart to obtain the target generated code. This invention is a high-precision, interpretable, and iterative chart code generation method based on charts.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

A method and system for improving completion parameters for understanding user input

PendingCN122114174AAccurately capture potential needsavoid misreadingDigital data information retrievalNatural language data processingUser inputEngineering
The application relates to the technical field of industrial manufacturing, and discloses a method and system for improving the completion parameters of understanding user input, which comprises the following steps: inputting an industrial standardized file, generating a semantic label, reasoning through the industrial standardized file and the semantic label, obtaining an optimized industrial standardized file, and obtaining an industrial rule set and a text parameter set input by a user. The application realizes accurate semantic label generation based on a WordNet synonym set, solves the polysemy ambiguity problem, calculates the similarity between a candidate parameter and a user conversation through an improved Levenshtein distance algorithm and a compensation mechanism, accurately captures the potential demand of the user, improves the consistency with the actual input intention of the user, multiplies the initial weight value, the click rate and the time decay factor to sort the candidate parameters, balances the instant production demand and the historical experience in the industrial scene, and significantly improves the accuracy, adaptability, efficiency and interpretability in four dimensions.
Owner:BEIJING INFORMATION TECH BOTE INTELLIGENT TECH CO LTD

Polarization code belief propagation decoding method for correcting insertion and deletion errors

The invention discloses a polarization code belief propagation decoding method for correcting insertion and deletion errors, and relates to the field of digital communication error control coding. The method comprises the following steps: mixing information bits with the length of 1 and all-zero frozen bits with the length of 1 according to position indexes of preset information bits and fixed bits to obtain a bit sequence; encoding the bit sequence to obtain a sending sequence with the length of 1; the sending sequence is transmitted through an IDS channel to obtain a receiving sequence with the length of the IDS channel; and carrying out insertion and deletion error correction processing on the receiving sequence through a BP decoder based on a weighted Lelwstein distance, and outputting an information sequence estimated value. According to the method, the drift distance is introduced, and the weighted Levinstein distance is adopted as probability measurement, so that the problem that insertion and deletion errors cannot be corrected by a traditional BP decoding algorithm is solved. The inherent problem of an SC framework is also avoided, and the error correction capability is improved; and through parallel operation, the decoding time delay is reduced, and the unit time processing capacity is improved.
Owner:TIANJIN NORMAL UNIVERSITY

Power Internet of Things Network Security Risk Prediction Method Based on Levenshtein Distance Algorithm

The present invention relates to the field of network security prediction, specifically a method for predicting power Internet of Things network security risks based on the Levenshtein distance algorithm, which includes the following steps: First, the attack source IP, attack behavior, and attack target IP in a single alarm event are used as a piece of valid alarm information. For each piece of alarm information, taking the current alarm event as the result, six closest alarm events before the occurrence time of this alarm event are found as the cause, thereby constructing a causal data, and storing the causal data in the database to form a causal database; Second, filter the causal database; Finally, use the Levenshtein distance algorithm to predict the alarm event. It solves the problem of poor long-term prediction effect of the existing technology for the system.
Owner:INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER

A knowledge graph retrieval and classification method based on an improved semi-supervised classification model

The present invention relates to the field of information retrieval technology, and in particular to a knowledge graph retrieval and classification method based on an improved semi-supervised classification model. The method comprises: performing a pre-processing classification operation on graph nodes in a knowledge graph based on the improved semi-supervised classification model to determine the node features of the graph nodes; when a query keyword is received, calculating the similarity score between the query keyword and the node features based on a semantic matching algorithm combining the Jaccard coefficient and the Levenshtein distance similarity; determining whether a graph entity is matched based on the similarity score and a preset similarity threshold; if not, determining the graph node with the highest matching degree based on the similarity score; and performing an extended query action based on the graph node with the highest matching degree to obtain graph nodes and their attribute information related to the query keyword. This method achieves the purpose of improving the classification accuracy of graph nodes, improving the matching accuracy between query keywords and graph nodes, and improving the coverage and accuracy of retrieval.
Owner:YUNNAN DAILY NEWSPAPER GRP

Text difference degree calculation method and system

ActiveCN115455933BNatural language data processingCalculation methodsLevenshtein distance
The application provides a text difference degree calculation method and system, and the method comprises the following steps: obtaining the length of a first text and the length of a second text; in the case that the length of the first text and the length of the second text satisfy a first preset condition, calculating the difference degree between the first text and the second text according to a target Levenshtein distance between the first text and the second text, the length of the first text and the length of the second text. The application calculates the difference degree between the texts to be compared based on the length of the texts to be compared and the target Levenshtein distance, solves the problems of slow manual comparison speed and complicated operation in the case that the amount of text data is very large, improves the calculation efficiency of the difference degree of the texts to be compared, and enables the user to quickly understand the difference degree between the texts to be compared.
Owner:TRANSN IOL TECH CO LTD

A cross-system document intelligent coding method and system based on a dynamic rule engine

The application discloses a kind of cross-system document intelligent coding method and system based on dynamic rule engine, by obtaining the metadata information of to-be-coded document, call dynamic rule engine from configurable rule base and match coding rule, rule engine supports multi-dimensional condition matching and priority dynamic calculation, and provide visual configuration interface to realize the real-time adjustment and version management of rule;Adopt three-level mapping strategy to convert the metadata of PDMS, BIM, SCADA and other heterogeneous systems into unified equipment and facility list ID without loss, combined with improved Levenshtein distance algorithm to calculate the similarity of equipment name, form the dual conflict detection mechanism of equipment ID accurate matching and equipment name similarity calculation, trigger artificial review when conflict is detected, review result is fed back to knowledge base to realize self-learning optimization, while establishing coding life cycle management;The application realizes the automation, intelligentization and traceability of cross-system document coding, significantly improves coding consistency and standard adaptation efficiency.
Owner:CHINA THREE GORGES CORPORATION

Method and system for discriminating AI code generation based on Levenshtein distance

The invention discloses a Levenshtein distance-based AI code generation distinguishing method and system, and the method comprises the steps: carrying out the fixed covering of an original code segment by taking a line as a unit, and then completing the covered part through a code completion model, and obtaining a plurality of disturbance code segment samples; respectively calculating the Levenshtein distance between the original code segment and each disturbance sample, and averaging the Levenshtein distance to obtain an average Levenshtein distance; the obtained average Levenshtein distance is compared with a set threshold value, if the average Levenshtein distance is larger than the threshold value, it is considered that the original code segment belongs to machine generated codes, and otherwise, it is considered that the original code segment belongs to manually written codes. The method is a new code detection scheme based on zero samples, the effectiveness during detection can be improved, and the time and hardware resources needed during detection can be greatly reduced.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

String comparison device and method

An electronic device for record linkage including a memory storing one or more instructions, and a processor that executes the one or more instructions to generate one or more vectors of one or more strings from a reference database based on a modified Levenshtein distance, and generate a vector database for spelling similarity based on the one or more vectors. The modified Levenshtein distance is based on one or more parameters, including at least one of: a first number of insertions, a second number of deletions, a third number of replacements, or a fourth number of matches, one or more of the one or more parameters including a predefined weight, and one or more fixed strings.
Owner:VIDOORI INC +2

Method for multilingual learning of language models using rlhf using synthetic eye gaze trajectories

FIELD: computer technology.SUBSTANCE: method for multilingual training of language models using reinforcement learning based on human feedback using synthetic gaze trajectories, comprising the steps of feeding multilingual text to a gaze prediction model based on a multilingual eye movement corpus and a multilingual BERT model to generate a fixation sequence, computing visual attention features including first-pass regression rate, skip rate, first-pass and total fixation counts, and normalized Levenshtein distance, generating a gaze-aware reward model by projecting features into a latent space through a fully connected network, distributing rewards to tokens proportional to fixation probabilities, and optimizing the policy using PPO or GRPO algorithms with a modified advantage that takes into account the Levenshtein distance.EFFECT: multilingual training in thirteen languages without collecting real eye tracking data, achieving the accuracy of reproducing the features of visual attention.6 cl
Owner:AVTONOMNAYA NEKOMMERCHESKAYA ORGANIZATSIYA VYSSHEGO OBRAZOVANIYA UNIV INNOPOLIS

Conversational artificial intelligence regression ensemble

PCT designated stage expiredWO2025122868A1Semantic analysisTransmissionData packRegression testing
A method of improving chatbot accuracy includes receiving an input test file containing chatbot interactions data including utterances, expected intents, and expected responses. For each interaction, the method may generate predicted intents and responses using an Al-based chatbot trained on previous interaction data. The method may determine similarity between predicted and expected responses / intents by calculating Levenshtein distances and comparing to predetermined thresholds. Failed interaction keywords may be extracted using natural language processing to identify failing topics. Machine learning decision tree classification may analyze patterns in failed interactions to generate resolution suggestions. The method may generate an interactive dashboard displaying historical performance trends, domain-specific accuracy metrics, visualizations of frequently keywords, and confidence distributions for failed interactions. This automated regression testing approach enables efficient identification and resolution of chatbot performance issues while maintaining response quality.
Owner:ELEVANCE HEALTH INC

Embedded model architecture search method and system based on double-layer proxy assisted genetic algorithm

The invention belongs to the technical field of knowledge graph completion, and discloses an embedded model architecture search method based on a double-layer proxy assisted genetic algorithm. A local semantic enhancement module, an interactive semantic enhancement module, a global semantic enhancement module and a convolution module are designed, and a pooling module extracts adjacent semantics, differentiated semantics under different relations and global compensation semantic features respectively; then, semantic information flow and complementation are promoted through semantic enhancement module cascading, and deep fusion of multi-granularity semantic features is achieved; an embedded model architecture search method SAGA-KGE based on a double-layer proxy assisted genetic algorithm is provided; in order to enhance the population diversity in the search process, an elite environment selection method based on Levenshtein Distance is provided, and the search range of the genetic algorithm is expanded under the condition that elite individuals are reserved; in order to reduce computing resource overhead in the search process, a'classification + regression 'double-layer agent model is designed to replace an expensive fitness evaluation function, and consumption of computing resources is reduced.
Owner:NORTHWEST A & F UNIV

Data processing method and device

The invention discloses a data processing method and device. The method comprises the steps of obtaining a first character string in a first text and a second character string in a second text; determining target characters respectively appearing in the first character string and the second character string; according to the target character, the first character string and the second character string are segmented to obtain at least two character string groups, each character string group comprises a first substring and a second substring, the first substring is derived from the first character string, and the second substring is derived from the second character string; aiming at each character string group, obtaining a Levinstein distance between the first substring and the second substring in the character string group; and according to the Levinstein distances corresponding to all the character string groups, the Levinstein distance between the first character string and the second character string is obtained, and the Levinstein distance is used for obtaining a comparison result between the first text and the second text.
Owner:SMARTER SILICON (SHANGHAI) TECH CO LTD

Knowledge graph retrieval and classification method based on improved semi-supervised classification model

The invention relates to the technical field of information retrieval, in particular to a knowledge graph retrieval and classification method based on an improved semi-supervised classification model. The method comprises the following steps: on the basis of an improved semi-supervised classification model, executing preprocessing classification operation on graph nodes in a knowledge graph, and determining node features of the graph nodes; when a query keyword is received, calculating a similarity score between the query keyword and a node feature based on a semantic matching algorithm combined with a Jaccard coefficient and a Levenshtein distance similarity; determining whether a map entity is matched or not according to the similarity score and a preset similarity threshold value; if not, determining the map node with the highest matching degree according to the similarity score; and executing an expansion query action according to the map node with the highest matching degree to obtain the map node related to the query keyword and the attribute information of the map node. The purposes of improving the classification precision of the atlas nodes, improving the matching precision between the query keywords and the atlas nodes and improving the coverage and precision of retrieval are achieved.
Owner:YUNNAN DAILY NEWSPAPER GRP

Multi-text field voice interaction method based on OCR (Optical Character Recognition) and dynamic mask labeling

The invention discloses a multi-text field voice interaction method based on OCR and dynamic mask labeling, and relates to the technical field of intelligent interface interaction.The method comprises the steps that screen content is captured in real time for OCR recognition, and an OCR recognition result is obtained; performing duplicate removal and clustering on an OCR recognition result to generate a text field group; creating a transparent mask layer to cover a current application interface, distributing digital tags to the text field groups in sequence, and sorting the text field groups according to a tag sorting rule; recognizing a voice input instruction of the user, correcting an error of the voice input instruction by using a Levenshtein distance algorithm, analyzing the instruction, judging a text field which the user intends to select, and determining a final matching result of the text field; and recording user behaviors and optimizing a label sorting rule. According to the method, the adaptability to a dynamic interface and changeable information is improved, the efficiency and accuracy of the system in real-time interaction are enhanced, and the user experience and the interaction efficiency are improved.
Owner:RIVOTEK TECH (JIANGSU) CO LTD