Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

50 results about "Edit distance" patented technology

In computational linguistics and computer science, edit distance is a way of quantifying how dissimilar two strings (e.g., words) are to one another by counting the minimum number of operations required to transform one string into the other. Edit distances find applications in natural language processing, where automatic spelling correction can determine candidate corrections for a misspelled word by selecting words from a dictionary that have a low distance to the word in question. In bioinformatics, it can be used to quantify the similarity of DNA sequences, which can be viewed as strings of the letters A, C, G and T.

Method for analyzing rationality of acquisition and bid evaluation information of power equipment

The invention discloses a rationality analysis method for acquisition and bidding evaluation information of power equipment, which comprises the following steps: constructing a regular expression to perform paragraph matching on preprocessed bidding document and bidding document texts, and extracting specific parameter contents; obtaining an absolute error and a relative error based on each bidding value and each bidding value, and comparing the absolute error and the relative error with a preset threshold to obtain a numerical deviation degree grade of the bidding file relative to the bidding file; the method comprises the following steps: splitting a bid invitation file and a bidding file into character sequences, creating a two-dimensional table, analyzing to obtain an editing distance value of each text, obtaining a comprehensive matching degree of the text through accurate matching and fuzzy matching based on a preset synonym library, and further obtaining a text deviation degree grade; constructing a deviation degree comprehensive evaluation rule matrix based on the numerical deviation degree grade and the text deviation degree grade of the bidding file relative to the bidding file, and obtaining a corresponding comprehensive evaluation result; according to the invention, the efficiency and accuracy of collection and bidding evaluation of the power equipment are improved, and personal errors and compliance risks are reduced.
Owner:GUIZHOU POWER GRID CO LTD

Natural language query analysis and database field matching method based on large model

The embodiment of the invention provides a large-model-based natural language query analysis and database field matching method, which comprises the following steps of: performing intention analysis and entity recognition on natural language query of a user based on a large language model (LLM), and extracting a potential keyword set; for each extracted keyword, binding an entity type through an STAM mechanism, and constructing an entity-type mapping relationship; a column description vector library is constructed, efficient retrieval is achieved in combination with an approximate nearest neighbor ANN algorithm, and results of semantic vector matching, editing distance matching and traceability enhancement matching are combined through a multi-strategy fusion matching mechanism to form a final candidate column set; a database metadata interface is connected, information is analyzed, a basic structure is constructed, semantic enhancement field description is generated, and output is organized in a tetrad structure; and reversely deducing the affiliated table through column matching, judging the table structure value by combining the information density in the table and the connectivity with other tables, eliminating redundant tables, and optimizing database representation.
Owner:数字郑州科技有限公司

Equipment fault diagnosis method based on dynamic knowledge graph and large model fine tuning technology

The invention discloses an equipment fault diagnosis method and device based on a dynamic knowledge graph and a large model fine tuning technology. The method comprises the following steps: firstly, identifying a core entity from multi-source heterogeneous equipment fault data through a named entity identification model for fine tuning of domain data and a relation extraction model for special fine tuning of a fault diagnosis domain corpus, mining deep semantic association, and injecting the deep semantic association into a graph database after cleaning to form an initial knowledge graph; receiving user natural language fault description, realizing term and standard entity linking through editing distance fuzzy matching and Sension-BERT semantic vector similarity calculation, and combining bidirectional retrieval and attention mechanism fusion to obtain an enhanced context; and finally, generating a structured diagnosis report containing thinking chain reasoning based on an enhanced context by utilizing a specialized fine-tuning fault diagnosis large language model. According to the method, the limitation of a traditional diagnosis method is effectively solved, high-precision and interpretable equipment fault diagnosis is realized, and the diagnosis efficiency and reliability are improved.
Owner:AIR FORCE UNIV PLA

Text-to-instruction system based on fuzzy pinyin matching and parameter robust analysis

The invention discloses a text-to-instruction system based on fuzzy pinyin matching and parameter robust analysis. The text-to-instruction system comprises an input preprocessing module, a text-to-instruction conversion module and a text-to-instruction conversion module, the fuzzy pinyin matching module is used for uniformly transferring the standardized text and a pre-stored instruction template into a silent pinyin sequence, and constructing a weighted editing distance matrix based on a preset pinyin character replacement cost; the parameter alignment search module is used for executing beam search on the weighted editing distance matrix and outputting a plurality of candidate matching paths with the minimum cost and a corresponding parameter fragment set; the parameter robust distribution module is used for performing legality and consistency checking on parameters in the candidate matching paths, completing parameter standardization and performing error correction in combination with a parameter candidate dictionary; the instruction output module is used for encoding the template identifier and the parameter key value pair obtained through robust distribution into a structured instruction and outputting the structured instruction; according to the method, text-to-instruction conversion can be completed in real time.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Similarity calculation method based on semantic editing fusion

The invention discloses a similarity calculation method based on semantic editing fusion. The method comprises the following steps: calculating global semantic similarity of a source text and a target text by utilizing a semantic coding model; obtaining a first keyword list of the source text and a second keyword list of the target text through word segmentation processing, taking each first keyword in the first keyword list as a target word, and respectively forming a plurality of to-be-compared word pairs with a corresponding word in the second keyword list and a plurality of adjacent second keywords; based on the to-be-compared word pairs, calculating the semantic similarity between each group of target words and the corresponding words by utilizing a semantic coding model, and determining the editing distance between the source text and the target text; and determining the local semantic similarity of the source text and the target text according to the editing distance, and combining the global semantic similarity to obtain the final text matching similarity. According to the method, the problem that text matching only depends on character-level surface matching and neglects semantic association between word pairs is solved, and the text matching precision is improved.
Owner:XIDIAN UNIV +1

Dynamic hybrid resolution of bond ambiguity instructions with deterministic verification system and method

This application relates to the field of fintech and discloses a dynamic hybrid parsing and deterministic verification system and method for fuzzy bond instructions. The system constructs structured data through an instruction preprocessing module; a routing decision module calculates complexity scores based on structure, semantics, and business characteristics, and generates path selection instructions using a resource-aware dual-threshold strategy; a hybrid parsing module responds to instructions, employing a rule engine to handle low-complexity tasks and a large-scale model with enhanced retrieval to handle highly ambiguous tasks; a fallback verification module performs confidence-weighted integration and logical consistency checks, and repairs anomalies through issuer set operations and edit distance matching; a mapping generation module constructs standard query statements using an abstract syntax tree; and a closed-loop optimization module dynamically updates parameters based on user interaction feedback. This invention effectively balances computational efficiency and parsing accuracy, eliminates model illusion risks, and achieves adaptive evolution of the system.
Owner:CFETS FINANCIAL DATA CO LTD

Lyrics alignment method based on automatic speech recognition and electronic device

This invention relates to the fields of audio signal processing and artificial intelligence, specifically to a lyrics alignment method and electronic device based on automatic speech recognition. The lyrics alignment method based on automatic speech recognition in this application uses a pre-trained automatic speech recognition model to recognize audio samples and generate recognized text. The edit distance sequence between the recognized text and the lyrics text is calculated, and spoken segments are determined by analyzing the distance change pattern through a sliding window. These spoken segments are then replaced with silent segments, effectively eliminating spoken segments and avoiding interference from non-lyric content in the alignment process, thus laying the foundation for subsequent alignment.
Owner:TIANJIN UNIV

Hallucination prevention for natural language insights

Methods and systems are provided for hallucination prevention for natural language insights. In embodiments described herein, a template-based insight with a set of facts is generated by a template-based insights engine. The set of facts are generated from a set of data and the template-based insight is generated based on a text template. A natural language insight is generated from the template-based insight using a language model. If a single fact of the template-based insight is missing from the natural language insight, the single missing fact is an integer and a remaining integer of the natural language insight is within a threshold edit distance of the integer, the hallucination of the natural language insight is corrected by replacing the remaining integer with the single missing fact.
Owner:ADOBE INC

Method, system and terminal for realizing dynamic report generation through visual configuration

The invention discloses a method, a system and a terminal for generating a dynamic report through visual configuration, and the method comprises the following steps: metadata configuration: defining a report name and a code through a visual interface, and establishing a mapping relation between a query condition attribute name and a display name; engine binding is executed, wherein an SQL mode and a storage process mode are included; the SQL mode comprises the step of automatically replacing an SQL script containing an attribute name placeholder with an actual parameter value; the storage process mode comprises the steps of performing standardization processing on an interface attribute name and a storage process parameter name, and matching parameters by adopting an editing distance algorithm allowing 1-2 character differences; and multi-table rendering: independently rendering the master table and the slave table according to the bound query result serial number, and configuring a column display rule and an exclusive script. According to the method, the development efficiency and the parameter matching accuracy are improved, and complex structures such as listing are supported; configuration takes effect in real time, and the operation and maintenance cost of restarting service in a traditional scheme can be avoided.
Owner:SHENZHEN AISHIDA INFORMATION TECHNOLOGY CO LTD

A low complexity forward-backward decoding method based on weighted edit distance

The application discloses a low-complexity forward-backward decoding method based on weighted edit distance, and the method comprises the following steps: b A code word sequence with a length of N L symbols is generated by an encoder of an LDPC code d A mark code w is uniformly inserted into the code word sequence d to generate a sending code word with a length of N c and output the sending code word x After the sending code word x passes through an insertion / deletion-substitution channel, a receiving sequence with a length of N y is generated; a watermark decoder decodes the receiving sequence y by using a low-complexity forward-backward decoding method and outputs a likelihood ratio sequence l ; and an LDPC decoder decodes the likelihood ratio sequence l and outputs the application stores the calculation result of an intermediate metric value in a lookup table, reduces the number of repeated calculations of the intermediate metric value, reduces the calculation complexity of the decoding algorithm and improves the decoding speed.
Owner:TIANJIN NORMAL UNIVERSITY

Code error feedback method combining static analysis and dynamic analysis

PendingCN122072607Atimely feedbackEasy to FeedbackFault responseDifference listRename
The invention relates to a code error feedback method combining static analysis and dynamic analysis, and mainly solves the problems of insufficient error feedback and low efficiency in an Online Judge system: reconstructing codes of correct and error versions, including alignment and renaming of variables, and rearranging statements on the premise of not influencing semantics; static and dynamic multi-level alignment is carried out, in the static alignment, an editing script with the minimum cost is calculated through a heuristic editing distance algorithm, an alignment block and a difference list are obtained, in the dynamic alignment, compiling execution is carried out on source codes of two versions, test cases which do not pass operation are dynamically printed through gdb, and execution track information of the test cases is dynamically printed through gdb; error positioning is achieved, data flow backtracking analysis and scoring are carried out on error positions, and error feedback is generated. The method has the characteristics of high execution efficiency, high accuracy and easier understanding of generated feedback, automatic feedback of error codes can be realized, and the learning efficiency of students in programming tasks is improved.
Owner:NANJING UNIV

Pinyin error correction method

This invention provides a Pinyin error correction method that simplifies the noise channel model used in most current Pinyin error correction algorithms by employing a real-time frequency counting method, effectively improving the efficiency and lightweighting of the algorithm. Furthermore, this invention uses a direct character substitution method instead of the traditional edit distance calculation method, avoiding frequent edit distance calculations in the Pinyin error correction algorithm. The dictionary database itself is localized to the individual user, resulting in a highly personalized, targeted, and small-scale dictionary that provides accurate candidate words. Compared to existing algorithms, this invention offers significant improvements in error checking rate, candidate word accuracy, and execution efficiency.
Owner:ZHONGKE FANYU (WUHAN) TECH CO LTD

Program error positioning and interpretation generation method and system based on reinforcement learning

The invention provides a program error positioning and explanation generation method and system based on reinforcement learning, and belongs to the field of AI auxiliary programming education and natural language processing, and the method comprises the steps: S1, based on student submission records of a real teaching scene, screening out error codes and correct code pairing samples through an editing distance, and constructing a training set; s2, sampling and filtering the training set by using a pre-training language model, and constructing a sampling rejection data set to supervise and finely adjust the pre-training language model to obtain a base model; s3, constructing a multi-signal fusion reinforcement learning reward function for performing reinforcement learning training on the base model to obtain a trained error positioning and explanation generation model; and S4, inputting the question description and the error code into the trained error positioning and interpretation generation model, and generating structured feedback. According to the method, the error positioning accuracy is remarkably improved, the false drop rate is effectively reduced, and a feasible technical path is provided for intelligent programming education.
Owner:BEIHANG UNIV

Natural language query analysis and database field matching method based on large model

The embodiment of the specification provides a natural language query analysis and database field matching method based on a large model, which comprises the following steps: performing intent analysis and entity recognition on a user natural language query based on a large language model (LLM) to extract a potential keyword set; for each extracted keyword, binding an entity type through an STAM mechanism and constructing an entity-type mapping relationship; constructing a column description vector library, combining an approximate nearest neighbor (ANN) algorithm to realize efficient retrieval, and combining a multi-strategy fusion matching mechanism to combine the results of semantic vector matching, edit distance matching and traceability enhanced matching to form a final candidate column set; connecting a database metadata interface, analyzing information and constructing a basic structure to generate an enhanced semantic field description, and outputting in a four-tuple structure; reversing the corresponding table through column matching, judging the table structure value in combination with the information density in the table and the connectivity with other tables, eliminating redundant tables, and optimizing the database representation.
Owner:数字郑州科技有限公司

String similarity determination method, apparatus, program product, and related device

This disclosure provides a method, apparatus, program product, and related equipment for determining string similarity, relating to the field of artificial intelligence technology. The method includes: acquiring a first string and a second string; selecting a word from the first word in the first string that has not been compared before as the current word to be compared; performing semantic comparisons between the current word to be compared and the second word in the second string; if the second word has a semantic opposite to the current word to be compared, then determining the similarity between the first string and the second string to be zero; otherwise, continuing to select a first word that has not been compared before as the current word to be compared and continuing semantic comparisons; if, after traversing the first word, it is determined that there is no word in the second string with a semantic opposite to the first word, then using the edit distance similarity between the first string and the second string as the similarity between the first string and the second string. This method can improve the accuracy of string similarity determination.
Owner:TENCENT CLOUD COMPUTING (CHANGSHA) CO LTD

Automatic process control system and method for low-code platform

The invention discloses an automatic process control system and method for a low-code platform, and relates to the technical field of computers.The method comprises the steps that source data of a user on the low-code platform are obtained according to the preset frequency, and snapshots are generated; the source data and predefined cue words are spliced to serve as large language model input; the large language model outputs structured suggestions, and the suggestions are classified and displayed for users to confirm; after the user confirms, recording related data as a training sample and adjusting the suggestion display frequency; constructing an event chain library, calculating a matching degree between an actual operation event chain and an operation event chain in the event chain library according to an editing distance algorithm, and recommending a component type; according to the method, the local network speed and computing power data of the user are obtained, the learning rate of the large language model is weighted and adjusted after normalization, training is completed in a local background, intelligentization and high efficiency of automatic process control of the low-code platform are achieved, and user operation experience and platform applicability are improved.
Owner:WISCOM SYSTEM CO LTD

Robustness verification method and system based on source code pre-training model and storage medium

The invention discloses a robustness verification method and system based on a source code pre-training model and a storage medium. The method comprises the steps of obtaining a model and extracting code data features to obtain a token sequence; performing slicing processing on the token sequence; sampling token sequence slices and recombining to obtain a random token sequence; designing a classification loss function and finely adjusting the large pre-training model; taking the random token sequence as the input of a large pre-training model to obtain a classification label predicted by the model; calculating self-confident degree of classification according to classification labels obtained by multiple times of model prediction, and obtaining output prediction of the large pre-training model; constructing a new token sequence based on the editing distance; and calculating the robust radius according to the boundary of the confidence value of the predicted label and a dichotomy. According to the method, the problem that the robustness is difficult to guarantee due to the fact that an existing random smoothing method is difficult to perform unified modeling due to complex disturbance in a source code scene in an existing method and a large pre-training language model cannot perform fine processing on various features is solved.
Owner:HANGZHOU DBAPPSECURITY CO LTD

Question and answer pair automatic generation method and device, computer equipment and storage medium

The invention relates to the technical field of natural language processing and information retrieval, and discloses a question and answer pair automatic generation method and device, computer equipment and a storage medium. The method comprises the following steps: firstly, extracting an initial question and answer pair from an input text through a large language model (LLM); slicing and vectorizing the original text, and constructing an original text vector database; vectorizing answers in the initial question and answer pair, matching the vectorized answers with an original text vector database, and positioning one or more original text fragments for each answer to form a traceable triple; performing quality verification on the triple by adopting a comprehensive verification mechanism based on integrity and accuracy evaluation and key information coverage calculation; and finally, de-duplication is carried out in combination with editing distance calculation and semantic similarity calculation, and high-quality question and answer pairs are output. According to the method, the problems of unstable generation quality, lack of answer basis and the like are solved, and automatic construction of the high-credibility knowledge base is realized.
Owner:山东齐鲁壹点传媒有限公司 +1

Log analysis method and computing device

The application discloses a log analysis method and a computing device. The method comprises the following steps: obtaining log content of a target line of a test log; determining whether the number of lines of the test log is greater than a first threshold value when the log content does not contain a target keyword; when the number of lines of the test log is greater than the first threshold value, calculating the jasccard similarity between the log content and each keyword; when the similarity is greater than or equal to a second threshold value, determining that the keyword corresponding to the similarity is the target keyword; when the number of lines of the test log is less than the first threshold value, calculating the edit distance between the log content and the target keyword; when the edit distance is less than or equal to a third threshold value, determining that the keyword corresponding to the edit distance is the target keyword; and then obtaining a log analysis result according to the target keyword. In this way, the analysis of the log content can be more efficient and accurate.
Owner:YUXIN ELECTRONIC TECHNOLOGY GROUP CO LTD

Using fuzzy matching to determine whether segment(s) of responsive content, that is generated using generative model(s), match segment(s) of additional data

Some implementations described herein relate to determining whether to modify segment(s) of responsive content, that is generated using a generative model (GM), based on a corresponding edit distance between the segment(s) of the responsive content and segment(s) of additional data. Processor(s) of a system can: receive user input associated with a client device, generate the responsive content using the GM and based on processing the user input, and determine whether to modify the segment(s) of the responsive content using the corresponding edit distance (e.g., fuzzy matching). Subsequent modification and / or processing of the responsive content can be dynamically adapted based on whether there is a match and, if there is a match, based on source(s) associated with the additional data that match. Further, the processor(s) can cause the responsive content, or modified responsive content, to be rendered at the client device.
Owner:GOOGLE LLC

Text data processing method, device, equipment and medium

The embodiments of the present application provide a text data processing method, apparatus, device, and medium, which relates to the field of artificial intelligence, and includes: obtaining a first text pair and a second text pair, obtaining a first subtext from the first text pair, and obtaining a second subtext from the second text pair; determining the edit distance between the first subtext and the second subtext, and if the edit distance satisfies a similarity condition, generating a first target subtext associated with the semantic information of the first subtext and belonging to a third language type, and generating a second target subtext associated with the semantic information of the second subtext and belonging to a second language type; generating a text sample pair based on the first text pair, the second text pair, the first target subtext, and the second target subtext. By using the present application, text sample pairs composed of different language types can be generated, thereby increasing the quantity of the corpus while ensuring the quality of the corpus.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Semantic generation method and device, equipment, medium and product

The invention discloses a semantic generation method and device, equipment, a medium and a product. The method comprises the steps that a labeled sentence pattern set corresponding to a text corpus of a vehicle-mounted service is generated through a sequence labeling model; constructing tree structures of all sentence patterns in the labeled sentence pattern set to obtain all tree structures corresponding to the labeled sentence pattern set; and combining all the tree structures based on the editing distance to obtain corpus semantics of the text corpus, and if the corpus semantics do not meet a preset condition, correcting the corpus semantics to obtain corrected corpus semantics. The annotated sentence pattern set corresponding to the text corpus of the vehicle-mounted service is generated through the sequence annotation model, the annotation efficiency is improved, and the sentence patterns are clearly presented by constructing the tree structures of all the sentence patterns in the annotated sentence pattern set, so that all the tree structures are conveniently combined through the editing distance, the combined tree structures are more accurate, and the annotation efficiency is improved. And thus, the obtained corpus semantics are more accurate.
Owner:IFLYTEK CO LTD

Text alignment method and apparatus, electronic device, and storage medium

ActiveCN116306557BText alignmentEngineering
Embodiments of the present application provide a text alignment method and device, electronic equipment and storage medium, belonging to the field of artificial intelligence. The method comprises: obtaining a preset reference text, obtaining and calculating a sliding window ratio according to an original text segment to obtain an original sliding window ratio; calculating an initial edit distance according to the original sliding window ratio and the reference text; screening the original sliding window ratio according to the initial edit distance to obtain an initial sliding window ratio; contracting the original sliding window ratio according to the initial sliding window ratio to obtain a target sliding window ratio; calculating a target edit distance according to the target sliding window ratio and the reference text; screening the target sliding window ratio according to the target edit distance to obtain a current sliding window ratio; and aligning the original text segment and the reference text according to the current sliding window ratio to obtain a target text. The embodiments of the present application realize text alignment under the condition of a small amount of resources.
Owner:PING AN TECH (SHENZHEN) CO LTD

A text content comparison matching method, device, equipment and medium

PendingCN122366407AEngineeringThresholding
This invention relates to the field of artificial intelligence technology, and discloses a text content comparison and matching method, apparatus, device, and medium. The method includes sequentially performing character adaptation and normalization, phonological encoding conversion, segmentation processing, fuzzy weighted edit distance calculation, peak analysis and path backtracking, and jump threshold traversal judgment on template text and speech recognition text to obtain text comparison and matching result data. In this invention, addressing the problem that existing text content comparison and matching schemes struggle to accurately identify the continuous inclusion of speech recognition text in template text, character adaptation and normalization and phonological encoding conversion can eliminate character form differences; segmentation processing and fuzzy weighted edit distance calculation can suppress accumulated errors; and peak analysis, path backtracking, and jump threshold traversal judgment are then applied. This method can be applied to the financial and medical fields, thus accurately identifying the continuous inclusion of speech recognition text in template text at the character granularity.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-stage spelling error correction method and device for medical image report and medium

The invention relates to a multi-stage spelling error correction method and device for a medical image report and a medium, in a non-term error correction stage, for each text segment, a pinyin error correction knowledge base and a polyphone probability base are called, screening and replacement are carried out according to text similarity and editing distance, candidate correction for general spelling errors is generated, and the spelling error correction efficiency is improved. Obtaining a text fragment after non-term error correction; in the term error correction stage, a medical term package is called, candidate correction for term spelling errors is generated and screened through candidate interval positioning, set similarity screening and sequence similarity calculation, and text segments after term error correction are obtained; and enabling the text segments after non-term error correction and the text segments after term error correction to correspond to each other at the text position, correcting the candidates with overlapping or conflicting, and performing resolution and combination according to a preset priority rule to obtain a final error correction result. Compared with the prior art, the method has the advantages of wide range, high accuracy, high reliability and the like.
Owner:SHANGHAI EBM MEDICAL INFORMATION SYST

Method for extending sound near-sensitive words

The application provides an extension method of homophonic sensitive words, comprising: combining two by two of pinyins in a legal pinyin table; obtaining an edit distance of each two-by-two combination result, and extracting a homophonic pinyin group according to the edit distance to construct a pinyin-homophonic pinyin table; replacing any character pinyin in a sensitive word in a sensitive word database based on the pinyin-homophonic pinyin table, and mapping the replaced any character pinyin into a character based on a pinyin-Hanzi table to construct a candidate homophonic word; and pre-judging the candidate homophonic word to realize the supplementary extension of the sensitive word database. By using the existing sensitive word library and homophonic word table, the homophonic character variants of the sensitive words that the black production may use are speculated, the characteristics such as the possibility of missing and the long time consumption of the whole link are solved in advance, and the effectiveness of the extracted keywords is improved.
Owner:SHENZHEN BAICHUAN SHUAN TECH CO LTD

Cross-language editing distance and multi-model fusion mixed word detection method and system

The invention discloses a cross-language editing distance and multi-model fused mixed word detection method and system, and belongs to the technical field of natural language processing. According to the method, firstly, preprocessing and candidate mixed word extraction are carried out on a text, a cross-language editing distance is calculated by utilizing dynamic programming, a large language model is called twice to carry out semantic scoring by adding a text window fragment without a context, then calibration weighted aggregation is carried out, and a final result is output through multi-model fusion and optimization. And high-precision recognition of Chinese-English mixed words is realized. By means of the scheme, effective combination of the structural features and the semantic features is achieved, the Chinese and English mixed words can be efficiently and accurately detected, interpretability and real-time performance are achieved, the accuracy and robustness of mixed word detection are improved, and the method can be widely applied to input methods, social platforms, network public opinion monitoring, news transmission analysis and other scenes.
Owner:ZHEJIANG UNIV

Translation quality classification method and apparatus

This application provides a method and apparatus for translation quality classification. The method includes: acquiring source text and machine translation; inputting the source text and machine translation into a translation quality classification model, and acquiring the classification result output by the translation quality classification model; the translation quality classification model is optimized based on a first edit distance score, which is obtained based on the machine translation and the proofread translation, wherein the proofread translation is the machine translation after proofreading. The translation quality classification method and apparatus provided in this application can optimize a pre-trained language model using the first edit distance score between the machine translation and the proofread translation to obtain a translation quality classification model. This allows the translation quality classification model to output a classification result based on the input source text and machine translation, thereby improving the accuracy of translation quality classification.
Owner:TRANSN IOL TECH CO LTD

Book name error correction method and system based on semantic embedding and editing distance fusion

The invention discloses a book name error correction method and system based on semantic embedding and editing distance fusion, and the method comprises the steps: carrying out the recall of two-stage candidate words from a candidate word bank for a to-be-corrected word, obtaining a target candidate word set, and carrying out the scoring and sorting of the candidate words in the target candidate word set; and based on the sorting result, selecting the target candidate word as an error correction result to be output. According to the technical scheme provided by the invention, firstly, the candidate words with similar fonts are ensured not to be omitted through the primary recall based on the editing distance, and then the similar candidate words are expanded through the secondary recall based on the semantic vector similarity, so that the coverage range and diversity of the candidate word set are remarkably improved; then multi-dimensional reordering is carried out on the candidate words through a comprehensive scoring model, decision making is carried out by fully combining font features, semantic relevance and context importance, and therefore the accuracy of error correction is greatly improved while the high recall rate is kept.
Owner:TIANJIN KAREL ROBOT TECH CO LTD