Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

63 results about "Edit distance" patented technology

In computational linguistics and computer science, edit distance is a way of quantifying how dissimilar two strings (e.g., words) are to one another by counting the minimum number of operations required to transform one string into the other. Edit distances find applications in natural language processing, where automatic spelling correction can determine candidate corrections for a misspelled word by selecting words from a dictionary that have a low distance to the word in question. In bioinformatics, it can be used to quantify the similarity of DNA sequences, which can be viewed as strings of the letters A, C, G and T.

Structured analysis and semantic template normalization method for unmanned vehicle operation logs

The invention provides a structured analysis and semantic template normalization method for an unmanned vehicle operation log. For the problems of unstructured unmanned vehicle log data, changeable formats, complex semantics and the like, logs are converted into semantic vectors by adopting a multi-language sentence vector coding model, and semantic grouping is performed by using a small-batch K-means clustering algorithm after dimension reduction is performed by adopting an incremental principal component analysis method; extracting a variable field from a clustering result by using a regular expression and generating a standardized template; and intelligent template combination is realized by calculating cosine similarity and editing distance between the templates. The method has the following three technical characteristics: 1) the template consistency is improved by adopting a strategy of clustering first and then normalizing; 2) introducing double similarity constraints to reduce template redundancy; and 3) retaining variable sequence and semantic information to support subsequent analysis. The method is suitable for log data analysis and intelligent operation and maintenance of various complex systems.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Method for analyzing rationality of acquisition and bid evaluation information of power equipment

The invention discloses a rationality analysis method for acquisition and bidding evaluation information of power equipment, which comprises the following steps: constructing a regular expression to perform paragraph matching on preprocessed bidding document and bidding document texts, and extracting specific parameter contents; obtaining an absolute error and a relative error based on each bidding value and each bidding value, and comparing the absolute error and the relative error with a preset threshold to obtain a numerical deviation degree grade of the bidding file relative to the bidding file; the method comprises the following steps: splitting a bid invitation file and a bidding file into character sequences, creating a two-dimensional table, analyzing to obtain an editing distance value of each text, obtaining a comprehensive matching degree of the text through accurate matching and fuzzy matching based on a preset synonym library, and further obtaining a text deviation degree grade; constructing a deviation degree comprehensive evaluation rule matrix based on the numerical deviation degree grade and the text deviation degree grade of the bidding file relative to the bidding file, and obtaining a corresponding comprehensive evaluation result; according to the invention, the efficiency and accuracy of collection and bidding evaluation of the power equipment are improved, and personal errors and compliance risks are reduced.
Owner:GUIZHOU POWER GRID CO LTD

Natural language query analysis and database field matching method based on large model

The embodiment of the invention provides a large-model-based natural language query analysis and database field matching method, which comprises the following steps of: performing intention analysis and entity recognition on natural language query of a user based on a large language model (LLM), and extracting a potential keyword set; for each extracted keyword, binding an entity type through an STAM mechanism, and constructing an entity-type mapping relationship; a column description vector library is constructed, efficient retrieval is achieved in combination with an approximate nearest neighbor ANN algorithm, and results of semantic vector matching, editing distance matching and traceability enhancement matching are combined through a multi-strategy fusion matching mechanism to form a final candidate column set; a database metadata interface is connected, information is analyzed, a basic structure is constructed, semantic enhancement field description is generated, and output is organized in a tetrad structure; and reversely deducing the affiliated table through column matching, judging the table structure value by combining the information density in the table and the connectivity with other tables, eliminating redundant tables, and optimizing database representation.
Owner:数字郑州科技有限公司

Equipment fault diagnosis method based on dynamic knowledge graph and large model fine tuning technology

The invention discloses an equipment fault diagnosis method and device based on a dynamic knowledge graph and a large model fine tuning technology. The method comprises the following steps: firstly, identifying a core entity from multi-source heterogeneous equipment fault data through a named entity identification model for fine tuning of domain data and a relation extraction model for special fine tuning of a fault diagnosis domain corpus, mining deep semantic association, and injecting the deep semantic association into a graph database after cleaning to form an initial knowledge graph; receiving user natural language fault description, realizing term and standard entity linking through editing distance fuzzy matching and Sension-BERT semantic vector similarity calculation, and combining bidirectional retrieval and attention mechanism fusion to obtain an enhanced context; and finally, generating a structured diagnosis report containing thinking chain reasoning based on an enhanced context by utilizing a specialized fine-tuning fault diagnosis large language model. According to the method, the limitation of a traditional diagnosis method is effectively solved, high-precision and interpretable equipment fault diagnosis is realized, and the diagnosis efficiency and reliability are improved.
Owner:AIR FORCE UNIV PLA

Text-to-instruction system based on fuzzy pinyin matching and parameter robust analysis

The invention discloses a text-to-instruction system based on fuzzy pinyin matching and parameter robust analysis. The text-to-instruction system comprises an input preprocessing module, a text-to-instruction conversion module and a text-to-instruction conversion module, the fuzzy pinyin matching module is used for uniformly transferring the standardized text and a pre-stored instruction template into a silent pinyin sequence, and constructing a weighted editing distance matrix based on a preset pinyin character replacement cost; the parameter alignment search module is used for executing beam search on the weighted editing distance matrix and outputting a plurality of candidate matching paths with the minimum cost and a corresponding parameter fragment set; the parameter robust distribution module is used for performing legality and consistency checking on parameters in the candidate matching paths, completing parameter standardization and performing error correction in combination with a parameter candidate dictionary; the instruction output module is used for encoding the template identifier and the parameter key value pair obtained through robust distribution into a structured instruction and outputting the structured instruction; according to the method, text-to-instruction conversion can be completed in real time.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Similarity calculation method based on semantic editing fusion

The invention discloses a similarity calculation method based on semantic editing fusion. The method comprises the following steps: calculating global semantic similarity of a source text and a target text by utilizing a semantic coding model; obtaining a first keyword list of the source text and a second keyword list of the target text through word segmentation processing, taking each first keyword in the first keyword list as a target word, and respectively forming a plurality of to-be-compared word pairs with a corresponding word in the second keyword list and a plurality of adjacent second keywords; based on the to-be-compared word pairs, calculating the semantic similarity between each group of target words and the corresponding words by utilizing a semantic coding model, and determining the editing distance between the source text and the target text; and determining the local semantic similarity of the source text and the target text according to the editing distance, and combining the global semantic similarity to obtain the final text matching similarity. According to the method, the problem that text matching only depends on character-level surface matching and neglects semantic association between word pairs is solved, and the text matching precision is improved.
Owner:XIDIAN UNIV +1

PCB schematic diagram component information extraction method based on text recognition and graph algorithm

The invention relates to the technical field of electronic design automation, and provides a PCB schematic diagram component information extraction method based on text recognition and a graph algorithm. The method aims at solving the problem that component models in a complex PCB schematic diagram are extracted inaccurately, and the method mainly comprises the steps that preprocessing of binaryzation, denoising and text area enhancement is conducted on the PCB schematic diagram, and an optimized image is generated; performing character direction detection and rotation processing on the obtained optimized image, inputting the fine-tuned OCR model to perform multi-scale character recognition, and outputting a recognition result containing character content, coordinates and confidence; constructing a graph model containing MCU position nodes and character block nodes, and establishing associated edges through a weight calculation module; performing global optimal matching on the graph model based on a Hungary algorithm, and screening associated pairs of the MCU and the character blocks; and performing editing distance calculation on a matching result and a preset MCU model template library, correcting character confusion errors through a multi-stage verification mechanism, and outputting final component information.
Owner:NINGBO QINGYUN CHUANGXIN TECHNOLOGY CO LTD

Dynamic hybrid resolution of bond ambiguity instructions with deterministic verification system and method

This application relates to the field of fintech and discloses a dynamic hybrid parsing and deterministic verification system and method for fuzzy bond instructions. The system constructs structured data through an instruction preprocessing module; a routing decision module calculates complexity scores based on structure, semantics, and business characteristics, and generates path selection instructions using a resource-aware dual-threshold strategy; a hybrid parsing module responds to instructions, employing a rule engine to handle low-complexity tasks and a large-scale model with enhanced retrieval to handle highly ambiguous tasks; a fallback verification module performs confidence-weighted integration and logical consistency checks, and repairs anomalies through issuer set operations and edit distance matching; a mapping generation module constructs standard query statements using an abstract syntax tree; and a closed-loop optimization module dynamically updates parameters based on user interaction feedback. This invention effectively balances computational efficiency and parsing accuracy, eliminates model illusion risks, and achieves adaptive evolution of the system.
Owner:CFETS FINANCIAL DATA CO LTD

Lyrics alignment method based on automatic speech recognition and electronic device

This invention relates to the fields of audio signal processing and artificial intelligence, specifically to a lyrics alignment method and electronic device based on automatic speech recognition. The lyrics alignment method based on automatic speech recognition in this application uses a pre-trained automatic speech recognition model to recognize audio samples and generate recognized text. The edit distance sequence between the recognized text and the lyrics text is calculated, and spoken segments are determined by analyzing the distance change pattern through a sliding window. These spoken segments are then replaced with silent segments, effectively eliminating spoken segments and avoiding interference from non-lyric content in the alignment process, thus laying the foundation for subsequent alignment.
Owner:TIANJIN UNIV

Hallucination prevention for natural language insights

Methods and systems are provided for hallucination prevention for natural language insights. In embodiments described herein, a template-based insight with a set of facts is generated by a template-based insights engine. The set of facts are generated from a set of data and the template-based insight is generated based on a text template. A natural language insight is generated from the template-based insight using a language model. If a single fact of the template-based insight is missing from the natural language insight, the single missing fact is an integer and a remaining integer of the natural language insight is within a threshold edit distance of the integer, the hallucination of the natural language insight is corrected by replacing the remaining integer with the single missing fact.
Owner:ADOBE INC

Method, system and terminal for realizing dynamic report generation through visual configuration

The invention discloses a method, a system and a terminal for generating a dynamic report through visual configuration, and the method comprises the following steps: metadata configuration: defining a report name and a code through a visual interface, and establishing a mapping relation between a query condition attribute name and a display name; engine binding is executed, wherein an SQL mode and a storage process mode are included; the SQL mode comprises the step of automatically replacing an SQL script containing an attribute name placeholder with an actual parameter value; the storage process mode comprises the steps of performing standardization processing on an interface attribute name and a storage process parameter name, and matching parameters by adopting an editing distance algorithm allowing 1-2 character differences; and multi-table rendering: independently rendering the master table and the slave table according to the bound query result serial number, and configuring a column display rule and an exclusive script. According to the method, the development efficiency and the parameter matching accuracy are improved, and complex structures such as listing are supported; configuration takes effect in real time, and the operation and maintenance cost of restarting service in a traditional scheme can be avoided.
Owner:SHENZHEN AISHIDA INFORMATION TECHNOLOGY CO LTD

A low complexity forward-backward decoding method based on weighted edit distance

The application discloses a low-complexity forward-backward decoding method based on weighted edit distance, and the method comprises the following steps: b A code word sequence with a length of N L symbols is generated by an encoder of an LDPC code d A mark code w is uniformly inserted into the code word sequence d to generate a sending code word with a length of N c and output the sending code word x After the sending code word x passes through an insertion / deletion-substitution channel, a receiving sequence with a length of N y is generated; a watermark decoder decodes the receiving sequence y by using a low-complexity forward-backward decoding method and outputs a likelihood ratio sequence l ; and an LDPC decoder decodes the likelihood ratio sequence l and outputs the application stores the calculation result of an intermediate metric value in a lookup table, reduces the number of repeated calculations of the intermediate metric value, reduces the calculation complexity of the decoding algorithm and improves the decoding speed.
Owner:TIANJIN NORMAL UNIVERSITY

Code error feedback method combining static analysis and dynamic analysis

PendingCN122072607Atimely feedbackEasy to FeedbackFault responseDifference listRename
The invention relates to a code error feedback method combining static analysis and dynamic analysis, and mainly solves the problems of insufficient error feedback and low efficiency in an Online Judge system: reconstructing codes of correct and error versions, including alignment and renaming of variables, and rearranging statements on the premise of not influencing semantics; static and dynamic multi-level alignment is carried out, in the static alignment, an editing script with the minimum cost is calculated through a heuristic editing distance algorithm, an alignment block and a difference list are obtained, in the dynamic alignment, compiling execution is carried out on source codes of two versions, test cases which do not pass operation are dynamically printed through gdb, and execution track information of the test cases is dynamically printed through gdb; error positioning is achieved, data flow backtracking analysis and scoring are carried out on error positions, and error feedback is generated. The method has the characteristics of high execution efficiency, high accuracy and easier understanding of generated feedback, automatic feedback of error codes can be realized, and the learning efficiency of students in programming tasks is improved.
Owner:NANJING UNIV

Pinyin error correction method

This invention provides a Pinyin error correction method that simplifies the noise channel model used in most current Pinyin error correction algorithms by employing a real-time frequency counting method, effectively improving the efficiency and lightweighting of the algorithm. Furthermore, this invention uses a direct character substitution method instead of the traditional edit distance calculation method, avoiding frequent edit distance calculations in the Pinyin error correction algorithm. The dictionary database itself is localized to the individual user, resulting in a highly personalized, targeted, and small-scale dictionary that provides accurate candidate words. Compared to existing algorithms, this invention offers significant improvements in error checking rate, candidate word accuracy, and execution efficiency.
Owner:ZHONGKE FANYU (WUHAN) TECH CO LTD

Program error positioning and interpretation generation method and system based on reinforcement learning

The invention provides a program error positioning and explanation generation method and system based on reinforcement learning, and belongs to the field of AI auxiliary programming education and natural language processing, and the method comprises the steps: S1, based on student submission records of a real teaching scene, screening out error codes and correct code pairing samples through an editing distance, and constructing a training set; s2, sampling and filtering the training set by using a pre-training language model, and constructing a sampling rejection data set to supervise and finely adjust the pre-training language model to obtain a base model; s3, constructing a multi-signal fusion reinforcement learning reward function for performing reinforcement learning training on the base model to obtain a trained error positioning and explanation generation model; and S4, inputting the question description and the error code into the trained error positioning and interpretation generation model, and generating structured feedback. According to the method, the error positioning accuracy is remarkably improved, the false drop rate is effectively reduced, and a feasible technical path is provided for intelligent programming education.
Owner:BEIHANG UNIV

Natural language query analysis and database field matching method based on large model

The embodiment of the specification provides a natural language query analysis and database field matching method based on a large model, which comprises the following steps: performing intent analysis and entity recognition on a user natural language query based on a large language model (LLM) to extract a potential keyword set; for each extracted keyword, binding an entity type through an STAM mechanism and constructing an entity-type mapping relationship; constructing a column description vector library, combining an approximate nearest neighbor (ANN) algorithm to realize efficient retrieval, and combining a multi-strategy fusion matching mechanism to combine the results of semantic vector matching, edit distance matching and traceability enhanced matching to form a final candidate column set; connecting a database metadata interface, analyzing information and constructing a basic structure to generate an enhanced semantic field description, and outputting in a four-tuple structure; reversing the corresponding table through column matching, judging the table structure value in combination with the information density in the table and the connectivity with other tables, eliminating redundant tables, and optimizing the database representation.
Owner:数字郑州科技有限公司

String similarity determination method, apparatus, program product, and related device

This disclosure provides a method, apparatus, program product, and related equipment for determining string similarity, relating to the field of artificial intelligence technology. The method includes: acquiring a first string and a second string; selecting a word from the first word in the first string that has not been compared before as the current word to be compared; performing semantic comparisons between the current word to be compared and the second word in the second string; if the second word has a semantic opposite to the current word to be compared, then determining the similarity between the first string and the second string to be zero; otherwise, continuing to select a first word that has not been compared before as the current word to be compared and continuing semantic comparisons; if, after traversing the first word, it is determined that there is no word in the second string with a semantic opposite to the first word, then using the edit distance similarity between the first string and the second string as the similarity between the first string and the second string. This method can improve the accuracy of string similarity determination.
Owner:TENCENT CLOUD COMPUTING (CHANGSHA) CO LTD

A string fuzzy matching method and system

The present invention discloses a string fuzzy matching method and system, which relates to the field of character recognition. The method comprises: obtaining multiple characters in different font sizes and constructing a glyph similarity matrix; weighting the cost of the replacement operation by character similarity within the edit distance method to construct a weighted edit distance method; applying a weighted edit distance to the two strings to be matched, normalizing the obtained weighted edit distance, and obtaining the similarity of the two strings to be matched; when the similarity is 1, it indicates a successful match; when the similarity is greater than a set similarity threshold and less than 1, it indicates a suspicious match result; and when the similarity is less than the set similarity threshold, it indicates a failed match. The present invention can solve the problem of character errors in similarity results identified by OCR technology.
Owner:XIANGTAN UNIV

Automatic process control system and method for low-code platform

The invention discloses an automatic process control system and method for a low-code platform, and relates to the technical field of computers.The method comprises the steps that source data of a user on the low-code platform are obtained according to the preset frequency, and snapshots are generated; the source data and predefined cue words are spliced to serve as large language model input; the large language model outputs structured suggestions, and the suggestions are classified and displayed for users to confirm; after the user confirms, recording related data as a training sample and adjusting the suggestion display frequency; constructing an event chain library, calculating a matching degree between an actual operation event chain and an operation event chain in the event chain library according to an editing distance algorithm, and recommending a component type; according to the method, the local network speed and computing power data of the user are obtained, the learning rate of the large language model is weighted and adjusted after normalization, training is completed in a local background, intelligentization and high efficiency of automatic process control of the low-code platform are achieved, and user operation experience and platform applicability are improved.
Owner:WISCOM SYSTEM CO LTD

Robustness verification method and system based on source code pre-training model and storage medium

The invention discloses a robustness verification method and system based on a source code pre-training model and a storage medium. The method comprises the steps of obtaining a model and extracting code data features to obtain a token sequence; performing slicing processing on the token sequence; sampling token sequence slices and recombining to obtain a random token sequence; designing a classification loss function and finely adjusting the large pre-training model; taking the random token sequence as the input of a large pre-training model to obtain a classification label predicted by the model; calculating self-confident degree of classification according to classification labels obtained by multiple times of model prediction, and obtaining output prediction of the large pre-training model; constructing a new token sequence based on the editing distance; and calculating the robust radius according to the boundary of the confidence value of the predicted label and a dichotomy. According to the method, the problem that the robustness is difficult to guarantee due to the fact that an existing random smoothing method is difficult to perform unified modeling due to complex disturbance in a source code scene in an existing method and a large pre-training language model cannot perform fine processing on various features is solved.
Owner:HANGZHOU DBAPPSECURITY CO LTD

Product information query method and device, computer equipment and storage medium

The invention provides a product information query method and device, computer equipment and a computer readable storage medium, and belongs to the field of natural language processing. The method comprises the steps that statement generation is conducted based on a dialogue text of a target product system, a generated query statement of the dialogue text is obtained, the tree form editing distance between a syntax tree of the generated query statement and a syntax tree of a target standard query statement is larger than a preset editing distance, and the target standard query statement is determined based on a query intention of the dialogue text; obtaining redundant information for generating the query statement; optimizing the generated query statement based on the redundant information to obtain a target query statement of the dialogue text; and obtaining a product information query result of the dialogue text from a database of the target product system based on the target query statement. The method can be applied to the field of finance, individual differences of product information query results on query requirements of different users can be improved, and the accuracy of the product information query results is improved.
Owner:PING AN FINANCE CO LTD

Question and answer pair automatic generation method and device, computer equipment and storage medium

The invention relates to the technical field of natural language processing and information retrieval, and discloses a question and answer pair automatic generation method and device, computer equipment and a storage medium. The method comprises the following steps: firstly, extracting an initial question and answer pair from an input text through a large language model (LLM); slicing and vectorizing the original text, and constructing an original text vector database; vectorizing answers in the initial question and answer pair, matching the vectorized answers with an original text vector database, and positioning one or more original text fragments for each answer to form a traceable triple; performing quality verification on the triple by adopting a comprehensive verification mechanism based on integrity and accuracy evaluation and key information coverage calculation; and finally, de-duplication is carried out in combination with editing distance calculation and semantic similarity calculation, and high-quality question and answer pairs are output. According to the method, the problems of unstable generation quality, lack of answer basis and the like are solved, and automatic construction of the high-credibility knowledge base is realized.
Owner:山东齐鲁壹点传媒有限公司 +1

Log analysis method and computing device

The application discloses a log analysis method and a computing device. The method comprises the following steps: obtaining log content of a target line of a test log; determining whether the number of lines of the test log is greater than a first threshold value when the log content does not contain a target keyword; when the number of lines of the test log is greater than the first threshold value, calculating the jasccard similarity between the log content and each keyword; when the similarity is greater than or equal to a second threshold value, determining that the keyword corresponding to the similarity is the target keyword; when the number of lines of the test log is less than the first threshold value, calculating the edit distance between the log content and the target keyword; when the edit distance is less than or equal to a third threshold value, determining that the keyword corresponding to the edit distance is the target keyword; and then obtaining a log analysis result according to the target keyword. In this way, the analysis of the log content can be more efficient and accurate.
Owner:YUXIN ELECTRONIC TECHNOLOGY GROUP CO LTD

Using fuzzy matching to determine whether segment(s) of responsive content, that is generated using generative model(s), match segment(s) of additional data

Some implementations described herein relate to determining whether to modify segment(s) of responsive content, that is generated using a generative model (GM), based on a corresponding edit distance between the segment(s) of the responsive content and segment(s) of additional data. Processor(s) of a system can: receive user input associated with a client device, generate the responsive content using the GM and based on processing the user input, and determine whether to modify the segment(s) of the responsive content using the corresponding edit distance (e.g., fuzzy matching). Subsequent modification and / or processing of the responsive content can be dynamically adapted based on whether there is a match and, if there is a match, based on source(s) associated with the additional data that match. Further, the processor(s) can cause the responsive content, or modified responsive content, to be rendered at the client device.
Owner:GOOGLE LLC

Text data processing method, device, equipment and medium

The embodiments of the present application provide a text data processing method, apparatus, device, and medium, which relates to the field of artificial intelligence, and includes: obtaining a first text pair and a second text pair, obtaining a first subtext from the first text pair, and obtaining a second subtext from the second text pair; determining the edit distance between the first subtext and the second subtext, and if the edit distance satisfies a similarity condition, generating a first target subtext associated with the semantic information of the first subtext and belonging to a third language type, and generating a second target subtext associated with the semantic information of the second subtext and belonging to a second language type; generating a text sample pair based on the first text pair, the second text pair, the first target subtext, and the second target subtext. By using the present application, text sample pairs composed of different language types can be generated, thereby increasing the quantity of the corpus while ensuring the quality of the corpus.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Semantic generation method and device, equipment, medium and product

The invention discloses a semantic generation method and device, equipment, a medium and a product. The method comprises the steps that a labeled sentence pattern set corresponding to a text corpus of a vehicle-mounted service is generated through a sequence labeling model; constructing tree structures of all sentence patterns in the labeled sentence pattern set to obtain all tree structures corresponding to the labeled sentence pattern set; and combining all the tree structures based on the editing distance to obtain corpus semantics of the text corpus, and if the corpus semantics do not meet a preset condition, correcting the corpus semantics to obtain corrected corpus semantics. The annotated sentence pattern set corresponding to the text corpus of the vehicle-mounted service is generated through the sequence annotation model, the annotation efficiency is improved, and the sentence patterns are clearly presented by constructing the tree structures of all the sentence patterns in the annotated sentence pattern set, so that all the tree structures are conveniently combined through the editing distance, the combined tree structures are more accurate, and the annotation efficiency is improved. And thus, the obtained corpus semantics are more accurate.
Owner:IFLYTEK CO LTD

Text alignment method and apparatus, electronic device, and storage medium

ActiveCN116306557BText alignmentEngineering
Embodiments of the present application provide a text alignment method and device, electronic equipment and storage medium, belonging to the field of artificial intelligence. The method comprises: obtaining a preset reference text, obtaining and calculating a sliding window ratio according to an original text segment to obtain an original sliding window ratio; calculating an initial edit distance according to the original sliding window ratio and the reference text; screening the original sliding window ratio according to the initial edit distance to obtain an initial sliding window ratio; contracting the original sliding window ratio according to the initial sliding window ratio to obtain a target sliding window ratio; calculating a target edit distance according to the target sliding window ratio and the reference text; screening the target sliding window ratio according to the target edit distance to obtain a current sliding window ratio; and aligning the original text segment and the reference text according to the current sliding window ratio to obtain a target text. The embodiments of the present application realize text alignment under the condition of a small amount of resources.
Owner:PING AN TECH (SHENZHEN) CO LTD

A text content comparison matching method, device, equipment and medium

PendingCN122366407AEngineeringThresholding
This invention relates to the field of artificial intelligence technology, and discloses a text content comparison and matching method, apparatus, device, and medium. The method includes sequentially performing character adaptation and normalization, phonological encoding conversion, segmentation processing, fuzzy weighted edit distance calculation, peak analysis and path backtracking, and jump threshold traversal judgment on template text and speech recognition text to obtain text comparison and matching result data. In this invention, addressing the problem that existing text content comparison and matching schemes struggle to accurately identify the continuous inclusion of speech recognition text in template text, character adaptation and normalization and phonological encoding conversion can eliminate character form differences; segmentation processing and fuzzy weighted edit distance calculation can suppress accumulated errors; and peak analysis, path backtracking, and jump threshold traversal judgment are then applied. This method can be applied to the financial and medical fields, thus accurately identifying the continuous inclusion of speech recognition text in template text at the character granularity.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-stage spelling error correction method and device for medical image report and medium

The invention relates to a multi-stage spelling error correction method and device for a medical image report and a medium, in a non-term error correction stage, for each text segment, a pinyin error correction knowledge base and a polyphone probability base are called, screening and replacement are carried out according to text similarity and editing distance, candidate correction for general spelling errors is generated, and the spelling error correction efficiency is improved. Obtaining a text fragment after non-term error correction; in the term error correction stage, a medical term package is called, candidate correction for term spelling errors is generated and screened through candidate interval positioning, set similarity screening and sequence similarity calculation, and text segments after term error correction are obtained; and enabling the text segments after non-term error correction and the text segments after term error correction to correspond to each other at the text position, correcting the candidates with overlapping or conflicting, and performing resolution and combination according to a preset priority rule to obtain a final error correction result. Compared with the prior art, the method has the advantages of wide range, high accuracy, high reliability and the like.
Owner:SHANGHAI EBM MEDICAL INFORMATION SYST

Image recognition, model training method and device

The present invention discloses an image recognition and model training method and apparatus. The method comprises: obtaining a character string in an image to be recognized; obtaining the edit distance of the character string, wherein the edit distance serves as a reward and penalty function; and performing a policy gradient calculation on the character string based on the reward and penalty function to obtain recognized text. This invention addresses the technical problem that, due to prior art practices where the accuracy of entire vocabulary prediction is not incorporated into OCR model training, the optimization objectives during the training and testing phases are inconsistent, resulting in reduced recognition performance.
Owner:ALIBABA GROUP HOLDING LTD