Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

32 results about "Text normalization" patented technology

Text normalization is the process of transforming text into a single canonical form that it might not have had before. Normalizing text before storing or processing it allows for separation of concerns, since input is guaranteed to be consistent before operations are performed on it. Text normalization requires being aware of what type of text is to be normalized and how it is to be processed afterwards; there is no all-purpose normalization procedure.

Intelligent composition quality evaluation method and system based on large language model

The invention relates to the technical field of artificial intelligence in the education industry, in particular to an intelligent composition quality evaluation method and system based on a large language model, and the method comprises the steps: carrying out the text normalization and semantic unit segmentation of a composition, extracting a semantic vector through a first large language model in combination with a context enhancement strategy, and positioning a semantic fracture risk position; recognizing composition core elements through a second large language model, and mapping the composition core elements back to the semantic unit sequence; constructing a demonstration logic diagram, extracting a core demonstration path and abstracting the core demonstration path into a logic role topological graph; in combination with a pre-constructed writing specification knowledge graph, comparing structural compliance, connection strength and an expected support relationship, identifying and demonstrating logic defects, and generating a global deduction item list; semantic clustering is carried out on illegal items to form an error label set, comprehensive weight is calculated in combination with historical data of students, and core weak items are positioned; according to the application, the logic analysis depth of intelligent evaluation of the argument is remarkably improved, and the pertinence and practicability of teaching feedback are improved.
Owner:DALIAN HOUREN EDUCATION TECH CO LTD

Self-adaptive text extraction method and system based on artificial intelligence

The invention discloses a self-adaptive text extraction method and system based on artificial intelligence, and the method comprises the steps: carrying out the analysis of the document structure entropy of an example document set, quantifying the noise density, geometric distortion degree and background complexity of the example document set, and carrying out the self-adaptive selection of a preprocessing assembly line intensity grade according to the above; dynamically configuring image preprocessing parameters and AI recognition model parameters, and generating a recognition engine instance to output a preliminary recognition text; after regularized coarse screening extraction is carried out based on key field description, a multi-candidate generation strategy is started for low-confidence-coefficient candidate text fragments, a multi-person cooperative verification process is triggered for lower-confidence-coefficient fragments, finally all the fragments are processed through a text standardization module, and structured text extraction information is output. According to the method, accurate adaptation of processing intensity is achieved through document quality quantitative evaluation, the extraction accuracy and system robustness of complex heterogeneous documents are effectively improved through a multi-level confidence coefficient verification mechanism, and the identification error risk caused by image quality fluctuation or rule solidification is reduced.
Owner:BEIJING VOCATIONAL COLLEGE OF ECONOMICS & MANAGEMENT (BEIJING MANAGER COLLEGE)

Streaming video subtitle generation method and system based on localized large model, and storage medium

The invention discloses a streaming video subtitle generation method and system based on a localized large model, and a storage medium, and belongs to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: constructing an asynchronous streaming audio extraction queue, and realizing parallel processing of millisecond-level response and background fragmentation of a first audio segment of a video; a decoupled voice activity detection model is adopted to carry out refined slicing on the audio stream, and noise filtering and muting are carried out; text decoding is carried out through a speech recognition model subjected to fine tuning of synthetic data, and term enhancement in professional fields is supported; carrying out multi-modal fusion and time sequence calibration on an identification result to realize sound and picture synchronization and text standardization; and finally, rendering subtitles through a real-time callback mechanism and a cross-platform interface. The method solves the problems of high delay of long video subtitle generation, time axis drift and low terminology recognition rate, is especially suitable for Linux / Windows cross-platform localization deployment, and has high privacy security and low-cost field adaptive capability.
Owner:WUHAN UNIV OF SCI & TECH

High-priority test case screening method and device, storage medium and computer equipment

The invention provides a high-priority test case screening method and device, a storage medium and computer equipment, and relates to the technical field of computers.The method comprises the steps that a demand document and a test case set are received; preprocessing the demand document, and constructing a cue word based on a preprocessing result; inputting the cue word and the test case set into a large language model to obtain an initial high-priority test case; after text normalization processing is conducted on the initial high-priority test case, a target high-priority test case is obtained through text matching processing, and the target high-priority test case is marked. According to the method, high-priority test cases can be automatically and accurately screened out in a rapid iteration software development environment, the coverage rate and efficiency of regression testing and smoking testing are improved, meanwhile, the correctness of core functions is guaranteed, and the reliability and applicability of case screening results are remarkably enhanced.
Owner:创优数字科技(广东)有限公司

Double-engine German inverse text standardization method based on FST + NN

The invention provides an FST + NN-based double-engine German inverse text standardization method, which comprises the following steps of: S1, sending a German spoken input text into an input and preprocessing layer to obtain an input representation which can be stably processed by a rule / FST and lightweight neural module; s2, the candidate fragments provided by the input and preprocessing layer are standardized through a rule and FST modular analysis layer, inverted-order composite word formation, word form change and cross-category combination of German are covered, and output consistency is ensured through a unified disambiguation strategy; s3, candidate fragments which cannot be recognized and converted by the rule and FST modular analysis layer are recognized and transferred through a lightweight neural network auxiliary layer; and S4, converting'types + standard values' given by the rule and FST modular analysis layer and the lightweight neural network auxiliary layer into texts which can be directly used through an output and post-processing layer. According to the method, a finite-state machine is taken as a core, modular modeling is performed on different types of expressions in German, and long digits are decomposed into a plurality of subunits to be identified and combined step by step through a hierarchical analysis mechanism, so that the processing difficulty is effectively reduced, and the method is provided for overcoming the defects of a rule method in flexibility and coverage. A lightweight neural network is introduced as an auxiliary module and is used for identifying unregistered words and special composite structures which are difficult to cover by rules.
Owner:BEIJING AISHU WISDOM TECH CO LTD

Question detection model for call transcript

Disclosed are some implementations of systems, apparatus, methods and computer program products for categorizing a sentence as a question. Rather than using a single model, several different models are leveraged to determine whether a sentence is a question. For example, the models can include an inverse text normalization (ITN) model, a sentence embeddings model, and a Term frequency inverse document frequency (TFIDF) model. The output of an ITN model is processed using a finite state transducer (FST) while the output of the sentence embeddings model and TFIDF model are processed using logistics regression (LR) models. A support vector machine (SVM) is then applied to the output of the FST and LR models to determine whether the sentence is a question.
Owner:SALESFORCE INC

Generating unified text using speech recognition models for conversational ai systems and applications

In various examples, generating unified text using speech recognition models for AI systems and applications is described herein. Systems and methods are disclosed that use a machine learning model that is trained to generate unified text associated with user speech, where the unified text includes punction marks, capitalizations of words, inverse text normalization formatting, end of sentence (EOS) detections, and / or end of utterance (EOU) detections. For instance, the machine learning model may receive audio data representing speech as input. The machine learning model may then process the audio data and, based at least on the processing, generate output data associated with the speech. In some examples, the output data may represent tokens, such as tokens associated with automatic speech recognition processing, punctuation and capitalization processing, EOS and / or EOU processing, and / or inverse text normalization processing. In such examples, the tokens may then be processed to generate the unified text.
Owner:NVIDIA CORP

Speaker diarization error correction

ActiveUS12555584B1Speech recognitionCorrection techniqueEncoder
Techniques for performing speaker error correction are described. In some examples, speaker error correction is a post-processing task to be performed on aligned predicted words and predicted one or more speaker identities to jointly perform speaker identities error correction and at least inverse text normalization, wherein the post-processing at least includes: predicting word and speaker contextual features from the aligned predicted words and predicted one or more speaker identities using an encoder, and predicting, from the word and speaker contextual features, inverse text normalization of the aligned predicted words and corrected speaker identities using a decoder.
Owner:AMAZON TECH INC

A speech synthesis method, device, computer equipment and storage medium

Embodiments of the present disclosure relate to the technical field of speech processing, and specifically relate to a speech synthesis method and device, computer equipment and a storage medium. The main steps of the foregoing method include: obtaining text data to be synthesized, performing language recognition on the text data, and determining at least one target language category to which the text data belongs. Based on the at least one target language category, the text data is subjected to text normalization processing to obtain normalized text. The unified phoneme sequence corresponding to the normalized text is input into a pre-trained shared acoustic model to obtain target acoustic features; and based on the target acoustic features, synthesized speech is obtained. By using a unified phoneme sequence defined based on cross-language pronunciation features as an intermediate identifier, a single shared acoustic model is used for processing, which significantly reduces the overall parameter quantity, storage space requirement and computing resource consumption of the system, thereby directly reducing the deployment and operation costs.
Owner:BEIJING BAILONG MAYUN TECH CO LTD

Systems and methods for providing a resolution recommendation service

Systems and methods for providing a resolution recommendation service. The method includes receiving at an interface a ticket; executing at least one of a plurality of processing procedures on the ticket, wherein the processing procedures include a text cleaning procedure, a text normalization procedure, or a text tokenization procedure; and processing the ticket using a pre-trained large language model to extract a context from the ticket, the context including one or more of an element, a relationship, or an intent. The method further includes analyzing one or more knowledge base articles using the model and the extracted context to generate a resolution recommendation and presenting to a user via a user interface the generated resolution recommendation.
Owner:VIRTUSA CORP

Software development application data processing method based on AI large model

The invention belongs to the field of software development application data processing, and particularly relates to a software development application data processing method based on an AI large model, which comprises the following steps: adopting a unified acquisition and text standardization processing method for first-line feedback and work order data, and cleaning and marking description information from each terminal; semantic analysis is carried out through an AI large model, and elements such as equipment, processes, functions and problem types are extracted; a hierarchical modeling and relation extraction method is adopted for equipment, process and function information, and an equipment-process-function relation network is constructed; performing path traversal and weighted summary on the text association nodes, and calculating influence scores of each software function; performing task template filling and historical case retrieval on the high-score function; performing effect evaluation and parameter updating on a running result after the development task is online; according to the method, multi-source application data can be uniformly processed, the software function influence is quantified, the development task is automatically generated, and research and development decision refinement is realized.
Owner:QIANCHUAN NETWORK TECHNOLOGY (SHANGHAI) CO LTD

Intelligent conversion method and system for natural language requirements to selenium script based on large language model

This invention discloses an intelligent conversion method and system for natural language requirements to Selenium scripts based on a large language model, relating to the field of software engineering technology. The method includes: S1, requirement text normalization; S2, test element structuring; S3, enhanced context construction; S4, script template and location strategy generation; S5, initial script generation and execution; and S6, script adaptive optimization. This invention can accurately map unstructured natural language requirements to structured test elements and generate executable scripts. Throughout the process, it preserves sentence-by-sentence traceability and semantic alignment, allowing each automated test case to trace back to the original requirement. This significantly improves requirement-test consistency and reduces the risk of misjudgment and omission of test cases due to misunderstanding biases. By enhancing context representation, location fingerprints, and execution feedback loops, this solution enables automatically generated scripts to have runtime adaptive correction and self-healing capabilities.
Owner:HEBEI GEOLOGICAL STAFF UNIV

A cross-cultural dynamic portrait-driven multi-agent content generation method and system

This invention discloses a multi-agent content generation method and system driven by cross-cultural dynamic profiling, belonging to the fields of artificial intelligence and cross-cultural information processing technology. The method includes: S1, collecting multilingual text data through social media interfaces and performing structural verification, language identification, text normalization, deduplication, and quality filtering to construct a cross-cultural corpus knowledge base; S2, dividing the corpus into groups based on clustering algorithms, extracting group features, and constructing structured user profiles; S3, scheduling multiple agents to collaboratively perform content generation, trend analysis, and cross-cultural interaction tasks according to user profile tags and rule-based routing; S4, constructing a virtual audience cluster and pre-evaluating the generation strategy using a simulation mechanism that decouples the generation model and the evaluation model; S5, optimizing and iterating the generation strategy based on the evaluation results to form a closed-loop update mechanism. This invention effectively improves the stability and computational accuracy of the cross-language data processing link.
Owner:BEIJING TECH & BUSINESS UNIV

An elderly-oriented consultation accompanying robot and a method thereof

This application discloses a consultation and companion robot and its method for the elderly, relating to the field of companion robot technology. It acquires user signals through an input interaction module and transcribes them into raw text; a text normalization module performs spoken language mapping and semantic reconstruction of the text based on a domain dictionary and user attributes; a semantic retrieval module performs a hybrid search in a vector library to obtain a candidate set; a context filtering module filters the target policy context based on metadata tags; an inference generation module concatenates the query and context and inputs it into a large language model for source-based inference; and a speech synthesis module outputs emotional audio. This application effectively solves the problems of semantic comprehension bias and information accuracy in policy consultation for the elderly, achieving precise, reliable, and emotionally caring intelligent interaction.
Owner:HANGZHOU YALING INTERACTIVE INTELLIGENT SERVICE ROBOT CO LTD +1

A parameter-efficient method for fine-tuning large language models with cross-lingual and cross-domain knowledge transfer

This invention provides a parameter-efficient method for bidirectional fine-tuning of a large recommendation model that integrates collaborative information. The method includes: acquiring and normalizing the input text; designing a MoE architecture with multiple query experts to handle different types of users and integrating user-specific collaborative information into the normalized text; training parameters in the normalized text; and selecting the highest-scoring item as the next item to recommend to the user. This invention can adapt a large recommendation model into an effective recommendation system in a parameter-efficient manner, significantly outperforming state-of-the-art methods.
Owner:BEIJING INST OF TECH

Generating unified text using speech recognition models for conversational AI systems and applications

This paper describes, through various examples, the generation of unified text using speech recognition models for AI systems and applications. Systems and methods are disclosed that utilize a machine learning model trained to generate unified text mapped to user speech, where the unified text includes punctuation, capitalization, inverse text normalization formatting, end-of-sentence (EOS) captures, and / or end-of-utterance (EOU) captures. For example, the machine learning model can receive audio data representing speech as input. The machine learning model can then process the audio data and, based at least on this processing, generate output data mapped to the speech.In some examples, the output data can be tokens, such as tokens associated with automatic speech recognition processing, punctuation and capitalization processing, EOS and / or EOU processing, and / or inverse text normalization processing. In such examples, the tokens can then be processed to generate the unified text.
Owner:NVIDIA CORP

A test question duplicate detection method based on semantic focus of a large language model

The application discloses a kind of based on the semantic focus of test question duplicate checking method of large language model, the method is based on corpus construction and text normalization, complete the construction of question corpus and metadata annotation;Using semantic vectorization representation strategy, the question in corpus is mapped to dense vector, and offline semantic vector library is constructed;Propose semantic vector recall→SimHash denoising→Reranker model rearrangement screening mechanism, solve the problem that traditional literal comparison cannot identify synonymous rewriting;Propose multi-level screening-large model deep judgment-online threshold self-adaptive cooperation architecture, through real-time feedback continuous iteration, realize semantic level accurate duplicate checking;By starting mechanism, according to the data in artificial review library Dynamic adjustment vector recall threshold and large model deep judgment threshold, to maximize the accuracy and efficiency of duplicate checking.The application solves the problem of synonymous rewriting, short text representation failure, static threshold false alarm / miss.
Owner:HARBIN INST OF TECH

An AI large model-based software development application data processing method

The application belongs to the field of software development application data processing, and particularly relates to a software development application data processing method based on an AI large model, which comprises the following steps: adopting a unified collection and text standardization processing method for first-line feedback and work order data, cleaning and labeling description information from each terminal; performing semantic analysis through an AI large model to extract elements such as equipment, process, function and problem type; adopting a hierarchical modeling and relationship extraction method for equipment, process and function information to construct an equipment-process-function relationship network; performing path traversal and weighted aggregation on text associated nodes to calculate the influence score of each software function; performing task template filling and historical case retrieval on high-score functions; and performing effect evaluation and parameter updating on the running results of the online development task; the application can uniformly process multi-source application data, quantify the influence of software functions and automatically generate development tasks, and realizes fine research and development decision-making.
Owner:QIANCHUAN NETWORK TECHNOLOGY (SHANGHAI) CO LTD

Multi-layer text representation and semantic reconstruction method and system for chinese scenario text

The application discloses a multi-layer text representation and semantic reconstruction method and system for Chinese scene text, and relates to the fields of computer vision and natural language processing. In view of the problems that the prior art cannot distinguish between creative expression and real error, the processing flow is rigid and lacks interpretability, the application first detects and identifies image text to generate original text; then performs spelling correction and grammar correction on the original text to obtain standardized text; then extracts and analyzes the differences between the original text and the standardized text to generate explanation information containing error types and sources; finally, according to the application scene requirements, the original text, the standardized text or the explanation information is flexibly output. The application improves the text processing accuracy, retains the creative expression, and enhances the interpretability and scene adaptability of the system, and is suitable for intelligent translation, advertising creativity and visual impairment assistance scenes.
Owner:HARBIN ENGINEERING UNIVERSITY SANYA NANHAI INNOVATION & DEVELOPMENT BASE +1

An AI-based after-sales service quality inspection analysis method and system

PendingCN122636217AText alignmentData pack
The application relates to the technical field of computers, in particular to an AI-based after-sales service quality inspection analysis method and system, which comprises the following steps: obtaining after-sales interaction data, including customer voice, voice transcription text, complaint work order text, online evaluation text, satisfaction score and business field; performing deduplication, missing arrangement, text normalization, turn division, voice text alignment and field normalization on the after-sales interaction data to obtain standardized quality inspection samples; performing text semantic understanding and acoustic state recognition based on the standardized quality inspection samples to obtain customer emotional states by fusion; identifying after-sales service items and quantifying emotional results, combining work order attributes and satisfaction scores to form a quality inspection sample set; generating complaint attribution labels, after-sales problem abstracts, business responsibility clues and processing urgency levels by using a semantic quality inspection model, training a satisfaction risk prediction model and calculating the global importance of each after-sales service item; and generating quality inspection analysis results accordingly.
Owner:SACCO (SHENZHEN) TECH CO LTD

Generating unified text using speech recognition models for conversational AI systems and applications

The invention relates to generating unified text using speech recognition models for conversational AI systems and applications. In examples, generating unified text for AI systems and applications using a speech recognition model is described herein. Systems and methods are disclosed for generating unified text related to user speech using a trained machine learning model, wherein the unified text includes punctuation marks, word capitalization, inverse text normalization formats, sentence end (EOS) detection, and / or utterance end (EOU) detection. For example, a machine learning model may receive, as input, audio data representing speech. The machine learning model may then process the audio data and generate output data related to the speech based at least on the processing procedure. In some examples, the output data may represent markings, such as markings related to automatic speech recognition processing, punctuation and capitalization processing, EOS and / or EOU processing, and / or inverse text normalization processing. In these examples, the tokens may be processed to generate unified text.
Owner:NVIDIA CORP

Network threat intelligence automatic extraction method based on multi-source fusion

PendingCN121809666ABiological modelsNatural language data processingCyber threat intelligenceLinguistic model
The invention provides an efficient and accurate network threat intelligence automatic extraction method, and aims to solve the problems of incomplete entity recognition, inaccurate relation reasoning, easy model illusion and the like when an existing intelligence extraction technology faces CTI texts with multi-source isomerism, complex safety terms and implicit relation expression. According to the method, the advantages of a deep learning model and a large language model are fused, and the structured understanding ability of complex threat intelligence is comprehensively improved. According to the specific technical scheme, firstly, multi-source data from security reports, technical blogs and the like are processed in a unified mode through a text standardization and entity preliminary screening module, and the basic quality of information extraction is improved; secondly, an entity-driven attention model is introduced, threat entity semantics are recovered through external knowledge enhancement and an entity-to-attention mechanism, a preliminary relation is recognized, and the extraction accuracy is improved; thirdly, capturing potential attack chain logic and implicit association by adopting an example retrieval mechanism based on relational logic driving and combining the analogy reasoning capability of a large language model; and finally, through a decision fusion and arbitration mechanism, consistency comparison, conflict verification and deletion completion are carried out on results of the deep model and the large model, so that the accuracy, integrity and robustness of network threat intelligence extraction are remarkably improved.
Owner:GUIZHOU UNIV

Dynamic modification and original definition backtracking method and system for database view

PendingCN121958272AFlexible and lossless structural modificationavoid cumbersomenessDatabase updatingSpecial data processing applicationsTable (database)Datasheet
The invention discloses a dynamic modification and original definition backtracking method and system for a database view. The method comprises the following steps: receiving a CREATE OR REPLACE VIEW command containing field addition and deletion and sequence or type change, and analyzing to generate a target column attribute list; an old column attribute list is obtained, after comparison, a deleted column is marked as a deleted column, a newly added column is added, and sequence / type column attributes are adjusted; and updating the total column number metadata of the view, traversing the dependency table to set the dependency object to be invalid, and triggering the first access to automatically recompile. Meanwhile, a complete original text containing an annotation format is captured through an analysis state machine, an original text-normalized text dual storage mechanism (two columns of a system metadata table are stored respectively) is adopted, the original text is returned when a user requests, and the normalized text is adopted when internal recompiling is carried out. According to the method, flexible lossless modification of the view, stable attribute, dependence on automatic management and accurate backtracking of original definition are realized, the operation and maintenance cost is reduced, and the system reliability is improved.
Owner:BEIJING VASTDATA TECH

Multi-agent collaborative dialogue data construction method and device, equipment and medium

The invention discloses a multi-agent collaborative dialogue data construction method, device and equipment and a medium, and is applied to the scene of the financial and medical field. The multi-agent collaborative dialogue data construction method comprises the following steps: acquiring an original call record and an automatic voice recognition text by connecting an external business system; sequentially carrying out speaker separation and stage division, text standardization and multi-dimensional structured labeling to form standard structured data; performing multi-dimensional quality evaluation and fine-grained scoring on the data in combination with a service achievement state, and screening high-quality dialogue fragments; generating strategy reasoning annotation data of each round of verbal skill based on the scoring result, and performing data expansion under the constraint of strategy consistency to generate an expansion sample containing positive and negative samples; and finally, integrating various types of data to construct a multi-format training sample set oriented to large language model fine tuning. According to the method, the conversion capability and strategy level of the intelligent sales model are remarkably improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Adaptive fitting model training method and device, computer device and storage medium

The application discloses a self-adaptive fitting model training method and device, computer equipment and a storage medium. The method comprises the following steps: obtaining a training scene text image; performing batch normalization processing on each text in the training scene text image through a text self-adaptive model to be trained to obtain synthesized text normalization features and real text normalization features; performing feature weight sorting on each real text normalization feature to obtain a first sorting result, and determining real text loss information according to the first sorting result; performing feature weight sorting on each synthesized text normalization feature to obtain a second sorting result, and determining synthesized text loss information according to the second sorting result; and adjusting model parameters of the text self-adaptive model to be trained according to the real text loss information and the synthesized text loss information until a preset training end condition is met, so that a trained text self-adaptive model is obtained. The method can make up for the gap between synthesized texts and real texts in scene texts.
Owner:SHENZHEN SMARTMORE TECH CO LTD

System and method for adaptive text sampling and summarization of qualitative responses in a communication exchange environment

A system and method for text summarization is described. A transformation computer receives thought objects containing text inputs and queries. The transformation computer performs text normalization, determines a dynamic token capacity threshold based on system requirements and text characteristics, and generates sampled subsets using random or stratified sampling techniques. The system combines text processing instructions with sampled texts to create structured prompts, processes them through a transformer, and outputs summarized content in predetermined formats with associated metadata.
Owner:FULCRUM MANAGEMENT SOLUTIONS

Systems and methods for providing a resolution recommendation service

Systems and methods for providing a resolution recommendation service. The method includes receiving at an interface a ticket; executing at least one of a plurality of processing procedures on the ticket, wherein the processing procedures include a text cleaning procedure, a text normalization procedure, or a text tokenization procedure; and processing the ticket using a pre-trained large language model to extract a context from the ticket, the context including one or more of an element, a relationship, or an intent. The method further includes analyzing one or more knowledge base articles using the model and the extracted context to generate a resolution recommendation and presenting to a user via a user interface the generated resolution recommendation.
Owner:VIRTUSA CORP

Custom display post processing in speech recognition

Solutions for custom display post processing (DPP) in speech recognition (SR) use a customized multi-stage DPP pipeline that transforms a stream of SR tokens from lexical form to display form. A first transformation stage of the DPP pipeline receives the stream of tokens, in turn, by an upstream filter, a base model stage, and a downstream filter, and transforms a first aspect of the stream of tokens (e.g., disfluency, inverse text normalization (ITN), capitalization, etc.) from lexical form into display form. The upstream filter and / or the downstream filter alter the stream of tokens to change the default behavior of the DPP pipeline into custom behavior. Additional transformation stages of the DPP pipeline perform further transforms, allowing for outputting final text in a display format that is customized for a specific user. This permits each user to efficiently leverage a common baseline DPP pipeline to produce a custom output.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Efficient hybrid text normalization

Methods and devices to efficiently normalize text by processing inputted text based on a text normalization model that includes processing the input text in a first stage including a statistical model as a first output, processing the first output in a second stage including a rule based model as a normalized text, and outputting the normalized text.
Owner:TENCENT AMERICA LLC

Automatic intellectual property document retrieval method and system based on machine learning

The invention provides an intellectual property document automatic retrieval method and system based on machine learning, and the method comprises the steps: carrying out the preprocessing of a user retrieval request, including text standardization, BERT word segmentation, stop word and punctuation filtering and word form normalization, and obtaining structured input; deep shared semantic features are extracted by using a multilayer Transform network, and multi-task training such as keyword intention, technical field and legal state prediction is executed in parallel; a gradient analysis and modulation mechanism is adopted, gradient conflicts are recognized and relieved based on gradient cosine similarity between tasks, and modulation parameters are dynamically optimized in combination with gradient memory and reinforcement learning agency, so that the training stability and the multi-task generalization ability of the model are effectively improved; the intellectual property document retrieval method is beneficial for improving multi-dimensional understanding and decision accuracy of intellectual property document retrieval.
Owner:HAINAN KEYING INFORMATION TECHNOLOGY CO LTD