Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1339 results about "Text corpus" patented technology

In linguistics, a corpus (plural corpora) or text corpus is a large and structured set of texts (nowadays usually electronically stored and processed). In corpus linguistics, they are used to do statistical analysis and hypothesis testing, checking occurrences or validating linguistic rules within a specific language territory.

Multi-source heterogeneous corpus fusion method and system based on government affair service data

The invention provides a multi-source heterogeneous corpus fusion method and system based on government affair service data, and the method comprises the steps: obtaining an original corpus set of a plurality of data sources in government affair service, carrying out the cross-modal semantic alignment processing of each corpus unit in the original corpus set, generating a normalized data block corresponding to each corpus unit, and carrying out the fusion of the data blocks; carrying out multi-modal semantic coding on the standardized data blocks to obtain semantic feature vectors of all corpus units, carrying out topological structure coding on association attribute sets among the standardized data blocks to generate a global structure relation graph, and carrying out dynamic weight distribution on the semantic feature vectors based on node connection weights in the global structure relation graph to obtain semantic feature vectors of all corpus units; and generating a fusion weight matrix, performing cross-modal feature fusion on the semantic feature vector to obtain a target semantic embedding representation, and generating a standardized corpus associated with the government affair service. According to the method, the semantic aggregation problem of the non-uniformly distributed corpus units is solved, and the government affair data governance efficiency and the cross-department cooperation capability are greatly improved.
Owner:GUANGDONG YIQI DATA IND CO LTD

Enhanced query processing using domain specific retrieval-augmented generation for financial services

Embodiments of the present invention provide an innovative Retrieval-Augmented Generation (RAG) system tailored for financial analysis, significantly enhancing the precision and contextual relevance of Large Language Models (LLMs). A part of the system is a query augmentation component that leverages a knowledge graph to semantically enrich user queries, ensuring comprehensive retrieval of pertinent financial documents. A noise filtering mechanism refines the search results, while a relevance ranking component prioritizes documents based on context (e.g., user and task). The system employs prompt engineering to guide the LLM in generating responses that meet the specific requirements of financial analysis. Additionally, the LLM is fine-tuned using a corpus of financial questions and answers, reinforced by human-in-the-loop feedback, to adapt the model to the financial domain's unique linguistic and structural nuances. This advanced RAG system offers financial professionals timely, reliable, and actionable insights, providing a competitive edge in a rapidly evolving financial landscape.
Owner:AUQUAN LTD

Model analysis method and device based on knowledge graph, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business scenes of medical health, financial science and technology, culture research and the like, and discloses a knowledge graph-based model analysis method, which comprises the following steps of: acquiring and processing target domain data, extracting and labeling keywords, constructing a knowledge graph comprising domain entities, a relational network and attribute characteristics, and analyzing the knowledge graph. Establishing a professional corpus and generating an analysis task, constructing a multi-dimensional analysis index system, inputting the analysis task to a to-be-analyzed model, obtaining a task output result, performing knowledge verification according to the knowledge graph, evaluating the knowledge verification result based on a multi-dimensional analysis index, and generating a final analysis result. According to the method, by fusing the knowledge graph and the multi-dimensional analysis indexes, the accuracy and comprehensiveness of evaluation are enhanced, the evaluation effect of the model in the aspects of professional knowledge understanding, reasoning ability and standard integrating degree is improved, the evaluation result is more reliable, and the application ability of the large model in the specific field is optimized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-modal heterogeneous model retrieval enhancement method and system

The invention provides a multi-modal heterogeneous model retrieval enhancement method and system, and the method comprises the steps: building a knowledge and application example double-corpus based on user multi-modal query, and designing a joint retrieval mechanism to obtain a result set; mapping and scheduling to obtain feature representation through special processing channels for texts, images and audios and a Spiking neural network with a segmented trapezoidal topological structure; constructing a three-stage cascade architecture of a basic model, an advanced model and human experts, and obtaining a decision path and answer candidate set in combination with a recursive and discarding decision mechanism; a Hamiltonian graph network is used for representing a multi-modal relation, and a gradient-free descent method is used for rapidly training and optimizing model parameters; an enhanced retrieval result is obtained through cross-modal semantic alignment and dynamic retrieval window adjustment; and high-quality response is obtained through context-aware sorting and retrieval enhanced reasoning. According to the method, the multi-modal information retrieval processing efficiency and the heterogeneous model reasoning response quality are improved.
Owner:贵州中汇科技发展有限公司

Copilot implementation: training an expansion machine learning tool

Apparatus and methods are disclosed for implementing a copilot as a network of microservices including specialized large language models (LLMs) or other trained machine learning (ML) tools. This architecture supports flexible, customizable, or dynamically determinable dataflow. Compared to much larger competing LLMs, comparable or superior performance is achieved for certain tasks, while significantly reducing computation time and hardware requirements, even to a single compute node with a single GPU. An expansion microservice converts client input tokens into associated tokens to diversify targeted tasks presented to a core microservice. An ML tool in the expansion microservice is pretrained in multiple stages for varying tasks using varying general and target-specific corpora. Synthesized training data derived from a knowledge graph can also be used. The expansion ML tool is subsequently fine-tuned in one or multiple stages. Variations and additional techniques are disclosed.
Owner:THIA ST CO

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Traditional Chinese medicine knowledge question-answering system based on fine-tuning large model and dual retrieval enhancement

The invention provides a traditional Chinese medicine knowledge question-answering system based on a fine-tuning large model and dual retrieval enhancement. The traditional Chinese medicine knowledge question-answering system comprises a large model fine-tuning module, a dual retrieval enhancement module, a prompt template module and an answer generation module. The large model fine tuning module constructs a high-quality corpus by using traditional Chinese medicine ancient books, clinical cases and the like, performs incremental pre-training and supervised fine tuning on a ChatGLM3-6B pre-training model, and adopts technologies such as low-rank adaptation (LoRA) and direct preference optimization (DPO) to improve the adaptability of the model to traditional Chinese medicine professional knowledge. The dual retrieval module enables the model to map user questions to related contents in traditional Chinese medicine classical literatures, guidelines and modern literatures through text retrieval based on a LangChain framework on one hand, and provides structured background knowledge through map retrieval on the other hand. The system effectively makes up for the knowledge blind area of a large-scale general model in the field of traditional Chinese medicine, realizes the improvement of the accuracy, continuity and speciality of question and answer results, and has wide application prospects and relatively high innovativeness.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Milk industry knowledge question-answering method and device based on large model and RAG and medium

The invention discloses a milk industry knowledge question-answering method and device based on a large model and RAG and a medium, and the method comprises the steps: collecting multi-source heterogeneous data of the milk industry, carrying out the term protection word segmentation processing of the multi-source heterogeneous data, and generating a labeled field corpus; based on a pre-training language model base, through a domain corpus injection and adversarial learning mechanism in a domain corpus, generating an enhanced language model adapted to the dairy industry terminology; receiving a milk industry problem of a user, executing semantic vector retrieval and term extension retrieval in parallel, and screening a multi-modal retrieval result through a dynamic sorting algorithm; and splicing the multi-modal retrieval result and the milk industry question into an enhanced prompt, inputting the enhanced prompt into an enhanced language model, generating a final answer corresponding to the milk industry question, and associating the final answer with a knowledge source.
Owner:浪潮(山东)农业互联网有限公司

Adapting embeddings for custom retrieval

A data processing system implements receiving, via a user interface of a client device, a query to a large language model (LLM) working in conjunction with an embedding model comprising a corpus of artifacts. The system further implements converting the query into a query embedding, transforming the query embedding into a transformed query embedding using a first transformation function that adds a first residual term to the query embedding, the first residual term being specific for a task associated with the query, measuring a first similarity between the transformed query embedding and corpus element embeddings in an embedding space of the embedding model, determining first top-rated corpus element(s) associated with the corpus element embeddings based on the first similarity, incorporating the first top-rated corpus element(s) into the query as a prompt to the LLM to generate a response to the query, and providing, via the user interface, the response.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Intention analysis and strategy generation method and device, equipment and medium

The invention relates to the technical field of semantic analysis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an intention analysis and strategy generation method, device, equipment and medium, and the method comprises the steps: obtaining interaction data, recognizing an interaction type according to metadata features, selecting a corresponding analysis model to analyze and process data, and generating a text corpus; current user information is extracted from the text corpus, user intention score analysis is executed based on an intention analysis strategy corresponding to the interaction type, and a current user intention score is generated; and generating a demand list according to the user information, the intention score and the interaction type, and generating and outputting an interaction strategy based on the user information and the demand list. Through interaction data processing and automatic intention scoring analysis, the demand list and the interaction strategy of the customer can be quickly and accurately generated, the customer information processing efficiency and accuracy are improved, the workload of manual input and analysis is reduced, and the customer service quality and the response speed are improved.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Method and system for generating indexed corpus for domain-driven knowledge augmented question answering

Existing question answering approaches have the disadvantages that they possess limited contextual understanding due to which the retrieval process they use is inefficient in nature. Embodiments disclosed herein provide a method and system for domain-driven knowledge augmented question answering. The system receives a raw corpus data as input, wherein the raw corpus data is a domain specific data. Further, an indexed corpus is generated from the raw corpus data, during which a document chunking approach is used. The indexed corpus is then used for processing received queries received, in order to generate response to the received user queries.
Owner:TATA CONSULTANCY SERVICES LTD

Fuzzy test Kubernete-based three-party component vulnerability mining method

The invention relates to a fuzzy test Kubernete-based three-party component vulnerability mining method, which comprises the following steps of: firstly, automatically identifying a user controllable function by combining a large language model based on an official document, a configuration list and source code information of a three-party component, and constructing a keyword dictionary to assist in large-scale identification of a potential target function; then analyzing context information of the target function and an interaction relation of the target function in a Kubernetes cluster, constructing a front dependency library required by fuzzy testing, and automatically generating a fuzzy testing driving program; integrating a fuzzy test driving program with a three-party component source code, performing code instrumentation, pre-compiling and corpus optimization, generating test input according to a set variation strategy, dynamically monitoring a coverage rate and an abnormal behavior, and continuously optimizing a test corpus; and finally, based on the crash behavior of the fuzzy test record, analyzing a trigger path of the vulnerability, performing vulnerability verification in a local Kubernetes cluster environment, and finally generating a vulnerability analysis report.
Owner:HANGZHOU DIANZI UNIV

Drilling and blasting method tunnel construction safety risk evolution analysis method based on mapping knowledge domain

The invention discloses a drilling and blasting method tunnel construction safety risk evolution analysis method based on a knowledge graph, and the method comprises the steps: obtaining and preprocessing the text data of a drilling and blasting method tunnel construction accident case, and obtaining a safety risk corpus; a mixed deep learning model is constructed, the safety risk corpus is trained through the mixed deep learning model, and the risk factors and the incidence relation of drilling and blasting method tunnel safety are obtained; constructing a drilling and blasting method tunnel construction safety risk coupling evolution analysis knowledge graph based on the risk factors and the incidence relation, and obtaining a key accident chain and a key edge in combination with a knowledge query and path reasoning algorithm; and quantitatively evaluating independent values and coupling effect intensity of the risk factors by using an interaction matrix algorithm, calculating comprehensive importance scores of the risk factors, and determining key risk points according to the comprehensive importance scores of the risk factors. The problems that in a traditional method, attention to the risk factor coupling effect is insufficient, and accident case data are not fully utilized are effectively solved.
Owner:CHINA UNIV OF MINING & TECH

Enhanced domain-specific language learning models

A method and system for creating an enhanced domain-specific language learning model is disclosed. In some embodiments, the method includes training a domain language model using domain-specific data. The method includes receiving input corpus for one or more downstream tasks. The method then includes using the domain language model using the input corpus to generate a first set of embeddings, and using a pre-trained large language model (LLM) using the input corpus to generate a second set of embedding. The method further includes combining the first and second sets of embeddings to form a combined set of embeddings and perform the one or more downstream tasks using the combined set of embeddings.
Owner:GENPACT USA INC

Ancient character image recognition and semantic analysis method

The invention relates to an ancient character image recognition and semantic analysis method, which comprises the following steps of: firstly, acquiring an original image containing irregular deformation and material diversity characteristics from the surface of a cultural relic, and eliminating noise and illumination interference through adaptive filtering; calculating a main inclination angle based on stroke feature distribution, and executing inclination angle correction to obtain a processed image; thirdly, separating independent characters by adopting a region growing algorithm based on stroke features, and extracting feature vectors by utilizing a deep convolutional network for identification; and finally, matching context information in combination with an ancient text corpus, and optimizing sequence labeling by adopting a conditional random field algorithm to obtain semantic output. And if the result is not ideal, the inclination angle correction parameter is adjusted through backtracking and iterative optimization is carried out, so that the overall identification accuracy is improved. According to the method, the problems of irregular deformation and material diversity in the ancient character image can be effectively solved, and the recognition precision and the semantic analysis reliability are improved.
Owner:SICHUAN NORMAL UNIV

Speech recognition method and system for constructing small language based on whispertoken

The invention provides a voice recognition method and system for constructing a small language based on whispertoken, and relates to the technical field of natural language processing and voice recognition, and the method comprises the steps: extracting all tokens related to a target small language in a whispertoken, and forming an initial candidate set; matching and analyzing the tokens in the initial candidate set and the collected training text corpus of the target small language, and counting the occurrence frequency of the tokens in the corpus; and screening high-frequency tokens according to a frequency statistical result, and supplementing low-frequency tokens to construct a dynamic vocabulary. According to the method, the vocabulary quality is improved, the model training efficiency is optimized, the speech recognition accuracy is enhanced, the model generalization ability is improved, and the model construction process is simplified, so that an efficient, accurate and easy-to-implement solution is provided for the field of minority language speech recognition.
Owner:BEIJING RUI KELUN INTELLIGENT TECH CO LTD

Method and system for generating mediation document

The invention provides a mediation document generation method and system, relates to the technical field of data processing, and comprises the step of forming structured case data through text preprocessing, named entity recognition and legal information association based on case text data and legal corpus information. Then, generating a preliminary mediation document with placeholders by utilizing case classification, semantic matching and a sequence-to-sequence model; after standard auditing and logic auditing are carried out on the documents, optimization is carried out based on reinforcement learning, language style migration and a text generative adversarial network, and the structural rationality, law term normativity and semantic integrity are improved. And finally, a mediator preference template is constructed through Few-shot Learning and meta learning, the format and language style of the document are adjusted in a personalized manner, and a mediation document conforming to the habits of a mediator is generated. According to the method, the intelligence and the accuracy of mediation document generation are improved, the law compliance and the logic preciseness are ensured, and the mediation working efficiency is improved.
Owner:SICHUAN XINYUNDIAO TECHNOLOGY SERVICE CO LTD

Dynamic retrieval enhancement generation and deep reasoning method and framework based on intelligent agent

The invention discloses a dynamic retrieval enhanced generation and deep reasoning method and framework based on an intelligent agent, and belongs to the technical field of natural language processing, and the method comprises the following steps: S1, inputting a problem into a big language model which analyzes the problem and divides the problem into three difficulty levels, the problems comprise a simple difficulty problem, a medium difficulty problem and a complex problem; s2, respectively processing the problems according to different difficulties of the problems, and retrieving related knowledge from a corpus; and S3, processing the knowledge obtained in the S2 by the large language model, and outputting a final result. According to the dynamic retrieval enhanced generation and deep reasoning method and framework based on the intelligent agent, the performance of a large language model in various complex tasks is enhanced, especially in the aspect of multi-hop reasoning, the overall efficiency and accuracy of a system are improved through the synergistic effect between modules, and the system is suitable for large-scale popularization and application. And the method is suitable for various knowledge-intensive application scenes.
Owner:SHENYANG AEROSPACE UNIVERSITY

Thinking instruction data generation method and system, electronic equipment and storage medium

The invention relates to the technical field of data processing, and discloses a thinking instruction data generation method and system, electronic equipment and a storage medium, and the method comprises the steps: integrating a multi-source heterogeneous nuclear industry data corpus, and forming a standardized data source; thinking chain instruction data are constructed through an instruction generation framework driven by a large language model, multi-level security constraint processing is performed on the thinking chain instruction data, and contents which do not conform to nuclear industry security specifications are identified and eliminated; executing multi-dimensional quality evaluation on the thinking chain instruction data passing the security constraint, and removing the thinking chain instruction data which does not reach a quality threshold value; and performing real-time updating and dynamic optimization on the nuclear industry data corpus and the thinking chain instruction data based on newly collected nuclear industry field data. The method can improve the flexibility, accuracy and timeliness of instruction generation.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Artificial intelligence English general recognition large model training method and system

The invention relates to the technical field of artificial intelligence, and discloses an artificial intelligence English general recognition large model training method and system, and the method comprises the steps: segmenting text data in a standardized multi-modal training set into semantic blocks, and constructing a semantic primitive-feature mapping table; associating the multi-modal English general recognition training corpus with the text semantic primitives in the semantic primitive-feature mapping table to obtain cross-modal alignment features; deep fusion features of the cross-modal alignment features after deep fusion are extracted; adjusting difficulty distribution of training samples in the training stage to obtain preliminary parameters; performing multiple rounds of knowledge distillation on the preliminary parameters based on entity relationship data in an English culture background knowledge graph to obtain culture enhancement model parameters; performing multi-dimensional evaluation on culture enhancement model parameters, and performing iterative optimization on the model according to an evaluation result to obtain a target artificial intelligence English general recognition large model; the training effect of the artificial intelligence English general recognition large model can be improved.
Owner:FOSHAN POLYTECHNIC

Network threat knowledge automatic extraction method, electronic equipment and storage medium

The invention discloses a network threat knowledge automatic extraction method, electronic equipment and a storage medium, and the method comprises the following steps executed by a computer hardware system: collecting threat intelligence data related to an APT organization from a multi-source network security text, and processing the threat intelligence data to generate a standardized corpus; using the pre-training sentence vector model to generate semantic embedding for a corpus input text and a manual annotation example library text, and retrieving similar examples to construct an ICL prompt template; inputting a large language model subjected to LoRA fine tuning, and extracting structured triples of multiple types of entities and semantic relationships; generating standardized entity nodes and updated relation information by adopting semantic aggregation; and constructing an APT organization network threat intelligence knowledge graph and outputting a structured file. The method provides key technical support for APT attack tracing, threat situation awareness and automatic security policy generation.
Owner:GUIZHOU UNIV

Neural architecture search of language models using knowledge distillation

A neural architecture search method, system, and computer program product that determines, by a computing device, a best fit language model of a plurality of language models that is a best fit for interpretation of a corpus of natural language and interprets, by the computing device, the corpus of natural language using the best fit language model.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Vision-language multi-mode-based ship license plate identification method

The invention discloses a ship license plate recognition method based on vision-language multiple modes. The method comprises the following steps: S1, acquiring an original image of a to-be-recognized area through image acquisition equipment; s2, super-resolution reconstruction and data enhancement preprocessing are carried out on the collected images, and a ship license plate recognition data set is generated; s3, constructing a vision-language multi-mode ship license plate recognition model, wherein the model comprises a vision module, a language module and a fusion module; s4, inputting the data set generated in the S2 into a visual module for pre-training; s5, performing language module pre-training by utilizing a ship brand corpus; and S6, loading the pre-training weights obtained in S4 and S5, inputting the data set generated in S2 into the ship license plate recognition model, dynamically weighting the visual features and the language features by adopting a gating fusion strategy, generating multi-modal joint representation, and optimizing the ship license plate recognition model through a multi-modal fusion loss function. According to the method, the accuracy of ship license plate identification is effectively improved, and the adaptability to shielded and blurred images is improved.
Owner:ZHEJIANG UNIV OF TECH

Semantic analysis method and device based on power grid domain knowledge mining and electronic equipment

The invention discloses a semantic analysis method and device based on power grid domain knowledge mining and electronic equipment, relates to the technical field of semantic analysis, and improves the analysis accuracy. The method comprises the following steps: constructing a pre-training corpus based on an original corpus related to the power grid field, and continuously updating the corpus of the pre-training corpus by adopting a dynamic knowledge graph and incremental learning fusion system; performing unified representation processing and cross-modal semantic alignment processing of multi-modal corpus data on corpora in the pre-training corpus by utilizing a multi-modal graph attention analysis network; based on the processed pre-training corpus, constructing an analytical model, and performing compression processing on the analytical model through a knowledge distillation technology; and deploying the compressed analytical model to edge equipment in the substation power grid, continuously collecting equipment state information of the edge equipment, performing semantic analysis on the equipment state information by adopting the compressed analytical model, and outputting an obtained analysis result.
Owner:QINZHOU POWER SUPPLY BUREAU OF GUANGXI POWER GRID CO LTD

Deep learning-based teaching corpus construction method and system, and medium

The invention discloses a teaching corpus construction method and system based on deep learning, and a medium, and belongs to the cross technical field of medical education, artificial intelligence and multi-modal data processing. The method comprises the following steps: a data acquisition and annotation stage: acquiring medical inquiry video data of different scenes; in the model training stage, a multi-modal deep learning model architecture is used for carrying out cross-modal alignment on multi-modal data; then semantic understanding and question classification are carried out based on the pre-training language model after fine tuning; in the corpus construction and optimization stage, multi-modal data are integrated, a corpus management system is built, dynamic updating and self-adaptive optimization are carried out on a corpus, meanwhile, a typical non-language posture library is constructed for a regional culture background, and teaching is carried out in combination with the corpus. According to the invention, in combination with medical professional knowledge, an efficient and intelligent medical inquiry video teaching corpus is constructed through deep fusion of natural language processing, computer vision, speech recognition and multi-modal data fusion technologies.
Owner:CHONGQING MEDICAL UNIVERSITY

Multimodal entity extraction, ontology mapping, and impact-based sentiment analysis using large language models

A method comprising retrieving one or more requirements of knowledge to be extracted; generating a prompt corresponding to the one or more requirements; validating the prompt by executing a large language model using the prompt and evaluating the response predicted by the large language model; fine-tuning the large language model using validation data generated as a result of validating the prompt; and executing the fine-tuned large language model using a text corpus to analyze one or more item reviews and generate a pair of at least one entity and a respective relationship sentiment value for the entity.
Owner:ZS ASSOCIATES INC

Medical decision-oriented multi-level knowledge graph construction and semantic reasoning method

The invention provides a medical decision-oriented multi-level knowledge graph construction and semantic reasoning method, and relates to the technical field of knowledge graphs, and the method comprises the steps: carrying out semantic segmentation and standardization processing on medical text corpora, extracting medical entities and attributes thereof, and constructing an incidence matrix; a reasoning path is mined based on recursion deep search, and an optimal path is selected by using an attention mechanism; and performing decision verification and optimization in combination with medical rules. According to the method, the accuracy and reliability of medical decision making can be improved, efficient mining and application of complex medical knowledge are achieved, and effective support is provided for clinical diagnosis and treatment.
Owner:BEIJING CORE HIGHLAND BIOTECHNOLOGY CO LTD

Intelligent translation calibration method based on RAG knowledge base

The invention discloses an intelligent translation calibration method based on an RAG knowledge base, and belongs to the technical field of artificial intelligence and natural language processing. In order to solve the problems that an existing large language model (LLM) is insufficient in professional term translation accuracy in text translation, low in RAG system retrieval efficiency and the like, an intelligent calibration triggering mechanism is constructed, so that the LLM can autonomously judge whether translation content needs knowledge base enhancement or not in the modes of field detection, term density analysis, user-defined triggering and the like; a full-amount retrieval strategy is changed, and the system efficiency is improved; an RAG knowledge base retrieval framework based on function calling is constructed, multiple types of knowledge sources such as a translation memory library, a term table and a professional corpus are integrated, and accurate domain knowledge is provided for translation calibration; flexible multi-knowledge-base configuration and priority management are supported, and a user can select knowledge base types and set retrieval strategies and fusion weights according to needs; through a progressive translation optimization process of initial translation, intelligent judgment, knowledge base retrieval and calibration optimization, accurate calibration of different granularities such as term level, sentence level and paragraph level is realized, and a standardized knowledge base interface specification is provided. According to the method, the defects of LLM in the aspects of translation accuracy and knowledge base utilization efficiency in the professional field are effectively overcome, and the accuracy and efficiency of knowledge-intensive translation tasks are remarkably improved.
Owner:QINGDAO WEIWEIYAN DATA INFORMATION TECHNOLOGY CO LTD