Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

914 results about "Text corpus" patented technology

In linguistics, a corpus (plural corpora) or text corpus is a large and structured set of texts (nowadays usually electronically stored and processed). In corpus linguistics, they are used to do statistical analysis and hypothesis testing, checking occurrences or validating linguistic rules within a specific language territory.

Multi-modal heterogeneous model retrieval enhancement method and system

The invention provides a multi-modal heterogeneous model retrieval enhancement method and system, and the method comprises the steps: building a knowledge and application example double-corpus based on user multi-modal query, and designing a joint retrieval mechanism to obtain a result set; mapping and scheduling to obtain feature representation through special processing channels for texts, images and audios and a Spiking neural network with a segmented trapezoidal topological structure; constructing a three-stage cascade architecture of a basic model, an advanced model and human experts, and obtaining a decision path and answer candidate set in combination with a recursive and discarding decision mechanism; a Hamiltonian graph network is used for representing a multi-modal relation, and a gradient-free descent method is used for rapidly training and optimizing model parameters; an enhanced retrieval result is obtained through cross-modal semantic alignment and dynamic retrieval window adjustment; and high-quality response is obtained through context-aware sorting and retrieval enhanced reasoning. According to the method, the multi-modal information retrieval processing efficiency and the heterogeneous model reasoning response quality are improved.
Owner:贵州中汇科技发展有限公司

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Thinking instruction data generation method and system, electronic equipment and storage medium

The invention relates to the technical field of data processing, and discloses a thinking instruction data generation method and system, electronic equipment and a storage medium, and the method comprises the steps: integrating a multi-source heterogeneous nuclear industry data corpus, and forming a standardized data source; thinking chain instruction data are constructed through an instruction generation framework driven by a large language model, multi-level security constraint processing is performed on the thinking chain instruction data, and contents which do not conform to nuclear industry security specifications are identified and eliminated; executing multi-dimensional quality evaluation on the thinking chain instruction data passing the security constraint, and removing the thinking chain instruction data which does not reach a quality threshold value; and performing real-time updating and dynamic optimization on the nuclear industry data corpus and the thinking chain instruction data based on newly collected nuclear industry field data. The method can improve the flexibility, accuracy and timeliness of instruction generation.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Artificial intelligence English general recognition large model training method and system

The invention relates to the technical field of artificial intelligence, and discloses an artificial intelligence English general recognition large model training method and system, and the method comprises the steps: segmenting text data in a standardized multi-modal training set into semantic blocks, and constructing a semantic primitive-feature mapping table; associating the multi-modal English general recognition training corpus with the text semantic primitives in the semantic primitive-feature mapping table to obtain cross-modal alignment features; deep fusion features of the cross-modal alignment features after deep fusion are extracted; adjusting difficulty distribution of training samples in the training stage to obtain preliminary parameters; performing multiple rounds of knowledge distillation on the preliminary parameters based on entity relationship data in an English culture background knowledge graph to obtain culture enhancement model parameters; performing multi-dimensional evaluation on culture enhancement model parameters, and performing iterative optimization on the model according to an evaluation result to obtain a target artificial intelligence English general recognition large model; the training effect of the artificial intelligence English general recognition large model can be improved.
Owner:FOSHAN POLYTECHNIC

Network threat knowledge automatic extraction method, electronic equipment and storage medium

PendingCN120930756AWeb data indexingSemantic analysisCyber threat intelligenceLinguistic model
The invention discloses a network threat knowledge automatic extraction method, electronic equipment and a storage medium, and the method comprises the following steps executed by a computer hardware system: collecting threat intelligence data related to an APT organization from a multi-source network security text, and processing the threat intelligence data to generate a standardized corpus; using the pre-training sentence vector model to generate semantic embedding for a corpus input text and a manual annotation example library text, and retrieving similar examples to construct an ICL prompt template; inputting a large language model subjected to LoRA fine tuning, and extracting structured triples of multiple types of entities and semantic relationships; generating standardized entity nodes and updated relation information by adopting semantic aggregation; and constructing an APT organization network threat intelligence knowledge graph and outputting a structured file. The method provides key technical support for APT attack tracing, threat situation awareness and automatic security policy generation.
Owner:GUIZHOU UNIV

Multimodal entity extraction, ontology mapping, and impact-based sentiment analysis using large language models

A method comprising retrieving one or more requirements of knowledge to be extracted; generating a prompt corresponding to the one or more requirements; validating the prompt by executing a large language model using the prompt and evaluating the response predicted by the large language model; fine-tuning the large language model using validation data generated as a result of validating the prompt; and executing the fine-tuned large language model using a text corpus to analyze one or more item reviews and generate a pair of at least one entity and a respective relationship sentiment value for the entity.
Owner:ZS ASSOCIATES INC

Medical decision-oriented multi-level knowledge graph construction and semantic reasoning method

The invention provides a medical decision-oriented multi-level knowledge graph construction and semantic reasoning method, and relates to the technical field of knowledge graphs, and the method comprises the steps: carrying out semantic segmentation and standardization processing on medical text corpora, extracting medical entities and attributes thereof, and constructing an incidence matrix; a reasoning path is mined based on recursion deep search, and an optimal path is selected by using an attention mechanism; and performing decision verification and optimization in combination with medical rules. According to the method, the accuracy and reliability of medical decision making can be improved, efficient mining and application of complex medical knowledge are achieved, and effective support is provided for clinical diagnosis and treatment.
Owner:BEIJING CORE HIGHLAND BIOTECHNOLOGY CO LTD

Translation ambiguity term accurate matching method based on fusion semantic vector space mapping

The invention discloses a fusion semantic vector space mapping-based translation ambiguity term accurate matching method, which comprises the following steps of: S1, obtaining source language ambiguity terms, context texts and a target language candidate translation list, and extracting domain tags and term matching features to form a multi-modal data set; s2, using improved XLM-R model coding to generate term-level, sentence-level and translation-level semantic vectors; s3, training a dynamic mapping matrix based on a bilingual parallel corpus, and aligning source side vectors to a shared semantic space; s4, fusing the source-side basic vector and the multi-dimensional features through a double-channel attention fusion network, and generating source-side and translation-side comprehensive semantic vectors; s5, introducing term-context attention weight to correct cosine similarity; and S6, outputting an optimal translation through normalized sorting and part-of-speech secondary judgment. According to the method, multi-field ambiguous term accurate matching is realized, the term translation precision and efficiency in professional fields are improved, and the requirements of high reliability of term translation in the fields of medicine, machinery, computers and the like are met.
Owner:XINJIANG DAWEIRAN BUILDING DECORATION GRP CO LTD

Speech synthesis method and device, vehicle and storage medium

The embodiment of the invention provides a voice synthesis method and device, a vehicle and a storage medium, and the method comprises the steps: obtaining multi-modal data which comprises a voice signal, a user instruction, a historical interaction log and bus data; determining model input parameters according to the multi-modal data; according to the model input parameters, the static corpus and a preset large model, generating a personalized script, the preset large model being used for dynamically generating the personalized script in combination with the model input parameters and the static corpus; and synthesizing the broadcast audio according to the personalized script. According to the invention, the technical problem that the voice synthesis technology in the related technology cannot adapt to the personalized demands of the user is solved.
Owner:GUANGZHOU AUTOMOBILE GROUP CO LTD

Machine Learning-Based Approach to Characterize, Triage, and Remediate Software Supply Chain Risk

PendingUS20260044609A1Platform integrity maintainanceUninitialized variableData stream
A software package is received and unpacked into multiple components comprising plural functions. Each function is lifted from machine code into static single-assignment intermediate representation and tokenized to produce semantics-preserving embeddings. Intermediate-representation data-flow features are extracted, including detection of constant static variables on a stack, stack reaching definitions, uninitialized variables, and intra-procedural aliases. For each component, the embeddings and features are input to a machine-learning model trained on semantic properties derived from a corpus of software packages to generate a software supply chain risk level. Data characterizing the risk level is provided to a consuming application. When the risk level satisfies a remediation criterion, a remediation action is initiated, including generation of a source-code patch recommendation for an identified root-cause function, insertion of a runtime guard into the component, or issuance of a security advisory for distribution to a security operations dashboard.
Owner:BINARLY INC

Methods and apparatuses for data integrity in retrieval-augmented generation (RAG) chatbots using original data sources for validation, segmentation, authorization, and monetization

A retrieval-augmented generation method for a large language model receives a user query and generates a vector embedding of the user query. The vector embedding of the user query is stored in an embedding vectors space. Also stored in the embedding vectors space are embeddings of documents retrieved from a corpus of information / A retrieval-based artificial intelligence (Al) model identifies which documents are relevant to the user query according to a similarity measure applied to the vector embedding of the user query and the vector embeddings of the documents. The generative Al model receives the user query and the documents identified as relevant to the user query and generates new content, based on the user query and the documents identified as relevant to the user query.
Owner:TECTONIQ INC

System and method for deterministically generating reproducible evaluative scores for a subject of analysis

The present invention relates to a system and method for deterministically generating reproducible evaluative scores for a subject of analysis (e.g., a security). The system comprises a processor and memory storing instructions to: receive verified data describing the subject; store this data in a fixed and version-controlled corpus to define a static analytical context; execute a large-language model (LLM) under a structured prompt framework that directs a controlled scratch-pad reasoning process for preliminary interpretations and evidence extraction; perform a multi-pass deterministic analysis of the fixed corpus to produce structured, synthesized statements as reproducible evidentiary outputs; and finally, apply a rubric-based scoring engine that converts these statements into calibrated alignment scores and aggregates them to generate a composite deterministic score. This architecture ensures reproducibility, transparency, and auditability by anchoring the flexible analysis of the LLM and the final scoring logic to a known, unchanging evidence corpus.
Owner:VIALAB

Electric power multi-mode corpus construction query method and system based on sliding window

The invention discloses an electric power multi-modal corpus construction query method and system based on a sliding window, which is applied to the field of electric power data query, and comprises the following steps: obtaining a structured document according to electric power multi-modal data, segmenting the structured document to obtain a plurality of segmented text blocks, and storing the segmented text blocks into a database; inputting each segmented text block into a large language model to generate a to-be-stored text vector and construct an electric power multi-mode corpus, when a user query request is received, generating a plurality of query variants according to query data, performing nearest neighbor search on each query variant in the electric power multi-mode corpus to obtain a corresponding nearest neighbor search result, and storing the nearest neighbor search result in the electric power multi-mode corpus. And fusing each nearest neighbor search result to generate an electric power related document set comprising a multi-modal association mark. According to the method, the semantic units can be accurately captured, semantic breakage is avoided, the retrieval continuity and coverage rate are improved, the comprehensiveness and context adaptability of retrieval results are improved, and the data retrieval requirement under the complex scene of the power industry is met.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +1

Document element rapid extraction system based on pre-training large model

The invention provides a document element rapid extraction system based on a pre-trained large model, and relates to the technical field of computer software application, the system comprises a parameter field adaptation module used for textualizing a document and constructing an industry standard corpus based on a textualized processing result, adjusting a preset language model by utilizing an industrial standard corpus; the dynamic document partitioning module is used for performing semantic segmentation processing on the industrial standard document to obtain a plurality of text blocks; the entity alignment module is used for carrying out entity and relation extraction on the text blocks and carrying out entity alignment in combination with a uniform manifold approximation and projection method; and the relation reasoning and knowledge graph completion module is used for performing completion processing on the preliminary knowledge graph and storing a completion result. According to the method, the element extraction efficiency can be directly improved without pre-defining a rule template or performing data annotation.
Owner:ANHUI BIAOXINCHA DATA TECH CO LTD

Robot operation and maintenance long thinking chain corpus generation method based on knowledge graph

The invention discloses a robot operation and maintenance long thinking chain corpus generation method based on a knowledge graph. The method comprises the steps of establishing an original corpus, generating a cleaning corpus, generating a deduplication corpus, performing intelligent blocking, generating question and answer pairs, performing entity recognition, performing entity mapping, determining a target reasoning path and generating a long thinking chain corpus. According to the method, a knowledge graph reasoning path is introduced as a generation constraint, so that model illusion is effectively inhibited, and explicit reasoning, logic completeness and full-link traceability of the reasoning process are realized; and meanwhile, in combination with puzzle-driven partitioning and domain model fine tuning, the semantic coherence and professional accuracy of the corpus are remarkably improved, and low-cost, large-scale and high-quality operation and maintenance corpus automatic production is realized.
Owner:SOUTH CHINA UNIV OF TECH

Equipment agent processing program conversion method and device based on cooperation of large model and small model

The invention discloses an equipment agent processing program conversion method and device based on cooperation of a large model and a small model. An equipment agent is constructed by analyzing control systems, kinematics structures and process constraint information of a source machine tool and a target machine tool, and line-level instructions are abstracted into system-independent unified semantic representation based on a multi-system numerical control corpus. And performing grammar analysis and semantic mapping on the source program, and generating a conversion task in combination with grammar rules, kinematics accessibility and process security constraints of a target machine tool. A small model is finely adjusted through supervised learning to realize basic generation capability, a large model is introduced as a patch evaluator, candidate results are jointly scored from different dimensions, knowledge migration is realized based on improved near-end strategy optimization, and a small model generation strategy is continuously optimized. And finally, high-quality results are screened through confidence degree sorting, cross-system G / M instruction and multi-axis track generation and verification are completed, and high-precision automatic conversion of numerical control programs is achieved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

A rule corpus-based text specification marking method and system

The application relates to the technical field of text label marking, and provides a text specification marking method and system based on a rule corpus, which comprises the following steps: analyzing a policy and regulation document, identifying and marking condition morphemes and conclusion morphemes in the policy and regulation document, constructing a logical relationship between the two by using a large language model, and forming a rule corpus composed of structured morpheme pairs; performing semantic embedding on the corpus to generate a semantic vector library; performing multi-label coding on a verification data set based on the rule corpus, and constructing a multi-label training data set; training a deep learning classification model by taking semantic vectors as features and multi-labels as targets, so that a text specification marking model is obtained; and automatically marking target text by using the model. The application significantly improves the accuracy, interpretability and business adaptability of text marking, improves the update quality of a system label data set, and reduces the system maintenance cost.
Owner:SSE INFORMATION NETWORK LTD

Domain-specific text labelling using natural language inference model

ActiveUS12524624B2Natural language translationSemantic analysisNatural language inferenceDisplay device
In an embodiment, a set of texts associated with a domain is received. A set of hypothesis statements associated with the domain is received. A pre-trained natural language inference (NLI) model is applied on each of the received set of texts and on each of the received set of hypothesis statements. A second text corpus associated with the domain is generated. The generated second text corpus corresponds to a set of labels associated with the domain. A few-shot learning model is applied on the generated second text corpus to generate a third text corpus associated with the domain. The generated third text corpus is configured to fine-tune the applied pre-trained NLI model, and the fine-tuned NLI model is configured to label an input text associated with the domain. A display of the labelled input text on a display device is controlled.
Owner:FUJITSU LTD

Power document keyword extraction method based on Prompt and knowledge graph

The invention provides an electric power document keyword extraction method based on Prompt and a knowledge graph, relates to the technical field of electric power document processing, and constructs a lightweight multi-level index knowledge graph in the electric power field by combining entity type and relation type division based on an electric power industry standard document and an electric power field corpus. The method comprises the following steps: performing vector modeling on a power document, constructing a multi-level index from an entity to a vector, realizing standardized semantic modeling and efficient hybrid retrieval of a power document field background, and obtaining a topic vector and a core paragraph of the power document in combination with power key information; according to the method, entity types are indexed in a knowledge graph by using subject vectors, similar entities are obtained to form knowledge sub-graphs, so that multilayer Prompt is obtained to guide a large language model to extract keywords, then knowledge graph similarity constraints are introduced to decode the output of the large language model, the keyword recognition capability in the power field is improved, and the keyword recognition efficiency is improved. And the accuracy of keyword type identification and the normalization of term naming are both considered.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD SHAOXING POWER SUPPLY CO

FPGA (Field Programmable Gate Array) netlist-level hardware Trojan horse detection method based on large language model

The invention discloses an FPGA (Field Programmable Gate Array) netlist-level hardware Trojan horse detection method based on a large language model, and relates to the technical field of integrated circuit safety and hardware Trojan horse detection, and the method comprises the following steps: constructing an original FPGA netlist into a text attribute graph containing textualized attributes, generating a path text sequence through bidirectional random walk, and carrying out two-way random walk on the path text sequence; constructing a corpus to pre-train a large language model; then, delimiting a local neighborhood by taking each node as a center, generating a path text set, extracting semantic vector representation of the path text set by utilizing a pre-training model, and further constructing a node-level training sample set; on the basis, supervised fine tuning is carried out by combining a classifier and CB-Focal Loss, and a final model is obtained; in the reasoning stage, representation construction and discrimination are carried out on nodes to be detected, and node-level hardware Trojan horse detection is achieved. According to the method, circuit topology and semantic information can be reserved at the same time, node-level hardware Trojan positioning is achieved, the detection precision and generalization ability are improved, and the automation degree is improved.
Owner:TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI

Empirical knowledge graph construction and question answering method based on heuristic self-question answering

The invention provides an experience knowledge graph construction and question answering method based on heuristic self-question answering, and the method comprises the steps: randomly dividing a text corpus into small batches, inputting the small batches into a large language model, and obtaining a heuristic rule set through induction processing and quality evaluation; constructing question and answer pairs by using the heuristic rule set and the text content cues, converting the question and answer pairs, and then carrying out confidence calculation to obtain a structured experience triple with confidence; constructing a knowledge graph by using the structured experience triad with confidence; extracting the core fragment of the original narrative query by using the heuristic rule set and generating a refined query text; obtaining an empirical path by using the embedded vector of the refined query text and the knowledge graph; and inputting the original narrative query and the experience path into the large language model to generate a final answer. According to the method, logic confusion and content contradiction possibly caused by a traditional RAG method are effectively avoided, the generated answer is more reliable, and the reasoning process is more interpretable.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Training method of large model in satellite communication field

The invention discloses a method for training a large model in the field of satellite communication, and aims to solve the problem that a general large model is insufficient in processing capability in the professional field of satellite communication. The method comprises the specific steps that multi-modal professional corpora (structured and unstructured data) in the satellite communication field are collected; generating more than ten thousands of incremental pre-training and fine-tuning corpora through data preprocessing, performing data enhancement on the corpora by adopting a plurality of mixing strategies, and dividing a training set and a test set according to a general proportion; based on low-rank self-adaptive fine tuning technologies such as LoRA, incremental pre-training and supervised fine tuning are carried out on the general large model; a professional evaluation data set is constructed manually, a self-established evaluation data set is used for comparing semantic comprehension and generation capabilities of the satellite communication large model and the general model in the satellite communication field, and model training hyper-parameters are iterated according to an evaluation result; an RAG framework is fused, a self-built private database, a self-defined text partitioning method and a knowledge base calling strategy are integrated, knowledge base data are accurately called, and question and answer content conforming to rules and specifications is generated.
Owner:THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION

Address data matching method and related equipment

The embodiment of the invention provides an address data matching method and related equipment, and belongs to the technical field of geographic information services. The method comprises the following steps: constructing an address annotation corpus according to input address information data and a preset address database; the method comprises the following steps: generating a geographic information embedding vector according to a preset geographic information knowledge graph, performing address element analysis in combination with an address annotation corpus to obtain an address element sequence so as to construct a dictionary tree, and performing similarity screening through a spatial hierarchical matching algorithm to obtain a similar address set; generating an address embedding vector matrix through a preset word embedding vector model, and performing feature extraction through a preset semantic feature extraction model to obtain semantic-level similar features; according to input address information data, multi-dimensional character similarity matching is carried out to obtain character-level similar features, then weighted fusion is carried out in combination with semantic-level similar features, and target matching address data is determined according to a weighted fusion result. According to the embodiment of the invention, the address data matching accuracy and efficiency can be improved.
Owner:CHINA TELECOM CORP LTD

Manufacturing system risk control knowledge matching method based on semantic embedding and clustering analysis

The invention relates to a manufacturing system risk control knowledge matching method based on semantic embedding and clustering analysis, and the method comprises the following steps: collecting and preprocessing risk control text data: collecting unstructured text data of a manufacturing system history record, and obtaining preprocessed risk control text data, constructing a professional corpus for a discrete manufacturing scene; text semantic embedding generation; semantic clustering modeling: performing unsupervised clustering modeling on all semantic vectors, mining semantic association and potential structures between texts, obtaining semantic representations of risk control knowledge through a clustering algorithm, and assisting in generating clustering tags; and a risk knowledge matching mechanism.
Owner:TIANJIN UNIV

DeepSeek-based role large model fine tuning corpus automatic generation method

The invention discloses a DeepSeek-based role large model fine tuning corpus automatic generation method. The method comprises the steps of receiving a role setting document describing target role features and an optional domain knowledge base; utilizing a Prompt engineering module to construct a generation Prompt containing deep role injection information and context perception knowledge fusion based on a role setting document and an optional domain knowledge base; the deep role injection information is used for enabling the generated content to accord with characters, mood, styles, knowledge backgrounds and value views defined in the role setting document; calling a preset basic large model core, and performing text generation by using the Prompt to obtain candidate corpora; performing automatic multi-dimensional quality evaluation on the candidate corpora by utilizing a quality evaluation module; wherein the assessment dimensions of the multi-dimensional quality assessment include role consistency, fact accuracy, diversity and security; and screening out qualified corpora based on an evaluation result of the multi-dimensional quality evaluation and a preset strategy, and forming a final high-quality corpus. The invention further discloses a system, electronic equipment and a computer readable storage medium.
Owner:CHINA ORDNANCE SCI INST

Multi-language code generation method based on self-supervised pre-training

The invention discloses a multi-language code generation method based on self-supervised pre-training, which comprises the following steps: acquiring and cleaning multi-language code data to form a training corpus; the method comprises the following steps: representing code data as an abstract syntax tree, extracting a control flow diagram and a data flow diagram of the code data, and obtaining unified semantic representation through combination of a diagram encoder and a sequence encoder; designing a self-supervised pre-training task, and pre-training the semantic representation based on the training corpus; constructing a multi-language pre-training model based on the structure-improved recurrent neural tensor network and the multi-language embedding matrix; when a user inputs a natural language, generating a target language code by using the multi-language pre-training model; and target language code correction is carried out through conventional function testing and grammar checking. According to the method, multi-channel recursive combination and a hierarchical recursive expansion mechanism are combined with self-supervised pre-training, so that accurate generation and performability improvement of cross-language codes are realized.
Owner:CLOUD HI-TECH (BEIJING) TECHNOLOGY CO LTD

Low-resource language translation model training method based on large language

The invention discloses a low-resource language translation model training method based on a big language, which comprises the following steps of: continuously pre-training a basic big language model by utilizing a multilingual text corpus, and storing a plurality of middle check points as candidate models; selecting a model with the optimal downstream translation task performance from the candidate models based on the performance of the verification set, and performing instruction supervision fine tuning by using the parallel instruction data set to obtain an intermediate model; and finally, training the intermediate model by using a preference optimization algorithm to obtain a target translation model by using a preference data set consisting of preferred translation and rejected translation. According to the method, through three-step progressive training, the problems of data sparsity, insufficient model capability, single training method and the like in low-resource language translation are effectively solved, and the translation quality is remarkably improved.
Owner:北京中科闻歌科技股份有限公司 +2

Computing systems and methods for generating a response to a query based on a corpus of documents

Systems and method for generating a response to a query. The method includes using a first large language model (LLM) to generate synthetic information related to a query; generating an amended query based on the synthetic information related to the query; using an information retrieval system to retrieve, from a plurality of chunks, a set of chunks that are relevant to the amended query, wherein each chunk of the plurality of chunks is all or a portion of a document in a corpus of documents; using a second LLM to rank the set of chunks based on a relevance to the query; selecting a subset of chunks from the set of chunks based on the ranking; and using a third LLM to generate a response to the query based on the subset of chunks.
Owner:THE TORONTO DOMINION BANK