Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

50 results about "Rewriting" patented technology

In mathematics, computer science, and logic, rewriting covers a wide range of (potentially non-deterministic) methods of replacing subterms of a formula with other terms. The objects of focus for this article include rewriting systems (also known as rewrite systems, rewrite engines or reduction systems). In their most basic form, they consist of a set of objects, plus relations on how to transform those objects.

Medical dialogue search term rewriting method and device, equipment and storage medium

The invention discloses a medical dialogue search term rewriting method and device, equipment and a storage medium. The method comprises the following steps: detecting whether current dialogue input is a question or has intention jump through a lightweight model for multi-round dialogue intention jump and question recognition; when it is detected that the current dialogue input is subjected to intention hopping or not questioning, the historical dialogue context is emptied, and the current input serves as independent query to be directly output; and when it is detected that the current dialogue input does not have intention hopping and the current dialogue input is a question, triggering the functions of substitution disambiguation and semantic completion, calling the fine-tuning and rewriting large model, and generating a standard medical question with complete semantics and standard terms, so that the context dependency relationship in the user dialogue can be accurately identified, and the user experience is improved. In addition, a large model can be called to complete deep semantic reconstruction when necessary, light model pre-judgment and large model accurate rewriting are combined, accuracy and efficiency are both considered, and the method is a key technical breakthrough for improving usability and specialty of a medical RAG system.
Owner:XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV

Automatic memory management method based on context awareness and hierarchical routing

An automatic memory management method based on context awareness and hierarchical routing belongs to the field of artificial intelligence and natural language processing, and comprises the following steps: step 1, scoring context importance; 2, mixed memory extraction; step 3, performing three-level automatic routing; 4, rewriting the mixed retrieval context; and 5, privacy compliance control is carried out. The dialogue importance is automatically evaluated through multi-dimensional features (entity density, emotional polarity and user active statement), and the problem that a traditional method depends on manual intervention is solved; according to the method, the rule engine and the fine-tuning small model are combined, accuracy and efficiency are both considered, and explicit and implicit memories can be processed at the same time; according to the invention, based on confidence and content type automatic hierarchical storage, intelligent management of short-term, medium-term and long-term memory is realized; the semantic similarity and the time decay weight are fused, and it is ensured that memory retrieval is relevant and timely.
Owner:BEIJING ZHONGKE SHENZHI TECH CO LTD

Two-stage large model cognitive enhancement method, system and device and storage medium

The invention relates to the technical field of large language models, in particular to a two-stage large model cognition enhancement method, system and device and a storage medium, and the method comprises the steps: receiving a user input request, and analyzing and decomposing the request into sequential cognition step sequences according to set cognition; based on the cognitive step sequence, using a large language model to generate intermediate representations and corresponding text candidates according to corresponding steps; evaluating the text candidates generated in each cognitive step, and when an evaluation result does not meet a preset threshold value, triggering a backtracking mechanism to guide the large language model to rewrite or correct the corresponding step; and storing the intermediate representation passing evaluation and verification and the corresponding final text in a track storage library, and finely adjusting the large language model by using a reinforcement learning method based on strategy optimization by taking data in the track storage library as a training sample. According to the method and the device, instability caused by full-text rewriting is avoided, so that the model can also produce the standard text when no prompt exists.
Owner:深圳阿丽塔数据科技有限公司

Method for compressing thinking chain of reasoning large model

The invention discloses an inference large model thinking chain compression method, which comprises the following steps of: firstly, generating answer sets with different detailed degrees by utilizing multiple rounds of sampling of a basic large model, and adaptively selecting an inference chain length by adopting a dynamic quantile algorithm based on task difficulty; secondly, performing diversity rewriting and compression on the reasoning step through KL divergence constraint, and generating the shortest expression on the premise of ensuring semantic consistency; constructing positive and negative samples to guide the model to learn simple expression, and training by adopting a composite loss function including supervised learning and length perception preference optimization; according to the method, external annotation data or a teacher model is not needed, adaptive matching of the reasoning depth and the problem difficulty can be achieved, the semantic integrity is guaranteed, meanwhile, the reasoning efficiency is remarkably improved, high expandability and good cross-task migration ability are achieved, and the method is particularly suitable for large-model lightweight deployment in a low-computing-power environment.
Owner:ZHEJIANG UNIV

Batch code upgrading processing method based on abstract syntax tree and terminal

The invention discloses a batch code upgrading processing method and terminal based on an abstract syntax tree, and belongs to the technical field of software engineering and code maintaining.The batch code upgrading processing method comprises the steps that new version information and old version information of a project are obtained, and the difference between the new version and the old version of the project is analyzed; analyzing the source code of the old version of the project, and automatically converting the source code of the old version into an abstract syntax tree; defining a processing rule according to the difference between the new and old versions of the project; and traversing each node of the abstract syntax tree, and matching and converting each node of the abstract syntax tree into a new version code according to a defined processing rule. According to the method, the low-version code can be automatically converted into the AST syntax tree, then the AST syntax tree is matched and converted into the high-version code through an algorithm, manual code rewriting is not needed, and a large amount of manpower is saved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

Application programming method based on ArtNet module

The invention relates to the technical field of intelligent control network application programming, and discloses an application programming method based on an ArtNet module, which is applied to an intelligent control network. The method comprises the steps that an ArtNet module storage medium is partitioned, firmware is received and verification information is stored, the validity of the firmware is judged, the firmware is transmitted to a main controller according to a preset upgrading protocol, the main controller distributes updates, and the ArtNet module clears temporarily stored data. According to the application, the intelligent control network master controller is an internal communication bus master device, the function nodes maintain the original access logic, only an ArtNet module is newly added as a slave device to access the internal bus, an independent communication link is established through a universal interface and the master controller, the Ethernet and the external control bus are externally connected, original hardware and an IAP protocol do not need to be changed, and the communication efficiency is improved. Equipment transformation and code rewriting cost caused by module expansion are avoided; based on the programming process, original IAP logic is not interfered, repeated burning caused by data residue is avoided, and the method is suitable for firmware upgrading of intelligent control networks such as moving head lamps, stage lighting and intelligent lighting.
Owner:GUANGZHOU YINGGUANG INTELLIGENT TECHNOLOGY CO LTD

Large language model retrieval enhancement generation method based on adaptive rewriting selection

The invention provides a large language model retrieval enhancement generation method based on adaptive rewriting selection, and is suitable for the field of natural language processing and information retrieval. According to the method, a pre-trained large language model is introduced to automatically generate diversified rewriting queries, and a self-supervised rewriting sequencer is combined to perform correlation evaluation and sequencing on candidate rewriting statements. Through a context multi-arm bandit selector, the optimal rewriting number is dynamically determined according to query semantics, a high-quality rewriting subset is selected in a self-adaptive mode, and the coverage degree and precision of information retrieval are effectively improved. And the rewriting-driven knowledge retrieval module utilizes a plurality of high-quality rewriting, integration and deduplication related knowledge blocks in parallel, and continuously optimizes a Bandit strategy based on a feedback signal to realize online self-learning. Different from a traditional RAG system depending on fixed parameters and static rewriting, the method can intelligently adjust the retrieval process for complex or variable queries, and the accuracy and practicability of a retrieval enhancement generation system in an open domain and a multi-hop reasoning scene are remarkably improved. According to the method, efficient and flexible technical support is provided for intelligent question answering and knowledge discovery in a complex environment, and the application effect and popularization value of the large language model are greatly enhanced.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Program vulnerability detection method based on redundant semantic compression and large language model enhancement

PendingCN121435239ABiological modelsPlatform integrity maintainanceAlgorithmFunctional semantics
The invention relates to a program vulnerability detection method based on redundant semantic compression and large language model enhancement, which comprises the following steps of: mapping an input source code into a semantic space, and dividing the semantic space into a functional semantic subspace and a redundant semantic subspace; guiding the large language model to automatically execute at least one of variable rewriting, redundant statement deletion and control structure replacement based on a predefined cue word; on the code samples subjected to semantic cleaning, enabling the large language model to generate function-related natural language descriptions through few sample prompt, encoding the generated natural language descriptions into vectors through a text embedding model, and performing feature fusion with original code embedding to obtain enhanced features; and inputting the enhanced features into a downstream deep learning vulnerability detection model, and realizing robust defense for backdoor attacks in training and reasoning. Compared with the prior art, the method has the advantages that the robustness of the vulnerability detection model to redundant semantics is improved under the condition that expert rules are not needed, the accuracy of the vulnerability detection model is effectively enhanced, and the success rate of backdoor attacks is reduced.
Owner:SHANGHAI JIAOTONG UNIV

Unsupervised power text grading rewriting method and system

The invention relates to the technical field of data processing, and provides an unsupervised power text grading rewriting method and system, and the method comprises the following steps: obtaining a to-be-rewritten original power knowledge point text and target user group identification information; selecting a continuous prompt vector group corresponding to the target group; inputting the original electric power knowledge point text and the continuous prompt vector group into a trained text rewriting model to obtain a rewritten text meeting the style requirement of the target population; in the text rewriting model training process, information reconstruction rewards and style conversion rewards are calculated through the estimated output probability of the large language model, the rewards and the comparison rewards form a total reward for unsupervised comparison learning, and a trained text rewriting model is obtained. An unsupervised comparative learning mechanism is introduced, information reconstruction and style conversion rewards are constructed based on the output probability of a large language model to train a text rewriting model, and personalized style rewriting of the electric power knowledge text is achieved.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO

A watermark embedding and detecting method and system based on a large model

The application discloses a watermark embedding and detection method and system based on a large model, wherein the method steps comprise: obtaining a real text sequence to be embedded with a watermark; constructing a watermark embedding and detection model based on a large model; and using the watermark embedding and detection model to complete watermark embedding and detection on the real text sequence. The application realizes efficient embedding and accurate detection of the watermark by directly embedding the watermark in the model training process and optimizing the detection mechanism. The application innovatively introduces a generative adversarial network and a binocular telescope structure, which not only ensures the natural fluency of the generated text, but also significantly enhances the robustness of the watermark under adversarial attacks and natural rewriting. At the same time, the low-rank adapter technology is adopted to reduce the computational cost, so that the method can be widely applied to various pre-trained large models and has strong adaptability.
Owner:BEIJING GUANGAN LIGHTING TECHNOLOGY CO LTD

Task planning method and device based on large model feedback optimization

The invention belongs to the technical field of artificial intelligence, and provides a task planning method and device based on large model feedback optimization, and the method comprises the steps: converting the natural language description of a problem into a PDDL symbol sequence, and calculating an output probability corresponding to the PDDL symbol sequence according to an LoRA parameter; calculating a loss value between the output probability and the target PDDL symbol sequence so as to update LoRA parameters, and training to obtain a large language model with a planning generation capability; and performing executable verification on the candidate PDDL planning scheme output by the model, adjusting input context information input into the model by using an error log, and obtaining an executable planning scheme through an iterative verification process. According to the method and the device, new errors caused by full-amount rewriting are avoided by limiting the modification range, and solvability and verification pass of a planner are taken as a stop condition during iteration of the model, so that a result has evaluability and reproducibility.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Method for automatic generation of frequently asked questions

Methods and systems for generating a frequently asked questions are provided, which include defining, by a computer program executed by a computer, a first large language model (LLM) with a user query and a feedback of the user query; refining, by the computer program, the user query to a question set based on the feedback, the question set comprising one or more sentences; defining, by the computer program, a second LLM to generate a first set of question and answer pairs from a source document; defining, by the computer program, a third LLM to generate a content set from the source document based on a rewriting of the source document; selecting, by the computer program, top questions from the content set to be provided to the second LLM; and generating, by the second LLM, a second set of question and answer pairs based on the top questions.
Owner:JPMORGAN CHASE BANK NA

Semantic focusing test question duplicate checking method based on large language model

The invention discloses a semantic focusing test question duplicate checking method based on a large language model, and the method comprises the steps: completing the construction of a question corpus and metadata labeling based on corpus construction and text standardization; mapping the topics in the corpus into dense vectors by adopting a semantic vectorization representation strategy, and constructing an offline semantic vector library; a semantic vector recall-SimHash denoising-Reranker model rearrangement screening mechanism is provided, and the problem that synonym rewriting cannot be recognized in traditional literal comparison is solved; a multi-level screening-large model deep judgment-online threshold value self-adaptive cooperation framework is provided, and semantic-level accurate duplicate checking is realized through real-time feedback continuous iteration; by starting a review mechanism, a vector recall threshold value and a large model deep judgment threshold value are dynamically adjusted according to data in a manual review library, so that the accuracy and efficiency of duplicate checking are maximized. According to the method, the problems of synonymous rewriting missing net, short text representation failure and static threshold false alarm / missing alarm are solved.
Owner:HARBIN INST OF TECH

Electric power material supply chain material semantic retrieval method based on large language model seed problem expansion

The invention discloses an electric power material supply chain material semantic retrieval method based on large language model seed problem expansion, and provides a seed problem text expansion technology in an electric power material supply chain material low-resource scene by utilizing semantic understanding and text generation capability of a large language model. Intelligent rewriting and expansion are performed by using a very small amount of seed problems, so that the problem of difficulty in semantic retrieval of complex proper nouns and domain terms is solved, and the adaptability of the system to specific fields such as the power industry is enhanced; according to the method, the text training data is expanded, manual intervention and manual labeling are avoided, and the automation level of the system is improved.
Owner:INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +1

Generative question sentence rewriting method and device for improving multi-round dialogues

The invention relates to the technical field of natural language processing, and discloses a generative question sentence rewriting method and device for improving multi-round dialogues, which realize accurate rewriting of question sentences in the multi-round dialogues and improve the modeling effect of the multi-round dialogues. According to the scheme, firstly, historical information of multiple rounds of conversations before the current session round and the problem of the current round of conversations are obtained; then, according to the obtained multi-round dialogue historical information and the round dialogue problem, an input feature vector is obtained; and finally, according to the input feature vector and the current round of dialogue problem, through a pre-trained encoder-decoder question sentence rewriting model, outputting the current round of rewritten dialogue problem.
Owner:PANOVASIC TECHNOLOGY CO LTD

Log analysis method and system based on large language model

The invention provides a log analysis method and system based on a large language model, and belongs to the technical field of log analys.The method comprises the steps that anti-fact rewriting is conducted on a to-be-analyzed log, and multiple log variants are obtained; log parameters are extracted from the log variants through the large language model to serve as intermediate variables, and the confidence coefficient of each intermediate variable is determined; performing correctness scoring on each intermediate variable based on statistical information of a historical log corpus knowledge base; determining a comprehensive score of each intermediary variable according to the confidence coefficient of the intermediary variable and the correctness score of the intermediary variable; and screening the medium variable with the highest comprehensive score as a log parameter of the log to be analyzed, and generating a log template of the log to be analyzed according to the log parameter of the log to be analyzed. According to the log analysis method, the stable semantic structure of the log can be more emphasized in the log analysis process, and the accuracy and generalization ability of log analysis are improved.
Owner:WUHAN UNIV

Multi-agent reward function automatic generation and optimization method and system

The invention provides a multi-agent reward function automatic generation and optimization method and system, and relates to the technical field of artificial intelligence and reinforcement learning. According to the method, state description information of an environment is extracted through an environment context construction module, and then a plurality of reward function candidates are evaluated at the same time through a parallelized reward function evaluation module by utilizing a reward function code generated by a large language model; afterwards, statistical information in the training process is converted into structured natural language feedback through a reflection report generation module, so that the large language model can understand the training dynamics and improve the reward function in a targeted manner, and then context-based directional rewriting of the reward function is realized; and finally, through an iterative optimization process, generating a reward function which not only conforms to a task target but also has good training characteristics. According to the method, the code generation capability of the large language model is combined with the reinforcement learning training process, so that automatic generation and iterative optimization of the reward function are realized.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Method for automatic generation of frequently asked questions

Methods and systems for generating a frequently asked questions are provided, which include defining, by a computer program executed by a computer, a first large language model (LLM) with a user query and a feedback of the user query; refining, by the computer program, the user query to a question set based on the feedback, the question set comprising one or more sentences; defining, by the computer program, a second LLM to generate a first set of question and answer pairs from a source document; defining, by the computer program, a third LLM to generate a content set from the source document based on a rewriting of the source document; selecting, by the computer program, top questions from the content set to be provided to the second LLM; and generating, by the second LLM, a second set of question and answer pairs based on the top questions.
Owner:JPMORGAN CHASE BANK NA

Language model watermarking method based on structured language features

The invention discloses a language model watermarking method based on structured language features, and aims to solve the problem that the existing watermarking technology depends on fixed keywords and is insufficient in stability in scenes of model fine tuning, pruning, text rewriting and the like. And generating a semantic-preserving trigger sample through structure controllable rewriting, so that the model forms a stable and distinguishable response to the structured language features. In a watermark embedding stage, a constraint mechanism based on model representation response is introduced, and joint optimization with an original training target of a language model is carried out, so that watermark related characteristics are embedded into a model representation layer in a dispersed form, and the robustness of the watermark under a parameter disturbance condition is improved. In the watermark verification stage, model parameters or intermediate representation do not need to be accessed, and watermark judgment can be completed only by comparing the output response difference of trigger input and common input.
Owner:BEIJING UNIV OF POSTS & TELECOMM

An open source dialogue model-oriented automatic jailbreak prompt word generation and attack method and system

The application discloses an open source dialogue model-oriented automatic jailbreaking prompt word generation and attack method and system. The application first screens a prompt word set from an original prompt word set; secondly, based on an attack question, the prompt word with the optimal attack efficiency is screened out from the original open source prompt word set, multi-path parallel testing is carried out by using a proxy model, and the attack success rate and the prompt word length of different prompt words in the prompt word set are evaluated in real time through a dynamic evaluation mechanism based on a greedy selection strategy, the final attack prompt word is output to a target model, and after being spliced with the attack question, the attack prompt word is returned and an analysis report is output; finally, the prompt word in the analysis report is adjusted by using an automatic mutation and expert modification strategy, and the prompt word is reconstructed and iterated according to the feedback result of each round of attack, so that a new prompt word is generated. The application combines offline evolution and online decision-making to automatically generate a high-success-rate jailbreaking prompt word, controls the length and overhead, and optimizes mutation by using semantic rewriting and logic skeleton extraction.
Owner:HANGZHOU DIANZI UNIV +1

Method and apparatus for extending data operation bit width

The embodiment of the application provides a data operation bit width expansion method and device, relates to the computer technical field, and comprises the following steps: reading variable vector information in a vector width control register, determining a maximum data operation bit width after expansion based on the variable vector information.If the width of to-be-processed data is greater than an original data operation bit width and less than or equal to the maximum data operation bit width, then based on the width of the to-be-processed data and the bit width of a single operation unit, the target number of operation units used for processing the to-be-processed data is determined.The target number of operation units is written into an operation unit control register after being reduced by one, so as to control the start of the target number of operation units, and the to-be-processed data is processed.Two control registers are newly defined to realize variable vector expansion without changing the original mode, and the data operation bit width is expanded from the original data operation bit width to a larger data operation bit width, so that the parallelization capability of data processing is improved without the need of instruction set expansion and code rewriting.
Owner:上海芯联芯智能科技有限公司

Tensor program functionalization method, tensor program functionalization equipment and tensor program product

The invention discloses a tensor program functionalization method, tensor program functionalization equipment and a tensor program product. The method comprises the following steps: acquiring a tensor program, and constructing a graph-level intermediate representation program of the tensor program; performing mutation rewriting and tensor version relation labeling on the graph-level intermediate representation program based on a memory dependency graph and an immutable operator of the graph-level intermediate representation program to obtain a rewritten graph-level intermediate representation program; the immutable operator is used for replacing a mutation operator in a graph-level intermediate representation program, and a new version tensor output by the immutable operator and an input source tensor have independent memories; and performing program optimization on the rewritten graph-level intermediate representation program to obtain a functional tensor intermediate representation program based on the static single assignment. The Graph-Level IR program is combined with the static single assignment SSA, so that the side effect of tensor mutation in the Graph-Level IR program is thoroughly eliminated, and the optimization performance of a deep learning compiler on the tensor program is improved.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Transcoding method, device and computer readable storage medium

The application provides a code conversion method, device and computer readable storage medium, the method comprising: obtaining a code project to be converted written in a first language; generating a first single file and a mapping relationship table according to at least one first code file included in the code project to be converted, the mapping relationship table being used for storing a corresponding relationship between a first identifier in the first single file and a path of a first code file defining the first identifier; performing conversion processing on the first single file to obtain a second single file written in a second language; and determining a target project according to the second single file and the mapping relationship table, the target project at least including at least one second code file. Through the method, the code conversion pass rate and efficiency can be improved, the determined target project is similar to the file structure of the code project to be converted, manual rewriting and reconstruction and the like are not required, manpower, material resources and financial resources can be saved, and the code can be quickly and accurately converted.
Owner:XFUSION DIGITAL TECH CO LTD

A method and device for constructing an incomplete speech rewriting model

The present application relates to a kind of incomplete speech rewriting model construction method and device, method includes: based on span dependency and insertion dependency dependency modeling and using node link resolution mode obtains the dependency graph of incomplete speech rewriting text editing operation;Using GPT model, the context similarity feature and / or rewriting consistency characteristic of current incomplete speech sentence are calculated, the context similarity feature and / or rewriting consistency characteristic are used to enhance the interactive inference of incomplete speech rewriting;The dependency graph score feature is fused with the context similarity feature and / or rewriting consistency characteristic, and the final feature after the feature fusion is pushed to rewrite result based on. More abundant semantic features can be provided for parsing model, and the speech rewriting effect is improved.
Owner:WUHAN UNIV

Method, device and equipment for training a rewriting model and storage medium

The present disclosure provides a method and device for training a rewriting model, a storage medium and a computer program product, relating to the technical field of computers. In the present disclosure, by using sample conversation data with judgment information, the model can not only learn which rewriting is correct, but also learn which rewriting is incorrect. The introduction of positive and negative samples helps the model learn from both correct and incorrect rewriting during the training process. Compared with the related art which only learns from the perspective of correct rewriting, the training effect of the rewriting model can be improved, thereby the accuracy of the rewriting model in generating rewritten sentences can be improved.
Owner:HUAWEI TECH CO LTD

Intelligent writing system and method based on multi-source knowledge base enhancement

The invention discloses an intelligent writing system and method based on multi-source knowledge base enhancement, and particularly relates to the technical field of intelligent writing, knowledge resources can be efficiently integrated in a writing task by comprehensively collecting structured, semi-structured and non-structured multi-source knowledge data and combining construction of a semantic vector library and a knowledge graph, and the writing efficiency is improved. The key points of each chapter can be identified and accurately extracted by decomposing a writing task, so that the understanding of a writing structure is optimized, the reliability and authority of quoted information are guaranteed by introducing retrieval quality information such as semantic relevancy, field matching degree and source credibility in a retrieval process, and the retrieval efficiency is improved. According to the method, the hierarchical structure and the causal relationship of the multi-source knowledge are combined, a writing outline can be more reasonably constructed, a full text first draft is generated, the quality of the full text can be evaluated in real time based on a self-correction mechanism of a quality optimization coefficient, a segment to be optimized is positioned, and it is ensured that the quality of the generated content meets a preset standard through iterative rewriting.
Owner:CHINA SOUTHERN POWER GRID DIGITAL GRID GROUP (GUANGDONG) CO LTD

A test question duplicate detection method based on semantic focus of a large language model

The application discloses a kind of based on the semantic focus of test question duplicate checking method of large language model, the method is based on corpus construction and text normalization, complete the construction of question corpus and metadata annotation;Using semantic vectorization representation strategy, the question in corpus is mapped to dense vector, and offline semantic vector library is constructed;Propose semantic vector recall→SimHash denoising→Reranker model rearrangement screening mechanism, solve the problem that traditional literal comparison cannot identify synonymous rewriting;Propose multi-level screening-large model deep judgment-online threshold self-adaptive cooperation architecture, through real-time feedback continuous iteration, realize semantic level accurate duplicate checking;By starting mechanism, according to the data in artificial review library Dynamic adjustment vector recall threshold and large model deep judgment threshold, to maximize the accuracy and efficiency of duplicate checking.The application solves the problem of synonymous rewriting, short text representation failure, static threshold false alarm / miss.
Owner:HARBIN INST OF TECH

Code retrieval data generation method and device and code retrieval method and device

PendingCN121957555AAddressing situations of relative scarcityeffective expansionDigital data information retrievalBiological modelsAlgorithmData retrieval
The invention provides a code retrieval data generation method and device and a code retrieval method and device, and relates to the technical field of computers.The method comprises the steps that basic code data and a plurality of preset retrieval scenes are obtained; respectively combining the basic code data with each retrieval scene to generate a plurality of pieces of first code retrieval data; and rewriting the first code retrieval data in at least one mode of replacing the basic code data, replacing the retrieval scene and replacing the language type to generate second code retrieval data. According to the embodiment, combination and rewriting are carried out on the basis of the preset retrieval scene and the basic code data, and the high-quality code retrieval data capable of being used for model training is obtained. Specifically, code rewriting is carried out on the basis of various modes of basic code data, retrieval scenes and language types, code retrieval data can be effectively expanded, and the number of samples of the code retrieval data is greatly increased.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Trusted splicing method and system for preventing context injection attack

The invention relates to the technical field of artificial intelligence, in particular to a credible splicing method and system for preventing a context injection attack. The method comprises the following steps: firstly, contextual fragments of three domain types of user input, a system preset instruction and external search content are obtained in parallel, and a unique corresponding immutable domain label is distributed to each fragment; key detection is carried out on the context segments marked as the external search content domain tags, and risk scores and grades are output; processing according to the risk level: normal use of low risk, safe rewriting or degradation of medium risk, and putting of high risk into a safe sandbox; carrying out dual-channel decoding on the processed context fragment, generating coherent content by a semantic channel, and carrying out safety protection on a strategy channel through attention shielding and weight control; and finally splicing the generated content and outputting a reliable splicing result. By applying the method and the device, fine management for context splicing can be realized, context injection attacks are prevented, and the system reliability is ensured.
Owner:RONGZHITONG TECH BEIJING