Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

57 results about "Perplexity" patented technology

In information theory, perplexity is a measurement of how well a probability distribution or probability model predicts a sample. It may be used to compare probability models. A low perplexity indicates the probability distribution is good at predicting the sample.

Jailbreak detection for language models in conversational ai systems and applications

In various examples, systems and methods are disclosed relating to language model jailbreak detection using length-perplexity metrics. A system can identify a prompt for a language model—such as an LLM, VLM, etc.—and generate a perplexity score for the prompt. The system can determine, based at least on the perplexity score and a length of the prompt, that the prompt is indicative of a jailbreak attempt for the large language model. The system can restrict the prompt from input to the large language model—or block an output generated based on the prompt from being shared—responsive to determining that the prompt is indicative of the jailbreak attempt.
Owner:NVIDIA CORP

Prompt anomaly detection method and apparatus, and device and storage medium

PCT designated stageWO2025213839A1Semantic analysisLinguistic modelAlgorithm
The embodiments of the present application relate to the technical field of artificial intelligence. Provided are a prompt anomaly detection method and apparatus, and a device and a storage medium. The method comprises: acquiring a prompt to be subjected to detection and the perplexity of said prompt; if the perplexity exceeds a second threshold value and does not exceed a first threshold value, using an adversarial perturbation mode to process said prompt, in order to acquire a perturbed prompt; on the basis of prediction by a large language model, obtaining response semantics of said prompt and perturbed response semantics corresponding to the perturbed prompt; and if the response semantics and the perturbed response semantics are inconsistent, said prompt being anomalous. Two threshold values are set for perplexity detection, and after perturbation processing is performed on a prompt to be subjected to detection, whether the prompt is anomalous can be determined again by means of the comparison between response semantics and perturbed response semantics, so that the accuracy of prompt anomaly detection can be effectively improved, thereby enhancing the security of large language models.
Owner:CHINA UNIONPAY

System and method for intelligent evaluation of artificial intelligence generated texts

A system and method for automatically evaluating computer generated content may include: calculating a plurality of metrics for an input text, where the plurality of metrics may include one or more perplexity scores describing a prediction of the input text by a large language model (LLM); determining, based on one or more of the calculated metrics, whether to accept or reject the input text; and performing an exchange of data between remotely connected computer devices based on the determining to accept or reject the text. In some embodiments, calculating of metrics and determining whether to accept or reject the input text may be performed without relying on any information received subsequent to the initial receiving of the input text. Some embodiments may perform automated computerized actions such as, e.g., deploy or discard an update to the LLM based on the determining whether to accept or reject the input text.
Owner:ACTIMIZE LIMITED

Earth and rockfill dam illness feature mining method and system based on historical text data

The invention discloses an earth and rockfill dam illness feature mining method and system based on historical text data, and the method comprises the steps: collecting earth and rockfill dam historical illness data, constructing an earth and rockfill dam illness diagnosis text corpus set, carrying out the structured preprocessing of the corpus set, and obtaining an illness feature and danger removal measure text corpus set; generating a structured word sequence subset through a word segmentation tool; processing the structured word sequence subset by adopting an LDA topic model, determining an optimal topic number through a confusion degree curve, and outputting a final topic and a corresponding topic word; constructing a visual network graph based on the subject term co-occurrence frequency; and carrying out centrality analysis on nodes in the visual network diagram, identifying key nodes in the visual network diagram, quantitatively analyzing association rules of the danger characteristics and danger removing measures, and completing feature mining of the earth and rockfill dam danger. The problems that traditional manual diagnosis is high in subjectivity and low in utilization rate of historical engineering data are solved, and intelligent auxiliary decision making of the earth and rockfill dam danger is achieved.
Owner:NANJING HYDRAULIC RES INST

AI text recognition method and device based on ensemble learning and advanced semantic statistical feature analysis

The invention provides an AI text recognition method and device based on ensemble learning and advanced semantic statistical feature analysis, and the method comprises the steps: 1, respectively sending a to-be-recognized text into a Bert detector and a high-order natural language statistical feature detector for recognition, the high-order natural language statistical feature detector comprises a word logarithm probability detector, a word ranking logarithm detector, an Entropy detector and a confusion degree detector; and 2, performing election on detection results output by the Bert detector and the high-order natural language statistical feature detector by using an election module to obtain an AI text recognition result. According to the method, an integrated learning strategy is adopted, and a pre-training language model subjected to fine tuning is combined with high-order natural language statistical characteristics, so that when the model detects a large language model to generate a text, the strong expression ability of the pre-training language model can be fully utilized, and a deep rule of the text can be captured through the high-order statistical characteristics; and the detection accuracy is improved.
Owner:ZHENGZHOU XINDA ADVANCED TECH RES INST

Model parameter compression method and device of large language model, equipment and storage medium

The embodiment of the invention provides a model parameter compression method and device for a large language model, equipment and a storage medium. The method comprises the steps of obtaining verification data and inputting the verification data into a large language model to obtain an input activation tensor received by each network layer; obtaining the reasoning confusion degree of the large language model for reasoning the verification data, and taking the minimization of the reasoning confusion degree as an optimization target to carry out iterative cutting decision to obtain a cutting decision vector; performing outlier clipping processing on the input activation tensor of the network layer based on the clipping decision vector to obtain a target activation tensor; for each network layer, calculating an importance score of a model parameter based on the target activation tensor, and pruning the initial model parameter tensor in combination with the importance score to obtain an intermediate model parameter tensor; and performing quantization processing on the intermediate model parameter tensor of the network layer to obtain a target large language model. Therefore, the storage and calculation complexity of the large language model can be reduced while the performance of the large language model is maintained.
Owner:PENG CHENG LAB

Model training method and device, electronic equipment and storage medium

The embodiment of the invention provides a model training method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the steps of obtaining preset question and answer training data, wherein the question and answer training data comprises input question data and a model output result corresponding to the input question data; reasoning trajectory generation is carried out based on each piece of input question data and the corresponding model output result, and a plurality of candidate reasoning trajectories are obtained; performing confusion degree evaluation based on the plurality of candidate reasoning trajectories to obtain a confusion degree score corresponding to each candidate reasoning trajectory; a target reasoning trajectory is determined from the plurality of candidate reasoning trajectories based on the confusion score to guide model training. According to the embodiment of the invention, stable model training can be carried out on a large language model used for business scenes such as intelligent investment advisers or insurance claims and the like in a low-cost and high-efficiency manner, the logic stability of the large language model during business processing is improved, and business services with high professionality and depth are provided.
Owner:PING AN TECH (SHENZHEN) CO LTD

Content security protection method and device based on semantic consistency

The invention belongs to the technical field of computer information security, particularly discloses a content security protection method and device based on semantic consistency, and aims to solve the problem that high-level cue word injection attacks are difficult to effectively recognize in the prior art. Carrying out real-time interception on dominant illegal texts through a content filtering module; quantizing generation rationality differences of prompt words between attack and normal language models by using a word vector confusion degree calculation module; evaluating the logic coherence of the sentence structure of the prompt word through a statement semantic consistency judgment module; integrating the multi-dimensional feature data and carrying out risk classification by adopting a machine learning algorithm; and routing the suspected attack request to a value fine tuning model for processing according to a judgment result. According to the technical scheme, high-precision recognition and response to complex cue word injection attacks are achieved, and the content safety protection capacity and compliance guarantee level of a large model in an interaction scene are remarkably improved.
Owner:ASPIRE TECH (SHENZHEN) LTD

Large model real-time safety protection method and system based on dynamic response

The embodiment of the invention discloses a large model real-time safety protection method and system based on dynamic response. The method comprises the following steps: acquiring text data input to a target large model by a user; determining the confusion degree, gradient and semantic feature vector of the text data under the target large model based on the standby large model; splicing the confusion degree, the gradient and the semantic feature vector into a detection feature vector; based on a multi-classifier integrated detector comprising a deep learning classifier, a rule detector and an abnormal behavior detector, performing adversarial detection on the detection feature vector to obtain a plurality of adversarial detection data; carrying out weighted fusion processing on the plurality of confrontation detection data to obtain a comprehensive confrontation confidence coefficient corresponding to the text data; and starting an intelligent grading protection mechanism for the target large model according to the comprehensive antagonism confidence coefficient. According to the embodiment of the invention, the robustness of the large model in a malicious environment is remarkably enhanced, and the output reliability and safety are guaranteed.
Owner:HANGZHOU ANQUAN DIGITAL INTELLIGENCE TECH CO LTD

Mobile edge collaborative multi-modal large language model speculative reasoning method

The invention discloses a speculative reasoning method for a multi-modal large language model based on mobile edge collaboration. The speculative reasoning method comprises the following steps: constructing an equipment-side double-clue probe model; task semantic distillation and feature alignment are carried out, joint optimization is carried out on double heads of a double-clue classifier, knowledge of an edge end MLLM is distilled to a double-clue probe model, and an equipment end minimum modal delay decision is executed; asynchronous mode screening and speculative reasoning are realized through a hierarchical gating network; and executing a rollback decision to obtain a corrected final reasoning result. According to the embodiment of the invention, a lightweight dual-clue probe model is constructed at a mobile client, task semantic features are extracted from the deep layer of an edge MLLM, joint estimation of reasoning confidence and modal sufficiency is realized, and a minimum modal delay decision of an equipment end is executed; a hierarchical gating mechanism is adopted at an edge end, a confusion-guided rollback mechanism is introduced, an early reasoning result with insufficient confidence is corrected, and early recognition and efficient reasoning of a minimum full subset in an asynchronous mode are achieved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Sparse large language model weight refinement method, system and device and storage medium

The invention discloses a sparse large language model weight refinement method, system and device and a storage medium, which are corresponding schemes, and the scheme realizes training-free and plug-and-play refinement of sparse weight after one-time pruning, significantly reduces the confusion degree, improves the zero sample task performance, and especially has an outstanding effect in a high sparse rate scene. The space is optimized in a unified mode through global soft constraint, heterogeneous Hessian inversion is avoided, pruning prior is effectively utilized in combination with historical momentum, Neumann series extrapolation solution is adopted, iteration time can be remarkably shortened, and expenses of computing resources and memory resources are reduced. The method is oriented to online and operation and maintenance processes of inference services such as text generation, question and answer, classification and retrieval enhanced question and answer and the like, and accesses the current network through hot update or gray release after completing intra-layer weight refinement; the existing word segmentation coding, service gateway and monitoring and rollback mechanisms are reused in the whole process, and it is ensured that the method can be implemented under real service distribution and resource budget.
Owner:UNIV OF SCI & TECH OF CHINA

A flood situation awareness method and device based on weighted LDA algorithm

This invention discloses a flood situation perception method and apparatus based on a weighted LDA algorithm, belonging to the field of flood disaster risk management technology. The method includes: acquiring posts related to flood events through web crawling technology, performing data cleaning and secondary processing to obtain multiple target words and storing them in a dataset; calculating the weight of each target word in the posts using a term frequency-inverse document frequency algorithm; comprehensively determining the number of flood topics using perplexity, consistency index, and Jensen-Shannon divergence; introducing the weights of each target word in the posts into the LDA model, sampling the LDA distribution using a weighted Gibbs sampling algorithm, and estimating the flood post-topic distribution and topic-word distribution to analyze and perceive the flood development trend. This invention can improve the accuracy and interpretability of topic identification, thereby achieving comprehensive perception and accurate understanding of the flood situation.
Owner:SOUTHWEST JIAOTONG UNIV

Applying a sparse-dense-sparse methodology to language models

Embodiments herein describe a sparse-dense-sparse (SDS) process that achieves a better pruning scheme that benefits from pruning-friendliness relative to one-shot pruning schemes. The SDS process performs a first pruning to generate a sparse ML model followed by reconstruction to generate a re-dense ML model, followed by a second pruning to generate another sparse ML model. By pruning a ML model and then re-constructing the ML model, the ML model can be made more pruning-friendly by performing data and / or weight regularization. As a result, performing the second pruning in the SDS process can result in a smoother weight distribution and lower perplexity relative to one-shot pruning.
Owner:XILINX INC

Evolutionary software vulnerability detection method based on large language model

The application discloses a kind of based on big language model's evolvable software vulnerability detection method, comprising: by regular pattern matching identification Source sentence, based on call graph traversal and data dependence analysis execution function level inter-process slice, build cross-function code context, input the big language model of parameter efficient fine-tuning, output the vulnerability propagation path from Source to Sink;When new vulnerability type needs to be extended, the parameter variation characteristics of old data are extracted by multi-step fine-tuning, and the representative core set is selected by random projection dimension reduction and hybrid distance hierarchical clustering, and the training is played back by mixing new data to alleviate catastrophic forgetting;In the actual use process of tool, the false alarm and the false alarm confirmed by user are collected as feedback signal, the core set is clustered and layered filtered and refined based on perplexity and error prediction analysis, the harmful old knowledge that leads to false alarm and false alarm is removed, and verified new mode is supplemented at the same time, to realize the closed-loop evolution of self-improvement.
Owner:NANJING UNIV

System and method for optimizing content positioning to influence LLM-based ai tools

PCT designated stageWO2026139961A1EngineeringData mining
A system for influencing outputs of a large language model includes at least one memory, at least one processor, a perplexity optimizer, and a corpus handler. The processor executes instructions stored in the memory to operate the optimizer and handler. The perplexity optimizer generates multiple candidate supporting texts based on a target concept. It computes a perplexity metric for each candidate within a context derived from a local corpus, using token likelihoods from a reference language model, and selects a supporting text based on the metric. The corpus handler identifies online editable corpora based on their likelihood of inclusion in language model training data. It then collects contextual information from these corpora to create the local corpus and inserts the selected supporting text into at least one of the identified online corpora.
Owner:WIX COM

Difficulty-aware learning based large language model personalization alignment method and device

PendingCN122114053AAvoid catastrophic forgettingImprove training stabilityBiological modelsPersonalizationLearning based
The application provides a large language model individualization alignment method and device based on difficulty perception learning, and belongs to the technical field of model training. The method comprises the following steps: obtaining multiple groups of training samples; each group of training samples comprises a user question, a user portrait, and a target reply corresponding to the user question under the user portrait; calculating the perplexity of the target reply in each group of training samples; dividing the multiple groups of training samples into high-perplexity training samples and low-perplexity training samples according to a preset perplexity threshold and the perplexity of the target reply in each group of training samples; performing first training on a preset large language model based on the low-perplexity training samples and a likelihood loss function, performing second training on the preset large language model after the first training based on the high-perplexity training samples and reinforcement learning, and obtaining a target large language model, so that the target large language model completes individualization alignment. The application can improve the individualization alignment effect of the model.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Text data processing method and device, equipment and storage medium

The present disclosure provides a text data processing method, device and equipment and storage medium. The present disclosure obtains a plurality of perplexity values corresponding to a plurality of sample texts respectively by performing inference on each of the plurality of sample texts through a plurality of pre-trained models. Further, a confidence interval is determined according to the plurality of perplexity values corresponding to the plurality of sample texts respectively, so as to determine an abnormal sample in the plurality of sample texts according to the confidence interval. After clustering processing of the plurality of sample texts is performed to obtain a plurality of clustering clusters, a clustering cluster containing an abnormal sample in the plurality of clustering clusters is removed, that is, all sample texts contained in the clustering cluster are removed. These sample texts are usually low-quality or irrelevant text data. The remaining sample texts constitute a high-quality data set, thereby improving the accuracy of subsequent processing or calculation processes that rely on the data set.
Owner:BEIJING CENTURY TAL EDUCATION TECH CO LTD

Pancreatic cancer prediction method and system based on local and global confusion weighted pruning

The invention discloses a pancreatic cancer prediction method and system based on local and global perplexity weighted pruning, and relates to the technical field of artificial intelligence natural language processing, and the method comprises the steps: obtaining electronic medical record text data of a to-be-diagnosed patient, and carrying out the data screening processing; according to the method, a structured cue word sequence is constructed, a large language model is guided to generate an initial reasoning path, a step which plays a key role in context coherence can be automatically identified by calculating a local confusion degree variable quantity, and evaluation of the contribution degree of the reasoning step to a final diagnosis conclusion is realized by calculating a global confusion degree variable quantity, so that the accuracy of reasoning is improved. By weighting local and global dual confusion degree variable quantities, a comprehensive importance scoring function is constructed, comprehensive analysis of reasoning steps is realized, redundant steps can be automatically eliminated by setting an adaptive pruning threshold value, an optimized thinking chain with clear logic is finally formed, and a pancreatic cancer risk prediction result and a key reasoning path are output based on the thinking chain. And the diagnosis process is completely clear.
Owner:QINGDAO UNIV

An ai-generated text detection method based on graph structure features

This invention presents an AI-generated text detection method based on graph structure features, belonging to the fields of artificial intelligence and natural language processing. The method includes: dataset construction, entity relation extraction and graph structure construction, graph structure feature extraction, graph structure feature model training, and text detection. It further incorporates traditional text feature extraction and model training, adaptively fusing the traditional text feature model and the graph feature model based on confidence-weighted entropy, and then performing text detection based on the fused model. This invention is the first to perform AI text detection from the perspective of graph structure features, breaking through the limitations of existing research that focuses on surface features such as vocabulary, syntax, and perplexity. The fusion strategy dynamically adjusts the fusion weights by quantifying the uncertainty of model predictions, maintaining a high level of performance on both original data and adversarial examples, achieving a balance between detection accuracy and adversarial robustness. It can be widely applied to the detection of AI-generated content such as news content and academic papers.
Owner:PEKING UNIV +2

RAG optimization method fusing bidirectional confusion blocking and adaptive selection

The invention discloses an RAG optimization method fusing bidirectional confusion degree blocking and adaptive selection, and relates to the technical field of retrieval enhancement generation. According to the segmentation method based on the bidirectional confusion degree, semantic boundaries in long documents are recognized in a self-adaptive mode by analyzing predictability of forward and backward marks, and compared with a traditional one-way PPL method, the method has the advantages that the coherence of language chunks is improved, and retrieval correlation is enhanced; on the basis of a hierarchical index structure, a self-adaptive selection method based on similarity gradient information is provided, and the optimal candidate number to be retrieved in each stage is determined in a self-adaptive mode by analyzing gradient changes in a similarity score sequence.
Owner:CHONGQING UNIV OF TECH

Private domain fine tuning corpus determination method and device

The embodiment of the invention provides a private domain fine tuning corpus determination method and device. The method comprises the following steps: cutting an original text corpus of an enterprise private domain according to a multi-level directory structure to obtain a plurality of target text fragments used for inputting a large language model; acquiring a plurality of private domain fine-tuning corpora through a large language model according to the target text fragment and in combination with a preset guide prompt word; and sorting the plurality of private domain fine-tuning corpora according to the confusion values corresponding to the plurality of private domain fine-tuning corpora, and determining the sorted plurality of private domain fine-tuning corpora as a target private domain fine-tuning corpora for training a private domain model of the enterprise. According to the embodiment of the invention, the problem that the quality of a private domain fine-tuning corpus based on manual calibration is low and the recovery of the question and answer ability of the private domain model is not facilitated in the related technology is solved, and the effect of improving the training accuracy of the private domain model is achieved.
Owner:ZTE CORP

Method and device for detecting intrusion in a computer system

Method and device for detecting intrusion in a computer system. This method uses a large language model (LLM), previously trained by machine learning on training data comprising at least training data representative of a normal state of network communications or of the operation of the computer system, and comprises the steps of: slicing (42) at least a subset of collected formatted data into a sequence of elementary fragments, applying (44) the LLM to the sequence of elementary fragments, providing as output, for each elementary fragment, a likelihood value of the elementary fragment as a function of a context comprising at least some of the other elementary fragments of said sequence of elementary fragments, calculating (46) a perplexity value of said sequence of elementary fragments as a function of said likelihood values,When the calculated perplexity value exceeds a normality threshold, detection (50) of a potential intrusion and issuance of an intrusion alert in said computer system. Figure for the abbreviation: Figure 2,
Owner:COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES

Power industry language large model evaluation method and system considering data pollution influence

The invention discloses an electric power industry language large model evaluation method and system considering data pollution influence, and belongs to the technical field of electric power industry language large model evaluation.The evaluation method comprises the steps that electric power elements in evaluation data are marked, and the confusion degree of the electric power elements in the evaluation data is calculated; comparing the calculated perplexity with a perplexity threshold, and if the perplexity of the power element in the evaluation data is determined, determining that the evaluation data is polluted; taking the meaning of not changing the sentence as a standard, rewriting the polluted evaluation data until the confusion degree of the electric power element in the evaluation data is greater than the confusion degree, and judging that the evaluation data meets the requirement; calculating the weight of the evaluation data meeting the requirements; and taking a weighted average value of the power industry language large model answer score by using the calculated weight to obtain an evaluation score. According to the method, the influence of data pollution on the evaluation result is considered, and the evaluation result can more accurately reflect the real capability of the power industry language large model.
Owner:CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD

Cross-domain feature alignment-based education scene large model revised text traceability detection method

The invention discloses a cross-domain feature alignment education scene large model revised text traceability detection method, and relates to the technical field of intelligent text detection, and the method comprises the specific steps: collecting a job, and generating a revised text by using a large model; multi-dimensional feature extraction: respectively constructing a student literary and sports library and defining machine revision features; performing confusion degree consistency optimization, quantifying man-machine text differences and performing preliminary judgment; performing orthogonal constraint and cross-domain alignment to realize accurate feature alignment; and finally, traceability detection reasoning is performed, a judgment result, a type and probability distribution are output, and a structured report is generated. According to the method, a dynamic data set and multi-dimensional feature extraction are matched with educational text features, man-machine text differences are accurately quantified through technologies such as confusion degree optimization, and the problem of machine revision recognition is solved; and then characteristic expression is enhanced by orthogonal constraint and the like, a structured report is output in combination with traceability reasoning, a learning condition basis is provided for teachers, and deep fusion of detection precision and teaching application is realized.
Owner:CHONGQING INST OF ENG

A method, device, medium and product for screening instruction data

The application relates to the technical field of information technology and discloses a kind of screening method, equipment, medium and product of instruction data.The method comprises the following steps: determining first perplexity and second perplexity according to a pre-trained lightweight general language model and a target sample set;The first perplexity is used to represent the predicted perplexity under the instruction condition;The second perplexity is used to represent the predicted perplexity under the non-instruction condition;According to the first perplexity and the second perplexity, determine the information gain score;Information gain score is used to quantify the contribution of instruction content to reduce the difficulty of response content generation;According to the information gain score and the target sample set, determine the training data set;The training data set is used to improve the generalization ability and instruction compliance accuracy of large language model in the instruction fine-tuning stage.The technical problems of high deployment threshold, high cost, supervision dependence, low training efficiency and poor model generalization ability can be solved.
Owner:SHANGHAI COOPERS TECHNOLOGY CO LTD

A pronunciation evaluation method, device, equipment and storage medium

Embodiments of the present application disclose a pronunciation evaluation method, device and equipment and a storage medium. The method comprises: obtaining to-be-evaluated audio and corresponding reference text, aligning the to-be-evaluated audio and the corresponding reference text through a preset acoustic model to obtain a first test text of the to-be-evaluated audio; merging continuous same letters in the first test text to obtain a second test text, calculating posterior probabilities of each letter in the second test text, and determining pronunciation accuracy of a corresponding letter according to the posterior probabilities; determining a missed letter according to letters in the second test text and letters in the corresponding reference text; deleting or replacing a blank symbol in the second test text with a pause symbol to obtain a third test text, calculating a language model perplexity of the third test text according to a preset pause language model, and determining pronunciation fluency of the to-be-evaluated audio according to the language model perplexity. The above technical means solve the problem of single evaluation dimension of the existing pronunciation evaluation method.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Retrieval method, system and equipment based on semantic confusion and medium

The invention discloses a semantic confusion-based retrieval method, system and device and a medium, and the method comprises the steps: before and after endowing retrieval information, sampling a given question based on a preset large language model, and obtaining a corresponding answer result and a generation probability; all the answer results expressing semantic equivalence are merged into the same semantic class set; calculating the sum of the generation probabilities of all the answer results in the semantic class set, and taking the sum as the semantic confusion of the large language model for the given question; measuring the contribution degree of the retrieval information by analyzing the gain between the determined semantic perplexity before and after endowing the retrieval information with the retrieval information; measuring the contribution degree of the retrieval information by analyzing the gain between the determined semantic perplexity before and after endowing the retrieval information with the retrieval information; and performing optimization training on the large language model according to an enhanced sample data set constructed by the contribution degree to improve the model output quality.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Code retrieval model training method, code completion method and related products

The invention provides a code retrieval model training method, a code completion method and a related product. The training efficiency of a code retrieval model can be improved. The method can be applied to a training device. Specifically, a training device obtains training data, the training data comprises a target code, preceding text information of the target code and a candidate code snippet set, then the preceding text information of the target code is input into a code retrieval model, and at least one candidate code snippet is retrieved from the candidate code snippet set; according to the at least one candidate code snippet, the context information of the target code and the target code, the confusion degree corresponding to the at least one candidate code snippet is obtained, and the confusion degree corresponding to the at least one candidate code snippet is used for representing the confusion degree of the at least one candidate code snippet under the condition that the previous text information of the target code is given; a probability of the target code is generated based on the candidate code snippet. And then, the training device adjusts parameters of the code retrieval model according to the confusion degree corresponding to the at least one candidate code snippet.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD +1

Diffusion covering-based reasoning prediction distributed distillation training method and system

The invention belongs to the technical field of artificial intelligence, and relates to an inference prediction distributed distillation training method and system based on diffusion covering, and the method comprises the steps: S1, obtaining an inference data set; s2, reasoning by the teacher model to obtain reasoning process texts, answers and reasoning prediction distribution sequences; s3, reasoning and checking the teacher model based on the standard answer; s4, segmenting the reasoning process text; s5, aiming at the reasoning step sequence, constructing reasoning input sequences under different perplexity levels; s6, the student model performs reasoning prediction to obtain a reasoning prediction distribution sequence and answer prediction probability distribution; and S7, on the premise that the reasoning steps are aligned, taking the reasoning prediction distribution sequence of each reasoning step of the teacher model as a distillation target, and performing distribution-level distillation training on the student model. The method enables the model to have higher reasoning robustness, stability and generalization ability in a complex reasoning task, and has high engineering practical value and popularization prospect.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

An Automatic Mining and Classification Method for Private Data

This invention discloses an automatic mining and classification method for private data, comprising the following steps: First, receiving private unstructured documents, performing preprocessing and semantic alignment, and using sliding window technology to segment the documents into continuous text blocks; then loading a locally deployed general basic model and an industry-specific model, calculating the perplexity of each text block, and determining the scarcity of the text block by comparing the differences in the model output; next, constructing a statistical model based on the perplexity distribution of the enterprise's historical documents and dynamically setting a threshold, calculating the value score for text blocks exceeding the threshold; finally, projecting the scarcity and value score onto a preset classification decision matrix to automatically determine the secret level of the text block and trigger corresponding security handling strategies. This invention achieves automated and intelligent classification of private data, improving the efficiency and accuracy of data security management.
Owner:JIANGSU DAOYUNYIN TECH CO LTD