Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

409 results about "Inference" patented technology

Inferences are steps in reasoning, moving from premises to logical consequences; etymologically, the word infer means to "carry forward". Inference is theoretically traditionally divided into deduction and induction, a distinction that in Europe dates at least to Aristotle (300s BCE). Deduction is inference deriving logical conclusions from premises known or assumed to be true, with the laws of valid inference being studied in logic. Induction is inference from particular premises to a universal conclusion. A third type of inference is sometimes distinguished, notably by Charles Sanders Peirce, distinguishing abduction from induction, where abduction is inference to the best explanation.

Multi-modal big language model reasoning optimization method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical health, and discloses a reasoning optimization method, device, equipment and medium for a multi-modal large language model.The method comprises the steps that an input long context sequence is obtained, and key value projection is conducted on the long context sequence to generate an initial key value cache; for each attention layer of the multi-modal large language model, calculating an attention matrix of the attention layer according to the vector dimension of the long context sequence and the initial key value cache; calculating a cross-modal attention entropy according to the attention matrix, and determining a cache size of an attention layer according to the cross-modal attention entropy; optimizing the initial key value cache based on a cumulative attention scoring mechanism and a window strategy to obtain a target key value cache; and reasoning the long context sequence according to the cache size and the target key value cache to generate a long context reasoning result. And the reasoning efficiency and the reasoning accuracy are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Scene application operation maintenance platform based on artificial intelligence model

The invention relates to the technical field of artificial intelligence, in particular to a scene application operation maintenance platform based on an artificial intelligence model, which comprises a reasoning state sensing module, a load balancing scheduling module, an intelligent resource allocation module, a deployment configuration adjustment module and a scene feedback integration module. According to the method, by sensing and reasoning process resource consumption and response time migration trend, abnormal tasks caused by resource fluctuation can be accurately recognized, the state monitoring sensitivity is improved, task allocation records are analyzed to construct conflict structures, the resource conflict recognition and visualization capability is enhanced, the model operation level and the GPU vacancy rate are combined, a task migration path is optimized, and the task migration efficiency is improved. Node resource balance is realized, storage frequency, memory margin and network delay are fused during deployment, migration stability is guaranteed, drift tasks are identified through response logs and call records, an operation and maintenance list is generated, operation and maintenance accuracy and multi-scene matching capability are improved, and reasoning task stability and resource scheduling intelligent level are enhanced.
Owner:SHENZHEN YUNHENG INTELLIGENT CO LTD

Inference acceleration optimization method and system applied to intelligent dialogue large model

The invention provides a reasoning acceleration optimization method and system applied to an intelligent dialogue large model, and belongs to the technical field of large models.Firstly, a to-be-reasoned dialogue sequence and reasoning environment configuration information are obtained, and the to-be-reasoned dialogue sequence comprises a user real-time input text and a historical interaction statement chain; the reasoning environment configuration information covers an operation node load state and cache resource occupation information, then joint process deconstruction processing is carried out on the operation node load state and the cache resource occupation information to obtain a reasoning node dependence graph and a resource elasticity demand list, reasoning link optimization processing is carried out based on the result, and a reasoning acceleration execution scheme is generated; the method comprises the steps of reasoning a node parallel scheduling rule and a resource pre-allocation strategy, regulating and controlling a reasoning operation process according to a reasoning acceleration execution scheme, generating a dialogue response sequence after acceleration processing, and finally pushing the dialogue response sequence after acceleration processing to a user interaction terminal to complete intelligent dialogue output. Therefore, the reasoning speed of the intelligent dialogue large model is effectively improved, and the dialogue interaction experience is optimized.
Owner:XINGFAN XINGQI (CHENGDU) TECH CO LTD

Large model intelligent reasoning method combining reinforcement learning and retrieval enhancement generation

The invention provides a large model intelligent reasoning method combining reinforcement learning and retrieval enhancement generation, and belongs to the technical field of artificial intelligence. Comprising the following steps: data input: data preprocessing; constructing a reinforcement learning environment; an RAG mechanism is integrated; training the model; and evaluating and iteratively optimizing. A reinforcement learning framework based on rules is introduced to guide a model to develop advanced reasoning skills such as reflection, verification and summarization. And in combination with a retrieval enhancement generation mechanism, the model can access and utilize a wide background knowledge base before answering questions. The synergistic effect between information retrieval and text generation is optimized. According to the method, the deficiency of knowledge of the model can be made up by introducing the external knowledge base, and the synergistic effect between information retrieval and text generation can be optimized in the reinforcement learning process, so that the capability of the model for processing complex reasoning tasks is remarkably improved, and formation of a more generalization reasoning strategy is promoted.
Owner:GUANGDONG UNIV OF TECH

Method and device for training reasoning model and method and device for processing reasoning problem

The invention provides a training method and device of an inference model and a processing method and device of an inference problem, and relates to the field of artificial intelligence, in particular to the technical field of natural language processing, large models and deep learning. Obtaining a training sample set, wherein training samples in the training sample set at least comprise sample reasoning questions, reference reasoning chains corresponding to the sample reasoning questions and reference answers; on the basis of a first training sample in the training sample set, performing supervised fine tuning SFT on a pre-training model to obtain a first inference model after fine tuning; and performing reinforcement learning RL training on the first reasoning model based on the second training sample in the training sample set to obtain a second reasoning model, thereby improving the accuracy and stability of the second reasoning model in the reasoning process.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Resource pre-allocation method and device for computing device cluster and electronic device

Embodiments of the invention provide a resource pre-allocation method and apparatus for a computing device cluster, and an electronic device. The method comprises the steps of obtaining task information of a reasoning task expected to be submitted to a reasoning model for reasoning; dividing the reasoning tasks into a plurality of types according to the input token number and the output token number of the reasoning tasks, and determining an expected concurrency number of each type of reasoning tasks; for each type of reasoning task, determining a first target model in the equipment models of the computing equipment included in the computing equipment cluster according to the number of input tokens and the number of output tokens of the type of reasoning task; and for each type of reasoning task, according to the expected concurrence number, the input token number and the output token number of the type of reasoning task, reserving a first reasoning instance in a computing device of a first target model determined for the type of reasoning task. By applying the embodiment of the invention, the overall reasoning efficiency of the reasoning model can be improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Model training and information replying method and device, storage medium and program product

The invention provides a model training and information replying method and device, a storage medium and a program product, and relates to the technical field of computers. The method comprises the steps of performing continuous pre-training on a base model based on a first training sample to obtain a basic model; performing cold start supervision fine tuning training on the basic model based on the second training sample to obtain a first supervision fine tuning model; performing multiple reasoning based on the third training sample, the target information and the to-be-trained model to obtain a reasoning result; in the Mth reasoning process, the target information comprises information obtained after a target tool determined by previous M-1 reasoning is called; optimizing the to-be-trained model based on the reasoning result to obtain a first reinforcement learning model; and based on the general recognition data and reasoning data output by the first reinforcement learning model, carrying out general recognition alignment training to obtain a target model. According to the method, the target tool can be called to obtain the required target information, so that the information output by the large model is more comprehensive.
Owner:RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD

Model training method and device, computer equipment, readable storage medium and program product

The invention relates to a model training method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: reasoning a sample problem through a reasoning model to obtain a reasoning result; determining a corresponding reasoning length control hyper-parameter according to the difficulty level category of the sample problem; constructing a reasoning length reward function according to the reasoning length control hyper-parameter and the reasoning length of the reasoning result, and constructing a reasoning accuracy reward function according to the reasoning result; and performing model training based on reinforcement learning on the reasoning model according to the reasoning length reward function and the reasoning accuracy reward function. The reasoning model trained by the method can give consideration to reasoning efficiency and accuracy, and a more efficient and accurate reasoning process can be realized.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method for realizing expansion and contraction of inference service instance, electronic equipment and storage medium

The invention provides an inference service instance expansion and contraction method, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: predicting a future business load based on historical operation data of a target inference service to generate an active expansion and contraction instance decision; evaluating the current operation state based on the real-time operation data of the target inference service to generate a passive scaling instance decision; and performing collaborative decision-making on the active expansion and contraction instance decision and the passive expansion and contraction instance decision to determine a final expansion and contraction instance instruction, and adjusting the instance number of the target inference service according to the final expansion and contraction instance instruction. According to the method, a double-engine cooperation mechanism combining active prediction and passive response is established, while prospective capacity expansion and contraction are realized by using historical data to reduce time delay, bottom correction is carried out by using real-time data to cope with burst load, the problem of response lag or resource waste of a single capacity expansion and contraction mode is effectively solved, and the method is suitable for large-scale popularization and application. And the resource utilization rate and the service quality stability of the inference service are obviously improved.
Owner:IFLYTEK CO LTD

Efficient and accurate regional explanation technique for NLP models

Herein are techniques for topic modeling and content perturbation that provide machine learning (ML) explainability (MLX) for natural language processing (NLP). A computer hosts an ML model that infers an original inference for each of many text documents that contain many distinct terms. To each text document (TD) is assigned, based on terms in the TD, a topic that contains a subset of the distinct terms. In a perturbed copy of each TD, a perturbed subset of the distinct terms is replaced. For the perturbed copy of each TD, the ML model infers a perturbed inference. For TDs of a topic, the computer detects that a difference between original inferences of the TDs of the topic and perturbed inferences of the TDs of the topic exceeds a threshold. Based on terms in the TDs of the topic, the topic is replaced with multiple, finer-grained new topics. After sufficient topic modeling, a regional explanation of the ML model is generated.
Owner:ORACLE INT CORP

Network security situation awareness method based on artificial intelligence

The invention discloses a network security situation awareness method based on artificial intelligence, and the method comprises the following steps: collecting multi-source heterogeneous data, and generating a standardized data set; spatial-temporal feature decoupling is carried out, and spatial-temporal dimension features are separated; fusing the time-space cross attention, and outputting a fused time-space feature vector; constructing a causal inference engine, and outputting a dynamic causal graph and an anti-fact inference result set; constructing a dynamic risk propagation model, and outputting a whole asset risk value matrix and a risk propagation path diagram; generating a situation quantization matrix, constructing an adversarial training decision network, and outputting a defense strategy set verified by adversarial training; automatically generating a strategy; and a man-machine cooperative verification closed loop is realized. According to the method, dynamic reconstruction of a threat propagation path is realized through spatial-temporal feature decoupling and a causal reasoning engine, and a risk positioning error is reduced; and the adversarial training decision network is combined, so that the misjudgment rate of the defense strategy in the simulation APT attack test is reduced.
Owner:BEIJING BEILONG YUNHAI NETWORK DATA TECH CO LTD

Refractory case question and answer sample acquisition method, model training method and related equipment

The invention provides a difficult case question and answer sample acquisition method, a model training method and related equipment. The difficult case question and answer sample obtaining method comprises the steps that a to-be-corrected question and answer sample is obtained, the to-be-corrected question and answer sample comprises a preset question, an annotated answer and a first reasoning link comprising a reasoning answer, and the reasoning answer included in the first reasoning link does not conform to the annotated answer; the to-be-corrected question and answer sample is input into a second language model, so that the second language model outputs first reflection content, and the first reflection content comprises an error point in the first reasoning link and a correction thought for the error point; inputting the preset question, the first reasoning link and the first reflection content into a second language model, so that the second language model outputs a second reasoning link including the reasoning answer; and if the inference answer included in the second inference link is consistent with the labeled answer, generating a difficult case question and answer sample based on the preset question, the labeled answer, the first inference link, the first reflection content and the second inference link.
Owner:ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD

Inference method and related device

The present application relates to the field of artificial intelligence, and in particular to an inference method and a related device. The method comprises: a first computing device acquiring an inference request statement of a user, and splitting the inference request statement so as to obtain a plurality of tokens; on the basis of a mapping table of parts of speech and devices, dividing the plurality of tokens into K groups of tokens, and determining the correspondence between the K groups of tokens and K second computing devices, wherein in the mapping table of parts of speech and devices, the parts of speech corresponding each of the K second computing devices comprise parts of speech of tokens in a token group corresponding to said second computing device; the first computing device respectively sending the K groups of tokens to the K second computing devices, such that each second computing device processes the received tokens on the basis of deployed experts, so as to obtain an inference response statement for responding to the inference request statement; and the first computing device receiving K inference response statements respectively sent by the K second computing devices. The use of the solution of the present application facilitates reduction of communication overhead.
Owner:HUAWEI TECH CO LTD

Multi-modal inference method and inference system based on error attribution

The invention belongs to the technical field of thinking chain reasoning, and particularly relates to a multi-modal reasoning method and system based on error attribution. The reasoning method comprises the following steps: on the basis of a current modal fusion weight, performing modal fusion on each piece of initial information in an initial information set, and then generating a thinking chain; after a reasoning dependency graph is constructed based on the thinking chain, check points are selected in the reasoning dependency graph; based on consistency, factuality and logicality, performing error possibility scoring on each check point, and if the error possibility scores of all check points in the current thinking chain are below a set threshold, outputting the current thinking chain; otherwise, marking the check points of which the error possibility scores exceed a set threshold value as error nodes; calculating relative contribution strength of different modes to error nodes; and on the basis of the relative contribution strength, updating the modal fusion weight, and regenerating the thinking chain. According to the invention, the accuracy of the reasoning result and the stability of the accuracy can be improved.
Owner:DATA SPACE RES INST

Large language model safety protection defense method and device based on dynamic regulation and control

The invention provides a large language model safety protection defense method and device based on dynamic regulation and control, and belongs to the field of artificial intelligence safety protection. The large language model security protection defense method comprises the following steps: constructing a non-security data set and a security data set jail break prompt data set, calculating a gradient average value of parameters of each layer of a large language model during back propagation of various data sets, calculating cosine similarity among gradients, determining a non-security layer, namely a layer most sensitive to non-security content, and performing security protection defense on the non-security layer. Therefore, the subsequent regulation and control are more accurate and effective; in the aspect of dynamic regulation and control of a specified non-security layer, a joint loss function is established to optimize and train a security offset vector, and the security offset vector is applied to a hidden state of the non-security layer in a big language model reasoning process to carry out intervention, so that output of a big language model is dynamically regulated and controlled; therefore, the robustness of the large language model is improved.
Owner:ZHEJIANG UNIV

Vehicle-mounted logistics intelligent control system based on artificial intelligence

The invention discloses a vehicle-mounted logistics intelligent control system based on artificial intelligence, and the system comprises the following steps: a state data collection module which is used for collecting information and preprocessing the information to form current state data; the causal inference module is used for generating causal inference data; the intelligent scheduling module is used for generating a vehicle-order matching scheme and a path planning scheme and forming an initial scheduling decision; the execution result acquisition module is used for issuing an initial scheduling decision and obtaining a scheduling execution result; the anti-fact inference module is used for inferring an anti-fact result when the initial scheduling decision is not executed; the causal contribution calculation module is used for determining causal contributions of the vehicles and the road sections; the causal credit updating module is used for updating the causal credit of the vehicle and the road section to form a causal credit account book; and the scheduling correction module is used for adjusting the scheduling output of the next period and updating. According to the invention, dynamic correction of the vehicle scheduling process is realized through causal inference and a causal credit account book.
Owner:CSCU (LIAONING) COMPUTER INTEGRATION CO LTD

Big language model-based reasoning service processing method, apparatus and device, and medium

The invention provides an inference service processing method and device based on a large language model, equipment and a medium, and relates to the technical field of artificial intelligence such as large language models and server-free architectures. The method comprises the steps of determining an actual demand number of model instances based on an inference service request received in a preset period; in response to the fact that the existing number of the loaded model instances is smaller than the actual demand number, through a first scheduling interface generated based on the CRD technology, selecting a target idle node from an idle node set which completes configuration of the service operation environment in advance; through a second scheduling interface generated based on the CRD technology, controlling model weight data pre-stored in a memory in the target idle node to generate new loading model instances which correspond to the missing number and run in the service running environment; and processing the received inference service request by using the newly loaded model instance and the loaded model instance. According to the scheme, thorough decoupling between the model service environment and the model ontology is realized.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Inference method and apparatus for large language model

An inference method and apparatus for a large language model, which method and apparatus are used for improving the inference calculation efficiency of a large language model. The method comprises: receiving an inference task, which carries a morpheme number, wherein the morpheme number is used for identifying a morpheme to be subjected to inference calculation, the morpheme comprising one or more sub-words determined on the basis of input text of a large language model (201); on the basis of the morpheme number corresponding to the inference task, querying a global index tree to determine a target inference node, wherein the global index tree comprises sub-trees corresponding to a plurality of inference nodes in a distributed inference cluster, and a sub-tree corresponding to the target inference node is a sub-tree in the global index tree, the matching degree of similarity of which sub-tree with the morpheme number is greater than a threshold, the matching degree of similarity indicating the number of pieces of reusable key-value cache data during inference calculation (202); and on the basis of the target inference node, executing the inference task (203).
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Visual language model target detection capability enhancement method

The invention discloses a visual language model target detection capability enhancing method, which comprises the following steps of: firstly, constructing an inference type target detection data set containing complex semantic tags such as attributes, interaction, orientation, negative and hard negative samples; and secondly, under a GRPO reinforcement learning framework, guiding the VLM to generate a reasoning process through a specific cue word, and then outputting a detection result. The present invention employs a composite reward function to evaluate a plurality of candidate outputs generated by the model, the function including a format reward that ensures that the output follows a preset thinking and answer structure, and an innovative ODLength reward. According to the ODLength reward, the average precision mean value is combined with a length penalty term, and redundant prediction is effectively restrained. And finally, updating the model strategy network according to the total reward value. According to the method, the target detection precision and generalization ability of the VLM in a complex reasoning scene can be remarkably improved, and the reasoning efficiency is improved.
Owner:HONGLONG TECH (HANGZHOU) CO LTD

Large language model output quality guarantee method, device and system based on chain deterministic reasoning verification

The invention discloses a large language model output quality guarantee method, device and system based on chain type deterministic reasoning verification, and belongs to the field of artificial intelligence safety, deterministic calculation and reasoning verification. The method comprises the following steps: classifying each reasoning step generated by LLM into a deterministic step (bT: deductive reasoning, verified facts and mathematical proof) or a non-deterministic step (bF: inductive reasoning, unverified references and creative guess), and accumulatively constructing a reasoning chain; chain certainty verification chain (chain) = foldr (AND, true, chain) and time complexity O (n) are carried out in each k steps; if the bF is returned, a chain pollution elimination theorem is applied to prove that once the chain contains non-deterministic steps, the whole chain loses deterministic guarantee; and a deterministic boundary report is generated, and a non-deterministic initial position is marked. The device comprises an inference classifier, a reference verifier, a mathematical checker and a chain verification engine, and all the modules can be independently realized (formalized verification / database query / rule engine / mixed mode).
Owner:GUANGZHOU KINGPIN IND CO LTD

Large model logical reasoning optimization method and system combined with knowledge graph

The invention provides a large model logical reasoning optimization method and system combined with a knowledge graph, and relates to the technical field of artificial intelligence, first, a logical reasoning task to be processed of a large model and a corresponding structured knowledge graph are obtained, and the logical reasoning task comprises knowledge units, association relationships and association strength; the logical reasoning task comprises a reasoning target, a premise and a constraint condition, then obtaining a large model reasoning guide rule set based on suitability analysis, injecting the large model reasoning guide rule set into a reasoning process, capturing and adjusting intermediate reasoning nodes in stages, correcting reasoning branches deviating from an incidence relation, and obtaining a large model reasoning guide rule set; then conflict resolution and propagation prediction are performed on the reasoning process after stage guidance, an inconsistent conclusion is identified, conflict rationality is verified, a propagation link is predicted, a multi-round correction scheme is generated, and finally reasoning path iterative optimization, integrating degree and efficiency evaluation and jump sequence and reference priority adjustment are performed based on the reasoning process after conflict resolution. Information is integrated to obtain an optimization result, and the logical reasoning quality of the large model is effectively improved.
Owner:XINGFAN XINGQI (CHENGDU) TECH CO LTD

Cascade speculation inference method and system based on hierarchical decline KV cache compression

The invention discloses a cascade speculation inference method and system based on hierarchical decline KV cache compression, and the method comprises the steps: firstly inputting a context prompt text into a target model for coding, generating a KV cache, and calculating an attention score between tokens; secondly, descending sorting is carried out based on the attention scores of the final input token, KV cache blocks corresponding to the first k attention scores are selected as cascading middle layer cache, a lightweight large language model is loaded to serve as a draft model, and a hierarchical decline KV cache compression strategy is adopted to maintain draft model cache. Then, based on all the caches, a double-layer cascade speculation reasoning framework is constructed, a target reasoning path is obtained, and the caches are updated; and finally, repeating the operation until target response data corresponding to the context prompt text is output according to the target reasoning path. According to the method, the KV cache proportion is reduced, meanwhile, the draft token acceptance rate of the target model of the full KV cache is improved, and reduction of precision is reduced.
Owner:HANGZHOU DIANZI UNIV

Emotion analysis method and system based on big language model reasoning chain generation

The invention discloses an emotion analysis method and system based on big language model reasoning chain generation, and the method comprises the steps: obtaining text data, classifying the features of the text data, and determining a graphic reasoning template; reasoning chain generation is carried out step by step through the reasoning template, and risk prediction is carried out on the generated reasoning chain so as to determine a generation strategy of candidate words in the reasoning chain, so that a preliminary reasoning chain result is obtained; and on this basis, calculating the aggregation support degree, finally adjusting the generation strategy again to obtain the inference result of the next step, and obtaining the final inference chain according to the graphical inference template through the cyclic adjustment. Therefore, according to the finally obtained reasoning chain, the illusion situation in the model operation process is greatly reduced, the credibility of the reasoning result is improved, and the stability and accuracy of the generated result are improved through multi-time interactive calculation of the reasoning chain and text data.
Owner:湖南工商大学

Managing inference models in view of reconstructability of sensitive information

Methods, systems, and devices for providing computer-implemented services are disclosed. To provide the computer-implemented services, inference models may be deployed to locations. Prior to deploying an inference model to a location, it may be determined whether the location is trustworthy. If the location is determined to not be trustworthy, an input data attack resistant inference model may be selected and deployed. The input data attack resistant inference model may be trained to impart reconstruction resistance to sensitive information during inference generation based on a schema for weighting the sensitive information for reconstruction resistance. The training process may decrease a likelihood of the inferences generated by the input data attack resistant inference model being usable to reconstruct the sensitive information.
Owner:DELL PROD LP

Dynamic scheduling reasoning method based on hybrid expert model and related equipment

The invention discloses a dynamic scheduling reasoning method based on a hybrid expert model and related equipment, and the method comprises the steps: determining a to-be-reasoned hybrid expert model which comprises a plurality of expert sub-models; performing nested weight quantization processing on the plurality of expert sub-models to obtain a quantized sub-model set; in response to an inference task, performing routing activation processing on the quantized sub-model set to obtain an activated sub-model set; performing dynamic bit width selection processing on the activation sub-model set according to the reasoning task to obtain an initial sub-model set; performing bit width sensing reordering processing on the initial sub-model set to obtain a target sub-model queue; and performing execution processing on the reasoning task according to the target sub-model queue to obtain a reasoning result. The embodiment of the invention can improve the efficiency of model reasoning, and can be widely applied to the technical field of artificial intelligence.
Owner:THE HONG KONG UNIV OF SCI & TECH +1

Large language model illusion detection method based on causal double chains and intelligent question and answer method

The invention relates to the technical field of natural language processing, and provides a causal double-chain-based large language model illusion detection method and an intelligent question answering method.The causal double-chain-based large language model illusion detection method comprises the steps that for each input question, multiple rounds of independent sampling are initiated, and a response set containing multiple sampling responses is formed; and for the response set, extracting an inference chain and a causal chain, calculating an inference chain conditional entropy, a causal chain conditional entropy and mutual information of the inference chain and the causal chain, and taking a weighted sum of the inference chain conditional entropy, the causal chain conditional entropy and the mutual information of the inference chain and the causal chain as an illusion score. And multi-dimensional accurate recognition of the composite illusion is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +3

Managing inference model training on an expanded knowledge base

Methods and systems for providing computer-implemented services using inference models are disclosed. To provide the computer-implemented services, supplemental training data may be obtained, the supplemental training data being usable to train a prototype inference model and the prototype inference model being based on an existing inference model. Performance of a training procedure may be initiated using at least the supplemental training data and a set of prompts based on the supplemental training data until performance criteria are met. If the performance criteria are met, the prototype inference model may be promoted to a production ready inference model and used to provide the computer-implemented services.
Owner:DELL PROD LP

Model performance monitoring for UE-based ai / ML positioning

A positioning model monitoring technique includes operations of obtaining an inference result output by an artificial intelligence / machine learning (AI / ML) model for determining a position of a user equipment (UE); obtaining a ground truth label for model monitoring corresponding to the inference result; and performing a monitoring function by comparing the inference result to the ground truth label.
Owner:APPLE INC +1

Inference model compression method and apparatus

Disclosed are an inference model compression method and apparatus, which relate to the technical field of computers. The method comprises: inputting inference data into an inference model, and executing the inference model; then, for a first network layer in the inference model, acquiring a weight matrix of the first network layer and input data of the first network layer, and calculating the product of the weight matrix of the first network layer and the input data of the first network layer to obtain a result matrix that indicates the importance of weight values in the weight matrix; and next, in descending order of the importance, selecting the weight values that rank in the last a% in the weight matrix of the first network layer to perform a sparsification operation, and executing the sparsification operation on a plurality of network layers in the inference model to obtain a sparsified inference model. The first network layer is any one layer among the plurality of network layers, the result matrix and the weight matrix have the same size, and the magnitude of the value at the same position in the result matrix as in the weight matrix indicates the importance of the weight value at the same position in the weight matrix.
Owner:HUAWEI TECH CO LTD

Managing untraining of inference models based on undesirable training data

Methods and systems for providing computer-implemented services using inference models are disclosed. To provide the computer-implemented services, an inference model may be untrained with respect to undesirable training data to obtain an updated inference model. If any other portions of the training data have embeddings similar to embeddings of the undesirable training data, a first testing process may be performed to determine whether the updated inference model provides consistent and accurate responses based on the other portions of the training data with the similar embeddings. If the inference model does not provide the consistent and accurate responses, the inference model may be re-trained to increase a likelihood that a re-trained updated inference model provides the consistent and accurate responses. If the re-trained updated inference model provides the consistent and accurate responses, the re-trained updated inference model may be a compliant inference model.
Owner:DELL PROD LP