Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

310 results about "Inference" patented technology

Inferences are steps in reasoning, moving from premises to logical consequences; etymologically, the word infer means to "carry forward". Inference is theoretically traditionally divided into deduction and induction, a distinction that in Europe dates at least to Aristotle (300s BCE). Deduction is inference deriving logical conclusions from premises known or assumed to be true, with the laws of valid inference being studied in logic. Induction is inference from particular premises to a universal conclusion. A third type of inference is sometimes distinguished, notably by Charles Sanders Peirce, distinguishing abduction from induction, where abduction is inference to the best explanation.

Multi-modal big language model reasoning optimization method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical health, and discloses a reasoning optimization method, device, equipment and medium for a multi-modal large language model.The method comprises the steps that an input long context sequence is obtained, and key value projection is conducted on the long context sequence to generate an initial key value cache; for each attention layer of the multi-modal large language model, calculating an attention matrix of the attention layer according to the vector dimension of the long context sequence and the initial key value cache; calculating a cross-modal attention entropy according to the attention matrix, and determining a cache size of an attention layer according to the cross-modal attention entropy; optimizing the initial key value cache based on a cumulative attention scoring mechanism and a window strategy to obtain a target key value cache; and reasoning the long context sequence according to the cache size and the target key value cache to generate a long context reasoning result. And the reasoning efficiency and the reasoning accuracy are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Large model intelligent reasoning method combining reinforcement learning and retrieval enhancement generation

The invention provides a large model intelligent reasoning method combining reinforcement learning and retrieval enhancement generation, and belongs to the technical field of artificial intelligence. Comprising the following steps: data input: data preprocessing; constructing a reinforcement learning environment; an RAG mechanism is integrated; training the model; and evaluating and iteratively optimizing. A reinforcement learning framework based on rules is introduced to guide a model to develop advanced reasoning skills such as reflection, verification and summarization. And in combination with a retrieval enhancement generation mechanism, the model can access and utilize a wide background knowledge base before answering questions. The synergistic effect between information retrieval and text generation is optimized. According to the method, the deficiency of knowledge of the model can be made up by introducing the external knowledge base, and the synergistic effect between information retrieval and text generation can be optimized in the reinforcement learning process, so that the capability of the model for processing complex reasoning tasks is remarkably improved, and formation of a more generalization reasoning strategy is promoted.
Owner:GUANGDONG UNIV OF TECH

Resource pre-allocation method and device for computing device cluster and electronic device

Embodiments of the invention provide a resource pre-allocation method and apparatus for a computing device cluster, and an electronic device. The method comprises the steps of obtaining task information of a reasoning task expected to be submitted to a reasoning model for reasoning; dividing the reasoning tasks into a plurality of types according to the input token number and the output token number of the reasoning tasks, and determining an expected concurrency number of each type of reasoning tasks; for each type of reasoning task, determining a first target model in the equipment models of the computing equipment included in the computing equipment cluster according to the number of input tokens and the number of output tokens of the type of reasoning task; and for each type of reasoning task, according to the expected concurrence number, the input token number and the output token number of the type of reasoning task, reserving a first reasoning instance in a computing device of a first target model determined for the type of reasoning task. By applying the embodiment of the invention, the overall reasoning efficiency of the reasoning model can be improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Method for realizing expansion and contraction of inference service instance, electronic equipment and storage medium

The invention provides an inference service instance expansion and contraction method, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: predicting a future business load based on historical operation data of a target inference service to generate an active expansion and contraction instance decision; evaluating the current operation state based on the real-time operation data of the target inference service to generate a passive scaling instance decision; and performing collaborative decision-making on the active expansion and contraction instance decision and the passive expansion and contraction instance decision to determine a final expansion and contraction instance instruction, and adjusting the instance number of the target inference service according to the final expansion and contraction instance instruction. According to the method, a double-engine cooperation mechanism combining active prediction and passive response is established, while prospective capacity expansion and contraction are realized by using historical data to reduce time delay, bottom correction is carried out by using real-time data to cope with burst load, the problem of response lag or resource waste of a single capacity expansion and contraction mode is effectively solved, and the method is suitable for large-scale popularization and application. And the resource utilization rate and the service quality stability of the inference service are obviously improved.
Owner:IFLYTEK CO LTD

Refractory case question and answer sample acquisition method, model training method and related equipment

The invention provides a difficult case question and answer sample acquisition method, a model training method and related equipment. The difficult case question and answer sample obtaining method comprises the steps that a to-be-corrected question and answer sample is obtained, the to-be-corrected question and answer sample comprises a preset question, an annotated answer and a first reasoning link comprising a reasoning answer, and the reasoning answer included in the first reasoning link does not conform to the annotated answer; the to-be-corrected question and answer sample is input into a second language model, so that the second language model outputs first reflection content, and the first reflection content comprises an error point in the first reasoning link and a correction thought for the error point; inputting the preset question, the first reasoning link and the first reflection content into a second language model, so that the second language model outputs a second reasoning link including the reasoning answer; and if the inference answer included in the second inference link is consistent with the labeled answer, generating a difficult case question and answer sample based on the preset question, the labeled answer, the first inference link, the first reflection content and the second inference link.
Owner:ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD

Large language model safety protection defense method and device based on dynamic regulation and control

The invention provides a large language model safety protection defense method and device based on dynamic regulation and control, and belongs to the field of artificial intelligence safety protection. The large language model security protection defense method comprises the following steps: constructing a non-security data set and a security data set jail break prompt data set, calculating a gradient average value of parameters of each layer of a large language model during back propagation of various data sets, calculating cosine similarity among gradients, determining a non-security layer, namely a layer most sensitive to non-security content, and performing security protection defense on the non-security layer. Therefore, the subsequent regulation and control are more accurate and effective; in the aspect of dynamic regulation and control of a specified non-security layer, a joint loss function is established to optimize and train a security offset vector, and the security offset vector is applied to a hidden state of the non-security layer in a big language model reasoning process to carry out intervention, so that output of a big language model is dynamically regulated and controlled; therefore, the robustness of the large language model is improved.
Owner:ZHEJIANG UNIV

Vehicle-mounted logistics intelligent control system based on artificial intelligence

The invention discloses a vehicle-mounted logistics intelligent control system based on artificial intelligence, and the system comprises the following steps: a state data collection module which is used for collecting information and preprocessing the information to form current state data; the causal inference module is used for generating causal inference data; the intelligent scheduling module is used for generating a vehicle-order matching scheme and a path planning scheme and forming an initial scheduling decision; the execution result acquisition module is used for issuing an initial scheduling decision and obtaining a scheduling execution result; the anti-fact inference module is used for inferring an anti-fact result when the initial scheduling decision is not executed; the causal contribution calculation module is used for determining causal contributions of the vehicles and the road sections; the causal credit updating module is used for updating the causal credit of the vehicle and the road section to form a causal credit account book; and the scheduling correction module is used for adjusting the scheduling output of the next period and updating. According to the invention, dynamic correction of the vehicle scheduling process is realized through causal inference and a causal credit account book.
Owner:CSCU (LIAONING) COMPUTER INTEGRATION CO LTD

Large language model output quality guarantee method, device and system based on chain deterministic reasoning verification

The invention discloses a large language model output quality guarantee method, device and system based on chain type deterministic reasoning verification, and belongs to the field of artificial intelligence safety, deterministic calculation and reasoning verification. The method comprises the following steps: classifying each reasoning step generated by LLM into a deterministic step (bT: deductive reasoning, verified facts and mathematical proof) or a non-deterministic step (bF: inductive reasoning, unverified references and creative guess), and accumulatively constructing a reasoning chain; chain certainty verification chain (chain) = foldr (AND, true, chain) and time complexity O (n) are carried out in each k steps; if the bF is returned, a chain pollution elimination theorem is applied to prove that once the chain contains non-deterministic steps, the whole chain loses deterministic guarantee; and a deterministic boundary report is generated, and a non-deterministic initial position is marked. The device comprises an inference classifier, a reference verifier, a mathematical checker and a chain verification engine, and all the modules can be independently realized (formalized verification / database query / rule engine / mixed mode).
Owner:GUANGZHOU KINGPIN IND CO LTD

Large model logical reasoning optimization method and system combined with knowledge graph

The invention provides a large model logical reasoning optimization method and system combined with a knowledge graph, and relates to the technical field of artificial intelligence, first, a logical reasoning task to be processed of a large model and a corresponding structured knowledge graph are obtained, and the logical reasoning task comprises knowledge units, association relationships and association strength; the logical reasoning task comprises a reasoning target, a premise and a constraint condition, then obtaining a large model reasoning guide rule set based on suitability analysis, injecting the large model reasoning guide rule set into a reasoning process, capturing and adjusting intermediate reasoning nodes in stages, correcting reasoning branches deviating from an incidence relation, and obtaining a large model reasoning guide rule set; then conflict resolution and propagation prediction are performed on the reasoning process after stage guidance, an inconsistent conclusion is identified, conflict rationality is verified, a propagation link is predicted, a multi-round correction scheme is generated, and finally reasoning path iterative optimization, integrating degree and efficiency evaluation and jump sequence and reference priority adjustment are performed based on the reasoning process after conflict resolution. Information is integrated to obtain an optimization result, and the logical reasoning quality of the large model is effectively improved.
Owner:XINGFAN XINGQI (CHENGDU) TECH CO LTD

Emotion analysis method and system based on big language model reasoning chain generation

The invention discloses an emotion analysis method and system based on big language model reasoning chain generation, and the method comprises the steps: obtaining text data, classifying the features of the text data, and determining a graphic reasoning template; reasoning chain generation is carried out step by step through the reasoning template, and risk prediction is carried out on the generated reasoning chain so as to determine a generation strategy of candidate words in the reasoning chain, so that a preliminary reasoning chain result is obtained; and on this basis, calculating the aggregation support degree, finally adjusting the generation strategy again to obtain the inference result of the next step, and obtaining the final inference chain according to the graphical inference template through the cyclic adjustment. Therefore, according to the finally obtained reasoning chain, the illusion situation in the model operation process is greatly reduced, the credibility of the reasoning result is improved, and the stability and accuracy of the generated result are improved through multi-time interactive calculation of the reasoning chain and text data.
Owner:湖南工商大学

Dynamic scheduling reasoning method based on hybrid expert model and related equipment

The invention discloses a dynamic scheduling reasoning method based on a hybrid expert model and related equipment, and the method comprises the steps: determining a to-be-reasoned hybrid expert model which comprises a plurality of expert sub-models; performing nested weight quantization processing on the plurality of expert sub-models to obtain a quantized sub-model set; in response to an inference task, performing routing activation processing on the quantized sub-model set to obtain an activated sub-model set; performing dynamic bit width selection processing on the activation sub-model set according to the reasoning task to obtain an initial sub-model set; performing bit width sensing reordering processing on the initial sub-model set to obtain a target sub-model queue; and performing execution processing on the reasoning task according to the target sub-model queue to obtain a reasoning result. The embodiment of the invention can improve the efficiency of model reasoning, and can be widely applied to the technical field of artificial intelligence.
Owner:THE HONG KONG UNIV OF SCI & TECH +1

Managing inference model training on an expanded knowledge base

Methods and systems for providing computer-implemented services using inference models are disclosed. To provide the computer-implemented services, supplemental training data may be obtained, the supplemental training data being usable to train a prototype inference model and the prototype inference model being based on an existing inference model. Performance of a training procedure may be initiated using at least the supplemental training data and a set of prompts based on the supplemental training data until performance criteria are met. If the performance criteria are met, the prototype inference model may be promoted to a production ready inference model and used to provide the computer-implemented services.
Owner:DELL PROD LP

Managing untraining of inference models based on undesirable training data

Methods and systems for providing computer-implemented services using inference models are disclosed. To provide the computer-implemented services, an inference model may be untrained with respect to undesirable training data to obtain an updated inference model. If any other portions of the training data have embeddings similar to embeddings of the undesirable training data, a first testing process may be performed to determine whether the updated inference model provides consistent and accurate responses based on the other portions of the training data with the similar embeddings. If the inference model does not provide the consistent and accurate responses, the inference model may be re-trained to increase a likelihood that a re-trained updated inference model provides the consistent and accurate responses. If the re-trained updated inference model provides the consistent and accurate responses, the re-trained updated inference model may be a compliant inference model.
Owner:DELL PROD LP

Processing method and system for rejecting and hanging work order data

PendingCN121766909AThe classification result is accurateSemantic analysisBiological modelsMulti-label classificationQuality data
The invention discloses a processing method and system for rejecting and hanging work order data, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining the rejecting and hanging work order data and a real label corresponding to the rejecting and hanging work order data; inputting the rejected work order data into a trained basic text model to obtain a reasoning result, the reasoning result comprising a prediction label and a confidence coefficient; screening out the rejected work order data of which the confidence coefficient is smaller than a preset value and the predicted tag is inconsistent with the real tag from the reasoning result as low-quality data; determining a quality problem type of the low-quality data based on a quality problem determination rule; performing iterative optimization on the basic text model by adopting a corresponding optimization strategy based on the quality problem type to obtain a multi-label classification model; and inputting the to-be-improved work order rejecting and hanging data into the multi-label classification model to obtain an optimal reasoning result, thereby facilitating solving the problem that the reasoning result of the work order rejecting and hanging data cannot be accurately obtained in the prior art.
Owner:CAPINFO CO LTD

Multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval

A multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval belongs to the field of natural language processing, and comprises the following steps: deconstructing a multi-hop reasoning process into a target-oriented sequence decision problem, carrying out dynamic reasoning guidance by using a large language model, generating a sub-problem sequence matched with a reasoning progress in real time, and carrying out multi-level self-feedback retrieval on the sub-problem sequence; target document retrieval is guided, and sub-questions are dynamically generated; according to the generated sub-questions, obtaining associated documents by adopting a three-level collaborative retrieval mechanism; and performing information refining on the associated document through a large language model, fusing the refined information into an inference chain, and performing inference to generate an answer. The invention further discloses a multi-hop reasoning system, a storage medium and a computer program product. The method aims at solving the complex multi-hop problem that multiple dispersed knowledge fragments need to be integrated, high-accuracy and high-efficiency reasoning is achieved, the retrieval requirement is dynamically generated through an explicit thinking chain guiding mechanism, and evidence obtaining is optimized and redundant information is filtered in combination with a three-level self-feedback retrieval mechanism.
Owner:XI AN JIAOTONG UNIV

Large model reasoning optimization method based on context increment updating

The invention relates to the technical field of data processing, in particular to a large model reasoning optimization method based on context incremental updating, which comprises the following steps: processing a multi-modal data stream through timestamp alignment and a filtering algorithm, extracting features by adopting a shared encoder and a private encoder, and realizing feature decoupling through a depth information bottleneck principle. A dynamic emotion map is constructed by using Gaussian process regression and a random process algorithm, and self-adaptive updating control is realized by combining meta-learning and Bayesian optimization. Incremental state management is realized by adopting a neural Turing machine, reasoning consistency is guaranteed through a generative adversarial network, model parameters are optimized in combination with a digital twin system and reinforcement learning, and a mental health service response is finally generated through a conditional generation model and hierarchical reinforcement learning. According to the method, the problem of asynchronism of multi-modal emotion feature dynamic evolution and context increment updating is effectively solved, accumulated drift of emotion state tracking is eliminated, and the continuity of reasoning logic is guaranteed.
Owner:LUSHAN COLLEGE OF GUANGXI UNIV OF SCI & TECH

Language model reasoning resource scheduling method and system based on multi-agent cooperation

The invention relates to a multi-agent cooperation-based language model reasoning resource scheduling method and system, and the method achieves the intelligent simulation of multi-view iterative thinking in a complex reasoning task through the construction of a multi-agent system which is clear in division of labor and is provided with a special knowledge base. View limitation and decision deviation of a single model in long-sequence and multi-step reasoning are effectively overcome; depending on the fusion of the general capability of the large language model and the task special knowledge base, the professionality and accuracy of the reasoning result are improved. Meanwhile, the problems of resource waste, redundant calculation and unstable convergence caused by cognitive overload in the cooperation process are solved by combining multi-round dynamic cooperation with a convergence mechanism of cognitive load perception and a dynamic token number limitation and low-confidence branch pruning strategy; therefore, on the premise that the reasoning quality is guaranteed, the utilization efficiency of computing resources is remarkably optimized, peak value occupation is reduced, the overall reasoning time delay is shortened, and reliable technical support is provided for efficient and stable deployment of a large language model in a complex task.
Owner:GUANGDONG SOUTH SMART MEDIA TECH CO LTD

Better inference pattern for long context retrieval

The present disclosure relates to a method and system for enhancing inference in large language models (LLMs) over long input sequences. A segmented inference strategy may be employed, wherein the long context can be divided and sequentially processed through a key-value (KV) cache of the LLM. At each step, the model may generate auxiliary outputs (or margins), which may include extractive summaries or intermediate signals based on the segment's relevance to an instruction. These margins may then be classified and selectively retained to guide final inference on the instruction. The retained margins may be prepended to the instruction to facilitate improved generation without modifying the model's internal weights. The disclosed approach provides efficient localization of relevant content, improves comprehension of extended contexts, and reduces computational overhead. Moreover, the disclosed technique is particularly effective for retrieval-based NLP tasks and supports long-context reasoning in LLMs while enhancing inference efficiency and user experience.
Owner:WRITER INC

Police file and image association reasoning method of multi-modal large model

The invention discloses a multi-mode large model police file and image association reasoning method, and belongs to the technical field of artificial intelligence and police information processing. The method comprises the following steps: preprocessing police affair texts and images; respectively extracting word-level and sentence-level features and local and global visual features by using a text encoder and an image encoder; realizing alignment and joint representation through a cross-modal fusion layer and an attention mechanism; generating a preliminary reasoning result by adopting a task output head; a heterogeneous evidence graph is constructed, and priori knowledge is injected in combination with the knowledge graph; and finally, joint reasoning is carried out through the graph neural network, and a cross-modal reasoning result and a complete evidence chain are output. According to the method, deep correlation analysis of police affair texts and images is realized, and the accuracy and interpretability of case research and judgment are improved.
Owner:JIANGSU LIANFENG GOLDEN SHIELD INTELLIGENT TECH CO

Tree safety risk assessment method, system and device based on key index priority mechanism and storage medium

The invention discloses a tree safety risk assessment method, system and device based on a key index priority mechanism, and a storage medium. The method comprises the following steps: constructing an inference model according to a fault tree model of a street tree and a fuzzy Bayesian network; based on the inference model, through reverse inference, sensitivity analysis and most approximate cause chain analysis, screening out a key index set and a conventional index set from a preset risk candidate index set; and based on a hierarchical risk assessment principle, respectively assessing the key index set and the conventional index set to obtain an assessment result. According to the inference model, key factors having significant contributions to the overall risk of the border tree are identified from a large number of risk indexes; a conjoint analysis method of reverse reasoning, sensitivity analysis and most approximate cause chain analysis is adopted, so that false alarm and missing alarm are effectively reduced; according to the hierarchical risk assessment principle, key factors are assessed preferentially, it is ensured that the street tree with structural fatal defects can be recognized and controlled in time, and the reliability and timeliness of assessment are improved.
Owner:SHANGHAI GREENING MANAGEMENT GUIDANCE STATION +1

Model training method and device, electronic equipment and storage medium

The embodiment of the invention provides a model training method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the steps of obtaining preset question and answer training data, wherein the question and answer training data comprises input question data and a model output result corresponding to the input question data; reasoning trajectory generation is carried out based on each piece of input question data and the corresponding model output result, and a plurality of candidate reasoning trajectories are obtained; performing confusion degree evaluation based on the plurality of candidate reasoning trajectories to obtain a confusion degree score corresponding to each candidate reasoning trajectory; a target reasoning trajectory is determined from the plurality of candidate reasoning trajectories based on the confusion score to guide model training. According to the embodiment of the invention, stable model training can be carried out on a large language model used for business scenes such as intelligent investment advisers or insurance claims and the like in a low-cost and high-efficiency manner, the logic stability of the large language model during business processing is improved, and business services with high professionality and depth are provided.
Owner:PING AN TECH (SHENZHEN) CO LTD

Bidirectional information retrieval enhancement generation method for large language model

The invention discloses a bidirectional information retrieval enhancement generation method for a large language model, and belongs to the technical field of artificial intelligence. In order to overcome the defects that noise is introduced and key evidences are omitted due to the fact that traditional RAG only executes'query-document 'one-way retrieval, a two-way semantic perception retrieval enhancement generation model and a two-stage training framework are constructed, wherein in the first stage, the positive / negative example distance is increased in an embedded space in a contrast learning self-supervision mode; in the second stage, fine-grained correlation discrimination is carried out on query-document bidirectional sentences through supervised dichotomy, and probabilistic correlation scores are output; in the reasoning stage, the bidirectional probabilities are fused according to Bayesian to obtain final relevancy, document reordering is carried out, and plug and play can be achieved without fine adjustment of LLM in the whole process. According to the method, the accuracy and consistency of single-hop and multi-hop questions and answers and fact checking tasks are remarkably improved, and the method has the advantages of light weight and low deployment cost.
Owner:中华人民共和国大连海关

Method for compressing thinking chain of reasoning large model

The invention discloses an inference large model thinking chain compression method, which comprises the following steps of: firstly, generating answer sets with different detailed degrees by utilizing multiple rounds of sampling of a basic large model, and adaptively selecting an inference chain length by adopting a dynamic quantile algorithm based on task difficulty; secondly, performing diversity rewriting and compression on the reasoning step through KL divergence constraint, and generating the shortest expression on the premise of ensuring semantic consistency; constructing positive and negative samples to guide the model to learn simple expression, and training by adopting a composite loss function including supervised learning and length perception preference optimization; according to the method, external annotation data or a teacher model is not needed, adaptive matching of the reasoning depth and the problem difficulty can be achieved, the semantic integrity is guaranteed, meanwhile, the reasoning efficiency is remarkably improved, high expandability and good cross-task migration ability are achieved, and the method is particularly suitable for large-model lightweight deployment in a low-computing-power environment.
Owner:ZHEJIANG UNIV

Case-based reasoning method for man-machine collaborative personalized exercise training driven by causal knowledge

The invention relates to the technical field of intelligent exercise training, and particularly discloses a causal knowledge-driven man-machine collaborative personalized exercise training case reasoning method, which comprises the following steps of: constructing a case library containing a plurality of cases, and performing causal knowledge reconstruction at different levels; searching a case most similar to the target case from a case library based on the weight of each description variable; the difference of training scheme variables between the target case and the retrieved similar cases is recognized, anti-fact intervention and inference of any training scheme variable are conducted on the target case, and one or more new cases are screened out; when an actual result generated after the new case is executed does not conform to expectation, executing an attribution process, and generating different processing suggestions according to attribution types; new cases which do not conform to expectation are stored in a case library, and causal knowledge reconstruction is periodically carried out. According to the method, causal science and human-computer interaction can be deeply fused to carry out personalized training scheme reasoning.
Owner:CHINA INST OF SPORT SCI

Retrieval enhancement generation reasoning method for reinforcement learning and thinking chain based on rules

The invention relates to a rule-based reinforcement learning and thinking chain retrieval enhancement generation reasoning method. The method comprises the following steps of: preprocessing data; model training; and performing actual reasoning by adopting the trained model. The method has the beneficial effects that the complex multi-step reasoning capability of a large model is remarkably improved; knowledge integration is carried out, and the efficiency of retrieval enhancement generation (RAG) is optimized; the training cost is reduced; the stability of the strategy is enhanced; dynamic decision support and logic specification constraints are enhanced; the method can perfectly accord with verification requirements for multi-step reasoning ability, and can better adapt to knowledge and language styles in specific fields; the problem that a traditional RAG retrieval result is disjointed with an inference chain is solved, and the capacity of the model for processing complex and multi-step inference tasks is improved.
Owner:JIANGSU ZHONGNONG IOT TECH CO LTD

Business execution method and device based on large model, medium and equipment

The embodiment of the invention discloses a business execution method based on a large model, and the method comprises the steps: taking data inputted into the large model as initial data, reversely determining a final step which needs to be executed to obtain an output target, forwardly determining an inference boundary condition which can be deduced by the initial data, and according to the inference boundary condition and the final step, carrying out the business execution according to the inference boundary condition and the final step. And constructing a reasoning process to obtain an output result, and executing a service according to the output result. According to the method, a global solution direction is determined through independent reverse planning, blindness and local traps caused by only forward reasoning are avoided, better robustness and adaptability are embodied in a complex problem scene through butt joint of forward derivation verification and reverse planning, computing resources can be saved, and meaningless search overhead is reduced.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Resource determination method and apparatus, and program product and storage medium

Disclosed in the embodiments of the present application are a resource determination method and apparatus, and a program product and a storage medium. The method in the embodiments of the present application comprises: receiving an inference request; determining a first computing resource required for when a large model performs inference on the inference request; acquiring first resource information of a heterogeneous cluster; and on the basis of the first computing resource and the first resource information, determining a first inference resource from the heterogeneous cluster, wherein the first inference resource is used by the large model to perform inference on the inference request. In this way, for an inference request received in real time, a computing resource required for performing inference on the inference request is determined in real time, which enables dynamic determination of computing resources corresponding to inference requests of different lengths. Thus, an inference resource corresponding to the inference request can be dynamically determined on the basis of the computing resource, and the corresponding inference resource can be more accurately determined from the heterogeneous cluster, so that inference is performed on the inference request on the basis of the inference resource, such that an inference process of the large model can avoid being affected.
Owner:HUAWEI TECH CO LTD

Knowledge-guided meter reading method and system based on visual language large model

The invention relates to the technical field of computer vision and artificial intelligence, in particular to a knowledge-guided meter reading method and system based on a visual language large model, and the method comprises the following steps: a template obtaining step: obtaining a meter image, and calling a structured reading prompt template containing metadata and an inference logic instruction based on the type; a multi-modal input step: taking the image and the template as multi-modal data and inputting the multi-modal data into the visual language large model; a reasoning generation step: extracting visual features, performing attention alignment on the visual features and the template, responding to an instruction to execute thinking chain reasoning, and generating a natural language description result containing an intermediate derivation basis; an analysis output step: performing structured analysis on the result, and extracting a meter reading value conforming to the format constraint; according to the invention, deep fusion of business logic and visual perception is realized, man-machine trust is established, and logic mistake of a pure visual model is avoided.
Owner:SHENZHEN LAIDA SIWEI INFORMATION TECH CO LTD

UE signalling applicatiblity or inapplicability of ai / ML functionalities based on network-provided interference configuration

PCT designated stageWO2026010549A1Spatial transmit diversityMachine learningUser deviceApplicatorful
In an embodiment, a method performed by a user equipment (UE) for signaling an applicability of Artificial Intelligence Machine Learning (AI / ML) functionalities comprises receiving, from a network node, a plurality of inference related configurations associated with at least one supported functionality; determining an applicability status for each inference related configuration of the plurality of inference related configurations, wherein the applicability status indicates whether the inference related configuration is applicable or non-applicable. The method also comprises transmitting a first indication comprising one or more of an applicability indication indicating that an AI / ML model functionality is determined to be applicable and an inapplicability indication indicating that the AI / ML model functionality is not determined to be applicable according to one or more inference related configurations.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Kubernetes-based large model reasoning optimization method, system and equipment

The invention provides a Kubernetes-based large model reasoning optimization method, system and equipment, and the method comprises the steps: constructing a resource scheduling assembly for a large model, and monitoring a real-time resource state at a current moment; constructing a batch processing agent component for the large model, performing feature analysis on the received reasoning request queue, and determining request features corresponding to the reasoning requests in the reasoning request queue; the real-time resource state is evaluated according to the request features, and the batch processing size is adjusted based on the evaluation result to generate the optimal batch processing length; batching the reasoning requests in the reasoning request queue according to the optimal batch processing length, and packaging each batch of reasoning requests into batch data; the resource scheduling component is utilized to distribute batch data to Pod in the K8s cluster for reasoning, a large model reasoning result is obtained, the quantity of the reasoning Pod is automatically adjusted according to the flow and the resource load required by reasoning, a GPU with poor performance is prevented from becoming a system bottleneck, and the responsiveness and the stability of service are improved.
Owner:广域铭岛数字科技有限公司 +1