Model reasoning method and device and related equipment

By performing reasoning on a second electronic device and probability calculation on a first electronic device, and by using validation models and knowledge distillation training methods, the problem of high computational resource consumption and low efficiency of large language models on a single device is solved, and more efficient reasoning results are output.

CN121860036APending Publication Date: 2026-04-14INNER MONGOLIA MOBILE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The problem of large language models consuming large computational resources and being inefficient when inferring on a single electronic device.

Method used

Reasoning is performed on a second electronic device, probability calculations are performed on a first electronic device, the accuracy of the reasoning results is determined using a verification model, a simpler reasoning model is generated by knowledge distillation training, and the optimal word segmentation strategy is used to improve efficiency during word segmentation processing.

Benefits of technology

It reduces the computational resource consumption of the first electronic device, improves the output efficiency and accuracy of inference results, and enhances the overall inference efficiency through collaborative operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860036A_ABST
    Figure CN121860036A_ABST
Patent Text Reader

Abstract

The invention provides a model reasoning method and device and related equipment, and relates to the technical field of artificial intelligence, the model reasoning method comprises the steps that a target cue word is sent to second electronic equipment, the target cue word is used for enabling a reasoning model deployed on the second electronic equipment to conduct reasoning according to the target cue word, and the target cue word is sent to the second electronic equipment; generating a first reasoning result; the first reasoning result sent by the second electronic equipment is received, the first reasoning result is input into a verification model for probability calculation, a target probability is obtained, the target probability is used for representing the probability that the first reasoning result is the reasoning result corresponding to the target prompt word, and the target prompt word is sent to the second electronic equipment. The reasoning model is determined according to the verification model; and under the condition that the target probability is greater than an acceptable probability, determining the first reasoning result as a reasoning result corresponding to the target cue word. Therefore, the output efficiency of the reasoning result corresponding to the target prompt word can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a model reasoning method, apparatus and related equipment. Background Technology

[0002] With the continuous development of artificial intelligence technology, large language models are being applied more and more widely in people's lives. Large language models in related technologies have not only significantly improved in accuracy and comprehension, but their application scope has also expanded from simple text generation to more complex fields, such as personalized recommendation, sentiment analysis, and multilingual translation. However, large language models in related technologies are typically deployed on a single electronic device, and all reasoning services are completed on this single device. Therefore, current large language models are prone to problems such as high computational resource consumption and low reasoning efficiency on a single electronic device during reasoning. Summary of the Invention

[0003] This application provides a model reasoning method, apparatus, and related equipment to address the problem that current large language models tend to consume a large amount of computing resources from a single electronic device and have low reasoning efficiency during reasoning.

[0004] To solve the above problems, this application is implemented as follows:

[0005] In a first aspect, embodiments of this application provide a model reasoning method, including:

[0006] Send a target prompt word to a second electronic device, the target prompt word being used to cause an inference model deployed on the second electronic device to perform inference based on the target prompt word in order to generate a first inference result;

[0007] The system receives the first inference result sent by the second electronic device and inputs the first inference result into the verification model for probability calculation to obtain the target probability. The target probability is used to represent the probability that the first inference result is the inference result corresponding to the target prompt word. The inference model is determined according to the verification model.

[0008] If the target probability is greater than the acceptable probability, the first inference result is determined as the inference result corresponding to the target prompt word.

[0009] Secondly, embodiments of this application provide a model inference apparatus, including:

[0010] A first sending module is configured to send a target prompt word to a second electronic device, the target prompt word being used to cause an inference model deployed on the second electronic device to perform inference based on the target prompt word in order to generate a first inference result;

[0011] The first receiving module is configured to receive the first inference result sent by the second electronic device, and input the first inference result into the verification model for probability calculation to obtain the target probability. The target probability is used to represent the probability that the first inference result is the inference result corresponding to the target prompt word. The inference model is determined according to the verification model.

[0012] The first determining module is used to determine the first inference result as the inference result corresponding to the target prompt word when the target probability is greater than the acceptable probability.

[0013] Thirdly, embodiments of this application also provide an electronic device, including: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the method described in the first aspect above.

[0014] Fourthly, embodiments of this application also provide a readable storage medium for storing a program, which, when executed by a processor, implements the steps of the method described in the first aspect above.

[0015] Fifthly, embodiments of this application also provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the method described in the first aspect above.

[0016] In this embodiment, a target prompt word can be sent to a second electronic device, causing the inference model deployed on the second electronic device to perform inference based on the target prompt word to generate a first inference result. This first inference result is then input into a verification model deployed on the first electronic device for probability calculation to obtain a target probability. If the target probability is greater than the acceptable probability, the first inference result is determined as the inference result corresponding to the target prompt word. By performing inference on the second electronic device and probability calculation on the first electronic device, the consumption of computing resources on the first electronic device is reduced. Furthermore, the efficiency of probability calculation on the first electronic device is far higher than that of inference. Simultaneously, by fully utilizing the computing resources of the second electronic device for inference, the output efficiency of the inference result corresponding to the target prompt word can be improved through the collaborative operation of the first and second electronic devices. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the structure of a network system to which the embodiments of this application can be applied;

[0019] Figure 2 This is a flowchart illustrating the model reasoning method provided in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of the structure of the model inference device provided in the embodiments of this application;

[0021] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.

[0024] Please see Figure 1 , Figure 1 This is a structural diagram of a network system to which the embodiments of this application can be applied, such as... Figure 1 As shown, it includes a first electronic device 11 and a second electronic device 12, wherein the first electronic device 11 and the second electronic device 12 can communicate with each other.

[0025] In practical applications, the first electronic device 11 can be a base station, an Access and Mobility Management Function (AMF), a relay, an access point, or other network elements, while the second electronic device 12 can be a mobile phone, a tablet personal computer, a laptop computer, a personal digital assistant (PDA), a mobile internet device (MID), a wearable device, or an in-vehicle device, etc.

[0026] The following describes the model reasoning method provided in the embodiments of this application.

[0027] See Figure 2 , Figure 2 This is a flowchart illustrating the model reasoning method provided in the embodiments of this application. Figure 2 The model inference method shown can be executed by a first electronic device, which can be referred to as a cloud server or cloud, and the corresponding second electronic device can be referred to as a terminal, terminal device or user terminal, etc.

[0028] like Figure 2 As shown, the model inference method may include the following steps:

[0029] Step 201: Send a target prompt word to the second electronic device. The target prompt word is used to cause the inference model deployed on the second electronic device to perform inference based on the target prompt word to generate a first inference result.

[0030] The reasoning model can be understood as a model that infers based on prompts to generate a reasoning result.

[0031] The specific length of the target prompt is not limited here. Optionally, the length of the target prompt can be less than or equal to the preset length, and thus the target prompt can be called a short prompt. Alternatively, the length of the target prompt can be greater than the preset length, and thus the target prompt can be called a long prompt. The length of the target prompt can be understood as the number of texts included in the target prompt. For example, when the number of texts is greater than the preset number, the length of the target prompt is greater than the preset length; when the number of texts is less than or equal to the preset number, the length of the target prompt is less than or equal to the preset length.

[0032] The specific content of the target prompt is not limited here. Optionally, the target prompt can include "generate a current status survey report for industry A", or alternatively, the target prompt can include "generate the optimal navigation path from location B to location C".

[0033] The first inference result can be the inference result obtained by reasoning from the target prompt word. For example, if the target prompt word includes the content "generate a current status survey report for industry A", then the first inference result is a current status survey report for industry A; if the target prompt word includes the content "generate the optimal navigation path from location B to location C", then the first inference result is the optimal navigation path from location B to location C.

[0034] Step 202: Receive the first inference result sent by the second electronic device, and input the first inference result into the verification model for probability calculation to obtain the target probability. The target probability is used to represent the probability that the first inference result is the inference result corresponding to the target prompt word. The inference model is determined according to the verification model.

[0035] The target probability is used to represent the probability that the first inference result is the inference result corresponding to the target prompt word. It can be understood as: the target probability can be used to represent the probability that the first inference result is the true inference result of the target prompt word. The larger the value of the target probability, the higher the probability that the first inference result is the true inference result of the target prompt word.

[0036] The validation model can be understood as a pre-trained model used to calculate the inference results. Optionally, the validation model can be a model trained on the basis of a general model using sample data. Alternatively, the validation model can also be a model trained on the basis of a self-built model using sample data.

[0037] The inference model is determined based on the validation model. This can be understood as the inference model and the validation model having a high degree of correlation. This results in a higher degree of fit between the inference model and the validation model, which allows the validation model to more accurately identify the first inference result generated by the inference model, thereby further improving the accuracy of the validation model in calculating the probability of the first inference result.

[0038] It should be noted that the specific method by which the inference model is determined based on the validation model is not limited here. Optionally, the inference model and the validation model can be models trained using the same sample data; alternatively, the above-mentioned inference model can be generated by fine-tuning the above-mentioned validation model.

[0039] As an optional implementation, before sending the target prompt word to the second electronic device, the method further includes:

[0040] The validation model is trained by distillation to obtain the inference model, wherein the number of structural layers in the inference model is less than the number of structural layers in the validation model.

[0041] The installation package of the inference model is sent to the second electronic device, the installation package being used to enable the second electronic device to deploy the inference model.

[0042] The number of structural layers in the inference model is less than that in the verification model. This can be understood as: the number of encoding and decoding layers in the inference model is less than the number of encoding and decoding layers in the verification model.

[0043] Since the number of structural layers in the inference model is less than that in the validation model, the validation model can be called the large model and the inference model can be called the small model. Specifically, the large model can be trained by knowledge distillation to generate a small model with the same structure but a significantly reduced parameter size (i.e., fewer structural layers). Optionally, during the distillation training process, the small model is trained by backpropagation with the objective function of minimizing the distillation loss. In this way, the number of encoder and decoder layers in the small model can be greatly reduced, thereby reducing the burden on the computing resources of the second electronic device, while maintaining the high inference accuracy of the inference model.

[0044] The specific type of the second electronic device is not limited here. Optionally, the second electronic device may include a terminal device or an edge device.

[0045] In this embodiment, the inference model is a model obtained by distilling the verification model, thereby making the inference model and the verification model more compatible, and thus making the verification model more accurate when verifying the first inference result generated by the inference model.

[0046] For example, the construction and optimization process of the above inference model can be described as follows:

[0047] 1. Obtain a validation model. The validation model can be a large language model, which may include generative pretrained Transformer (GPT) models, bidirectional encoder representations from Transformers (BERT) models, and other large text models. These large text models typically use Transformer models. Assume the expression of the large text model is... The encoder has n layers and the decoder has m layers.

[0048] 2. Initialize the inference model: Distill the validation model to form the inference model. With validation model They have the same structure, except for the reasoning model. The encoder has g layers and the decoder has q layers, where... , Random initialization parameter.

[0049] 3. Distillation Loss Function: The loss function guides the small terminal model to imitate the behavior of the large cloud model.

[0050] (1) Soft tag loss;

[0051] The softmax output of the validation model is used as the objective to assist the inference model in learning the probability distribution relationship of the validation model's inference. The formula for calculating the soft label loss can be found below:

[0052] ;

[0053] in, For soft label loss, To verify the model's next lexical ( The predicted probability of ). For the next inference model The predicted probability; T is the temperature parameter, initially set to 3; C is the output distribution. To illustrate this implementation, assume the input text to be inferred (i.e., the target prompt) is "China Mobile is the world's largest telecommunications operator, and the company's core advantage is...", the next... For "5G network", =0.6, =0.5, Next For "industry", =0.3, =0.3, next For "finance", =0.1, =0.2.

[0054] (2) Hard label loss;

[0055] We directly use the final output of the validation model as the objective to calculate the hard-label loss. The formula for calculating the hard-label loss is as follows:

[0056] ;

[0057] in, For hard label loss, These are training samples for the inference model. The predicted probability of the inference model is obtained from the training samples. The actual output result can be obtained. For example, if the actual output result of the sample is "5G network", then... =1, the rest are 0.

[0058] (3) Distillation loss function;

[0059] By combining soft-label and hard-label losses, a weighted optimization inference model is calculated using the distillation loss function. The formula for calculating the distillation loss function can be found in the following description:

[0060] ;

[0061] Where L represents the distillation loss, These are the weighting coefficients, and The value of is not limited here; optionally, The value can be 0.5.

[0062] It should be noted that, in order to minimize the distillation loss function mentioned above, the inference model can be obtained iteratively through backpropagation based on the training sample data. And inference models can be deployed on each second electronic device. Preliminary reasoning can be provided on the second electronic device.

[0063] Step 203: If the target probability is greater than the acceptable probability, determine the first reasoning result as the reasoning result corresponding to the target prompt word.

[0064] The specific value of the acceptable probability is not limited here. Optionally, the acceptable probability can be a preset fixed value. Alternatively, the acceptable probability can be a value that is updated in real time according to the content of the target prompt word. For example, when the content of the target prompt word includes the first content, the acceptable probability is the first value. When the content of the target prompt word includes the second content, the acceptable probability of disclaimer is the second value.

[0065] Specifically, when the target probability is greater than the acceptable probability, it indicates that the accuracy of the first inference result is relatively high. Therefore, the first inference result can be determined as the inference result corresponding to the target prompt word. This ensures that the accuracy of the determined inference result corresponding to the target prompt word is relatively high. At the same time, it reduces the consumption of computing resources on the first electronic device.

[0066] In this embodiment, steps 201 to 203 send a target prompt word to a second electronic device, causing the inference model deployed on the second electronic device to perform inference based on the target prompt word to generate a first inference result. The first inference result is then input into a verification model deployed on the first electronic device for probability calculation to obtain a target probability. If the target probability is greater than the acceptable probability, the first inference result is determined as the inference result corresponding to the target prompt word. By performing inference on the second electronic device and probability calculation on the first electronic device, the consumption of computing resources on the first electronic device is reduced. Furthermore, the efficiency of probability calculation on the first electronic device is much higher than that of inference. Simultaneously, by fully utilizing the computing resources of the second electronic device for inference, the output efficiency of the inference result corresponding to the target prompt word can be improved through the collaborative operation of the first and second electronic devices.

[0067] As an optional implementation, before sending the target prompt word to the second electronic device, the method further includes:

[0068] Receive initial prompt words;

[0069] If the length of the initial prompt word is determined to be greater than the preset length, a preset word segmentation strategy is obtained, which is a strategy for segmenting the prompt word.

[0070] The initial prompt word is segmented according to the preset word segmentation strategy to obtain multiple segmented prompt words, and the target prompt word includes at least a portion of the multiple segmented prompt words.

[0071] The initial prompt can be understood as a prompt entered by the user, or it can be a prompt received by the first electronic device from other electronic devices.

[0072] When the target prompt word includes at least two sub-prompt words, the at least two sub-prompt words can be sequentially input into the inference model for inference, thereby generating multiple inference sub-results in sequence. The multiple inference sub-results are then aggregated to generate the first inference result.

[0073] Alternatively, the number of inference models can be at least two. These at least two sub-prompts can be input into different inference models for inference, thereby generating multiple inference sub-results. These multiple inference sub-results are then aggregated to generate the first inference result. In this way, the inference processes of the at least two inference models can be executed in parallel, further improving the efficiency of generating the first inference result.

[0074] Since the preset word segmentation strategy is to segment the prompt words, the initial prompt words are segmented according to the preset word segmentation strategy, which can improve the accuracy of the word segmentation results and avoid the phenomenon of splitting multiple texts that originally belong to one word group into multiple word groups.

[0075] In this embodiment, when the length of the initial prompt word is greater than the preset length, it indicates that the initial prompt word is too long, i.e., the initial prompt word is a long prompt word. Since the inference model has low recognition accuracy and processing efficiency for long prompt words, the accuracy of the generated first inference result is also low. Therefore, the initial prompt word can be segmented according to the preset word segmentation strategy to obtain multiple segmented prompt words. The target prompt word includes at least a part of the segmented prompt words. In this way, the target prompt word input into the inference model is a short prompt word, thereby enhancing the inference model's recognition accuracy and processing efficiency for the target prompt word, and thus improving the accuracy and generation efficiency of the generated first inference result.

[0076] As an optional implementation, before obtaining the preset word segmentation strategy when the length of the initial prompt word is determined to be greater than a preset length, the method further includes:

[0077] Obtain multiple initial word segmentation strategies;

[0078] The multiple initial word segmentation strategies are input into the pre-trained inference latency prediction model to predict inference latency, thereby obtaining the inference latency prediction value corresponding to each initial word segmentation strategy.

[0079] The initial word segmentation strategy corresponding to the smallest inference delay prediction value among the multiple inference delay prediction values ​​is determined as the preset word segmentation strategy.

[0080] The initial word segmentation strategy can be understood as the strategy for segmenting the target prompt word. For example, if the target prompt word includes the content "China Mobile is the world's largest telecommunications operator", it can be segmented into "China Mobile", "is", and "the world's largest telecommunications operator" according to one initial word segmentation strategy, or into "China", "Mobile", "is", "the world's largest", and "the telecommunications operator" according to another initial word segmentation strategy.

[0081] Among them, the inference delay prediction value can be used to represent the length of the inference delay. The larger the inference delay prediction value, the longer the inference delay, the greater the delay when the initial word segmentation strategy segments the prompt words, and the lower the efficiency of word segmentation of the prompt words.

[0082] It should be noted that the latency corresponding to each word segmentation process can be summed to obtain the aforementioned inference latency prediction value. Alternatively, a weighted sum of the latency corresponding to each word segmentation process can be calculated to obtain the aforementioned inference latency prediction value. The specific method is not limited here.

[0083] In this embodiment, the initial word segmentation strategy corresponding to the smallest inference delay prediction value among multiple inference delay prediction values ​​can be determined as the preset word segmentation strategy. This ensures that the delay corresponding to the initial prompt word segmentation processing according to the preset word segmentation strategy is small, thereby ensuring that the delay in generating the first inference result is also small.

[0084] As an optional implementation, before inputting the plurality of initial word segmentation strategies into a pre-trained inference latency prediction model to predict inference latency and obtain the inference latency prediction value corresponding to each initial word segmentation strategy, the method further includes:

[0085] Obtain sample parameters, which include at least one of the following: computing power resources of the second electronic device, network resources of the second electronic device, computing power resources of the first electronic device, network resources of the first electronic device, and word segmentation results of historical sample prompt words with a length greater than a preset length;

[0086] The sample parameters are input into the initial model to be trained for the Nth training iteration, and the Nth time delay prediction value is output, where N is an integer greater than 1.

[0087] If the difference between the Nth predicted latency value and the target latency value is less than a preset difference, the initial model trained for the Nth time is determined as the inference latency prediction model, and the target latency value is the actual latency value corresponding to the sample parameters.

[0088] The computing resources may include at least one of the following: computing resources of a central processing unit (CPU), memory, and a graphics processing unit (GPU).

[0089] Network resources can be understood as data interfaces, network protocols, etc.

[0090] The word segmentation results of historical sample prompt words with a length greater than the preset length can be understood as the word segmentation results of long prompt words included in the historical samples. The word segmentation results can include the number of word segments, word size, and number of words.

[0091] In this embodiment of the application, the inference latency prediction model can be obtained by training with sample parameters. Since the sample parameters include at least one of the following: computing resources of the second electronic device, network resources of the second electronic device, computing resources of the first electronic device, network resources of the first electronic device, and word segmentation results of historical sample prompts with a length greater than a preset length, the training time of the inference latency prediction model can be shortened and the training rate of the latency prediction model can be improved by using the above sample parameters.

[0092] As an optional implementation, determining the first inference result as the inference result corresponding to the target prompt word when the target probability is greater than the acceptable probability includes:

[0093] If the target probability is greater than the acceptable probability, random sampling is performed based on the target prompt word to obtain a random sample value;

[0094] If the target probability is greater than the random sampling value, the first inference result is determined as the inference result corresponding to the target prompt word.

[0095] Here, the random sample value can be understood as the probability that the randomly generated inference result is the inference result corresponding to the target prompt word. The method of generating the randomly generated inference result can be seen as follows: the target prompt word is randomly generated according to the random generation algorithm, thereby obtaining the above-mentioned randomly generated inference result.

[0096] It should be noted that only when the target probability is greater than the random sample value can it be said that the reliability and accuracy of the first inference result are high. If the target probability is less than or equal to the random sample value, it means that the first inference result generated by the inference model is not much different from the randomly generated inference result, and may even be less accurate than the randomly generated inference result.

[0097] In this embodiment, when the target probability is greater than the acceptable probability, the size of the target probability and the random sample value can be further determined. Only when the target probability is greater than the random sample value is the first inference result determined as the inference result corresponding to the target prompt word. This can further improve the accuracy of the determined inference result corresponding to the target prompt word and reduce the occurrence of misjudgment.

[0098] As an optional implementation, after receiving the first inference result sent by the second electronic device and inputting the first inference result into the verification model for probability calculation to obtain the target probability, the method further includes:

[0099] If the target probability is less than or equal to the acceptable probability, the target prompt word is input into the verification model for inference to generate a second inference result;

[0100] The second reasoning result is determined as the reasoning result corresponding to the target prompt word.

[0101] The specific method of inputting the target prompt word into the validation model for reasoning to generate the second reasoning result is not limited here. Optionally, the target prompt word can be input into the validation model for reasoning, so that the validation model restarts the reasoning to generate the second reasoning result. Since the number of structural layers in the validation model is greater than the number of structural layers in the reasoning model, the validation model generates the second reasoning result more efficiently, thus ensuring that the determination efficiency of the reasoning result corresponding to the target prompt word is also higher.

[0102] Alternatively, by calculation, if the target probability of the target result in the first inference result is less than or equal to the acceptable probability, the target prompt word and the target result can be input into the verification model for inference. The verification model can then infer the content corresponding to the target result, generate corrected content, and replace the target content in the first inference result with the corrected content to generate the second inference result. In this way, there is no need to perform inference again, thereby improving the generation efficiency of the second inference result.

[0103] In this embodiment, when the target probability is less than or equal to the acceptable probability, it indicates that the accuracy of the first inference result generated by the inference model is low. Therefore, the target prompt word can be input into the verification model for inference to generate a second inference result. The second inference result is then determined as the inference result corresponding to the target prompt word. In this way, the accuracy of the inference result corresponding to the target prompt word can always be guaranteed to be high.

[0104] It should be noted that the training process of the inference latency prediction model can be described as follows:

[0105] This application argues that during the segmentation of long prompt words, the operating status of the first electronic device, the operating status of the second electronic device, and the size of the prompt word block can all affect the inference latency of the large model. Therefore, optionally, the operating status of the first and second electronic devices, the segmentation size of the prompt word (i.e., the segmentation results of historical sample prompt words) can be used as independent variables, and the predicted latency value can be used as the dependent variable. A neural network framework can then be used to train an inference latency prediction model. The specific steps are as follows:

[0106] (1) Input layer of block inference latency model;

[0107] The implementation method of this application uses a real-time monitoring tool to monitor the computing power and network resources of the first electronic device (which can be referred to as a cloud server) and the second electronic device (which can be referred to as a terminal). At the same time, it combines the segmentation of long prompt words (i.e., the word segmentation results of historical sample prompt words) as input information for the inference latency prediction model, as shown in Table 1:

[0108] Table 1 Input information of the input layer of the inference delay prediction model

[0109] Serial Number variable Variable description 1 CPU status of the i-th terminal 2 The i-th terminal memory 3 Disk status of the i-th terminal 4 Network status of the i-th terminal 5 Cloud server CPU status 6 Cloud server memory status 7 Cloud server disk status 8 Cloud server network status 9 Prompt word segmentation strategy

[0110] (2) Hidden layer of the inference delay prediction model;

[0111] The hidden layer consists of one layer and contains eighteen neurons. For any number of neurons in the hidden layer... Parameters, hidden layer activation function You can choose The function, the output of the hidden layer is The calculation method can be found in the following formula:

[0112] ;

[0113] Among them, w 1,j w 2,j to w 9,j All are weighting coefficients, β j It can be a preset fixed value.

[0114] (3) Output information of the inference delay prediction model

[0115] The inference latency prediction value output by the inference latency prediction model The calculation method can be found in the following formula:

[0116] ;

[0117] (4) Construct a loss function and train the inference delay prediction model;

[0118] The loss function for constructing the inference latency prediction model, assuming there are n samples in the system, can be calculated using the following formula:

[0119] ;

[0120] Where Loss is the loss function. y represents the predicted value, and y represents the actual value. The model is most accurate only when the loss value is minimized. w, v, β, and λ are all weight coefficients of the model. First, w, v, β, and λ are initialized, all parameters are randomly generated, and the optimal solution w, v, β, and λ are obtained using backpropagation, thus obtaining the inference delay prediction model.

[0121] It should be noted that when a long prompt word is received, the inference latency prediction model is used to predict the inference latency prediction value of different initial word segmentation strategies. The initial word segmentation strategy corresponding to the smallest inference latency prediction value among multiple inference latency prediction values ​​is selected as the preset word segmentation strategy. Based on the preset word segmentation strategy, the inference model deployed on the second electronic device executes the word segmented prompt word in parallel to improve inference efficiency.

[0122] It should be noted that the specific process by which the inference model deployed on the second electronic device performs inference based on the target prompt words to generate the first inference result can be described in the following description:

[0123] 1. The service request (i.e., the target prompt word) received by the first electronic device is: ,in Include indivual For example, a service request could be something like, "China Mobile is the world's largest telecommunications operator, and the company's core strengths are":

[0124] ;

[0125] 2. Validate the model based on Input, generate the first Represented as :

[0126] ;

[0127] in, This represents the validation model, which is based on the input. Generate the first for For example: the first It can be "5G network", and Connection generation The expression "China Mobile is the world's largest telecommunications operator, and its core advantage is its 5G network" is as follows:

[0128] ;

[0129] 3. Deploy the inference model on the second electronic device. , As input to the inference model, the next Based on different output probabilities, k types of results are generated, and each inference model outputs m results. The i-th type of result (i.e., the first inference result) can be represented as: ;

[0130] The specific method by which the verification model verifies the first inference result can be found in the following description:

[0131] First, the model is validated in parallel, verifying the first result of each of the k classes. And calculate the The probability of the i-th result is calculated. probability Similarly, the first result of the other k categories is calculated sequentially. The probability is calculated, and the result with the highest probability is selected as the optimal result. For example, suppose the inference model generates two types of inference results (i.e., two first inference results): one is "The characteristic of 5G networks is fast network speed". =0.7, =0.8; secondly, "the characteristic of 5G networks is low cost". =0.3, =0.2. Since the probability of "5G network is characterized by fast speed" is higher than that of "5G network is characterized by low cost", the verification model can prioritize verifying "5G network is characterized by fast speed".

[0132] Next, the model will be validated sequentially for all... The acceptable probability can be calculated using the formula described below:

[0133] ;

[0134] in, The inference probability of the inference model. To verify the inference probability of the model, as in the example above, "The characteristic of 5G networks is high speed." .

[0135] Finally, random sampling is generated. ,if If yes, accept and continue verifying the next one. If all of this type of result If all verifications pass, the result is acceptable. If... If the result fails, calculate the results for other classes, using the same validation method. If all classes fail, use the validation model to validate the failed classes. Begin, continue generating This continues until a result is generated (i.e., the second reasoning result).

[0136] Through the implementation method of this application, the verification efficiency of the verification model is much higher than that of the generation model. Therefore, the efficiency of inference can be greatly improved by adopting the implementation method of this application.

[0137] See Figure 3 , Figure 3 This is a structural diagram of the model inference device provided in the embodiments of this application, such as... Figure 3 As shown, the model inference device 300 includes:

[0138] The first sending module 301 is used to send a target prompt word to the second electronic device. The target prompt word is used to cause the inference model deployed on the second electronic device to perform inference based on the target prompt word in order to generate a first inference result.

[0139] The first receiving module 302 is used to receive the first inference result sent by the second electronic device, and input the first inference result into the verification model for probability calculation to obtain the target probability. The target probability is used to represent the probability that the first inference result is the inference result corresponding to the target prompt word. The inference model is determined according to the verification model.

[0140] The first determining module 303 is used to determine the first inference result as the inference result corresponding to the target prompt word when the target probability is greater than the acceptable probability.

[0141] As an optional implementation, the model inference device 300 further includes:

[0142] The second receiving module is used to receive the initial prompt word;

[0143] The first acquisition module is used to acquire a preset word segmentation strategy when it is determined that the length of the initial prompt word is greater than the preset length. The preset word segmentation strategy is a strategy for segmenting the prompt word.

[0144] The word segmentation processing module is used to segment the initial prompt word according to the preset word segmentation strategy to obtain multiple segmented prompt words, wherein the target prompt word includes at least a portion of the multiple segmented prompt words.

[0145] As an optional implementation, the model inference device 300 further includes:

[0146] The second acquisition module is used to acquire multiple initial word segmentation strategies;

[0147] The latency prediction module is used to input the multiple initial word segmentation strategies into the pre-trained inference latency prediction model to predict inference latency, and obtain the inference latency prediction value corresponding to each initial word segmentation strategy.

[0148] The second determining module is used to determine the initial word segmentation strategy corresponding to the smallest inference delay prediction value among the multiple inference delay prediction values ​​as the preset word segmentation strategy.

[0149] As an optional implementation, the model inference device 300 further includes:

[0150] The third acquisition module is used to acquire sample parameters, which include at least one of the following: the computing power resources of the second electronic device, the network resources of the second electronic device, the computing power resources of the first electronic device, the network resources of the first electronic device, and the word segmentation results of historical sample prompt words with a length greater than a preset length.

[0151] The training module is used to input the sample parameters into the initial model to be trained for the Nth training iteration and output the Nth time delay prediction value, where N is an integer greater than 1.

[0152] The third determining module is used to determine the initial model trained for the Nth time as the inference latency prediction model when the difference between the Nth latency prediction value and the target latency value is less than a preset difference. The target latency value is the actual latency value corresponding to the sample parameters.

[0153] As an optional implementation, the first determining module 303 includes:

[0154] The random sampling submodule is used to perform random sampling based on the target prompt word to obtain a random sample value when the target probability is greater than the acceptable probability.

[0155] The first determining submodule is used to determine the first inference result as the inference result corresponding to the target prompt word when the target probability is greater than the random sampling value.

[0156] As an optional implementation, the model inference device 300 further includes:

[0157] The reasoning module is used to input the target prompt word into the verification model for reasoning when the target probability is less than or equal to the acceptable probability, so as to generate a second reasoning result;

[0158] The fourth determining module is used to determine the second reasoning result as the reasoning result corresponding to the target prompt word.

[0159] As an optional implementation, the model inference device 300 further includes:

[0160] A distillation training module is used to perform distillation training on the verification model to obtain the inference model, wherein the number of structural layers in the inference model is less than the number of structural layers in the verification model.

[0161] The second sending module is used to send the installation package of the inference model to the second electronic device, the installation package being used to enable the second electronic device to deploy the inference model.

[0162] The model inference device 300 can realize the embodiments of this application. Figure 1 The various processes in the method embodiments, and the ways to achieve the same beneficial effects, will not be repeated here to avoid repetition.

[0163] This application also provides an electronic device. Please refer to [link to relevant documentation]. Figure 4 The electronic device may include a processor 401, a memory 402, and a program 4021 stored in the memory 402 and executable on the processor 401. When the electronic device is a first electronic device, the program 4021, when executed by the processor 401, can achieve... Figure 1 Any steps in the corresponding method embodiments and the achievement of the same beneficial effects will not be repeated here.

[0164] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by hardware related to program instructions, and the program can be stored in a readable medium. This application also provides a readable storage medium storing a computer program, which, when executed by a processor, can implement the above-described methods. Figure 1 Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.

[0165] The storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0166] This application also provides a computer program product, including computer instructions, which, when executed by a processor, can perform the above-described functions. Figure 1 Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.

[0167] The above description represents the preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A model reasoning method applied to a first electronic device, characterized in that, The method includes: Send a target prompt word to a second electronic device, the target prompt word being used to cause an inference model deployed on the second electronic device to perform inference based on the target prompt word in order to generate a first inference result; The system receives the first inference result sent by the second electronic device and inputs the first inference result into the verification model for probability calculation to obtain the target probability. The target probability is used to represent the probability that the first inference result is the inference result corresponding to the target prompt word. The inference model is determined according to the verification model. If the target probability is greater than the acceptable probability, the first inference result is determined as the inference result corresponding to the target prompt word.

2. The method according to claim 1, characterized in that, Before sending the target prompt word to the second electronic device, the method further includes: Receive initial prompt words; If the length of the initial prompt word is determined to be greater than the preset length, a preset word segmentation strategy is obtained, which is a strategy for segmenting the prompt word. The initial prompt word is segmented according to the preset word segmentation strategy to obtain multiple segmented prompt words, and the target prompt word includes at least a portion of the multiple segmented prompt words.

3. The method according to claim 2, characterized in that, Before obtaining the preset word segmentation strategy when the length of the initial prompt word is determined to be greater than the preset length, the method further includes: Obtain multiple initial word segmentation strategies; The multiple initial word segmentation strategies are input into the pre-trained inference latency prediction model to predict inference latency, thereby obtaining the inference latency prediction value corresponding to each initial word segmentation strategy. The initial word segmentation strategy corresponding to the smallest inference delay prediction value among the multiple inference delay prediction values ​​is determined as the preset word segmentation strategy.

4. The method according to claim 3, characterized in that, Before inputting the multiple initial word segmentation strategies into the pre-trained inference latency prediction model to predict inference latency and obtaining the inference latency prediction value corresponding to each initial word segmentation strategy, the method further includes: Obtain sample parameters, which include at least one of the following: computing power resources of the second electronic device, network resources of the second electronic device, computing power resources of the first electronic device, network resources of the first electronic device, and word segmentation results of historical sample prompt words with a length greater than a preset length; The sample parameters are input into the initial model to be trained for the Nth training iteration, and the Nth time delay prediction value is output, where N is an integer greater than 1. If the difference between the Nth predicted latency value and the target latency value is less than a preset difference, the initial model trained for the Nth time is determined as the inference latency prediction model, and the target latency value is the actual latency value corresponding to the sample parameters.

5. The method according to any one of claims 1 to 4, characterized in that, The step of determining the first inference result as the inference result corresponding to the target prompt word when the target probability is greater than the acceptable probability includes: If the target probability is greater than the acceptable probability, random sampling is performed based on the target prompt word to obtain a random sample value; If the target probability is greater than the random sampling value, the first inference result is determined as the inference result corresponding to the target prompt word.

6. The method according to any one of claims 1 to 4, characterized in that, After receiving the first inference result sent by the second electronic device and inputting the first inference result into the verification model for probability calculation to obtain the target probability, the method further includes: If the target probability is less than or equal to the acceptable probability, the target prompt word is input into the verification model for inference to generate a second inference result; The second reasoning result is determined as the reasoning result corresponding to the target prompt word.

7. The method according to any one of claims 1 to 4, characterized in that, Before sending the target prompt word to the second electronic device, the method further includes: The validation model is trained by distillation to obtain the inference model, wherein the number of structural layers in the inference model is less than the number of structural layers in the validation model. The installation package of the inference model is sent to the second electronic device, the installation package being used to enable the second electronic device to deploy the inference model.

8. A model inference device, applied to a first electronic device, characterized in that, The device includes: A first sending module is configured to send a target prompt word to a second electronic device, the target prompt word being used to cause an inference model deployed on the second electronic device to perform inference based on the target prompt word in order to generate a first inference result; The first receiving module is configured to receive the first inference result sent by the second electronic device, and input the first inference result into the verification model for probability calculation to obtain the target probability. The target probability is used to represent the probability that the first inference result is the inference result corresponding to the target prompt word. The inference model is determined according to the verification model. The first determining module is used to determine the first inference result as the inference result corresponding to the target prompt word when the target probability is greater than the acceptable probability.

9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; characterized in that the processor is configured to read the program from the memory to implement the steps in the model reasoning method as described in any one of claims 1 to 7.

10. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps in the model reasoning method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps in the model inference method as described in any one of claims 1 to 7.