Large and small model collaborative reasoning optimization method and device
Through collaborative inference between small client models and large server models, the problem of low resource utilization in the large model inference method is solved, more efficient computing resource utilization and lower inference delay are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510351879.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, large-model inference methods rely on the computing power of a single device, resulting in low computing resource utilization, high inference delay, and affecting user experience.
Preliminary reasoning is performed through the client's small model, the original reasoning results are obtained, and then the server-side large model is verified to realize collaborative reasoning of large and small models, improve resource utilization and reduce inference delay.
It improves the resource utilization rate of computing devices, reduces inference delay, and improves user experience.
Smart Images

Figure CN120450030A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for optimizing collaborative reasoning of large and small models. Background Art
[0002] With the rapid development of artificial intelligence, large models have demonstrated outstanding performance in fields such as natural language processing and computer vision. However, during the inference phase, large models typically consume significant computing resources and graphics memory, resulting in high inference latency, limited throughput, and extremely demanding performance requirements for computing devices. Traditional large model inference methods often rely on the computing power of a single device, which can easily lead to low computing device resource utilization, high inference latency, and a poor user experience. Summary of the Invention
[0003] The present invention provides a method and device for optimizing collaborative reasoning of large and small models, which is used to solve the defects of the large model reasoning method in the existing technology that relies on the computing power of a single device, has low computing resource utilization and high reasoning delay.
[0004] In a first aspect, the present invention provides a method for optimizing collaborative reasoning between large and small models, comprising: Obtain the target text sequence to be inferred from the user input; Inputting the target text sequence into the target small model to obtain the original inference result of the target text sequence output by the target small model; Sending the target text sequence and the original inference result to a server, wherein the server is used to verify the original inference result based on a large model, determine the inference result of the target text sequence, and send the inference result to the client; Receive the inference result sent by the server.
[0005] In some embodiments, the original inference result includes a plurality of candidate word-grams and a first predicted probability of each candidate word-gram; Correspondingly, inputting the target text sequence into the target small model to obtain the original inference result of the target text sequence output by the target small model includes: Based on the target small model, multiple autoregressions are performed according to the target text sequence to generate multiple candidate word units, and a first prediction probability of each candidate word unit is determined; The verifying the original inference result based on the large model to determine the inference result of the target text sequence includes: Inputting the target text sequence and multiple candidate word units into a large model, and having the large model perform parallel verification on the multiple candidate word units to obtain a second prediction probability for each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, the multiple candidate word-units are sampled using a greedy method and / or a kernel sampling method to obtain the inference result.
[0006] In some embodiments, sending the target text sequence and the original inference result to the server includes: Encoding and encapsulating the target text sequence and the original inference result to construct an HTTP request message; Send the HTTP request message to the server.
[0007] In some embodiments, before inputting the target text sequence into the target mini-model, the method further includes: Obtaining attribute information and resource utilization information of multiple small models deployed on the client; Based on the attribute information and resource utilization information of the multiple small models, the target text sequence is matched with the multiple small models, and according to the matching results, the target small model is determined from the multiple small models.
[0008] In a second aspect, the present invention further provides a large-scale model collaborative reasoning optimization method, which is applied to the server side and includes: Receiving a target text sequence and an original inference result sent by a client, where the original inference result is obtained by inferring the target text sequence by a target small model of the client; Verifying the original inference result based on the large model to determine the inference result of the target text sequence; The inference result is sent to the client.
[0009] In some embodiments, the original inference result includes a plurality of candidate word-grams and a first predicted probability of each candidate word-gram; Correspondingly, verifying the original inference result based on the large model to determine the inference result of the target text sequence includes: Inputting the target text sequence and multiple candidate word units into a large model, and having the large model perform parallel verification on the multiple candidate word units to obtain a second prediction probability for each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, the multiple candidate word-units are sampled using a greedy method and / or a kernel sampling method to obtain the inference result.
[0010] In some embodiments, sending the inference result to the client includes: Encoding and encapsulating the inference result to construct an HTTP response message; Send the HTTP response message to the client.
[0011] In some embodiments, a plurality of small models are deployed on the client, and the target small model is determined by the client after matching the target text sequence with the plurality of small models based on attribute information and resource utilization information of the plurality of small models.
[0012] In a third aspect, a large-scale model collaborative reasoning optimization device includes: A first acquisition unit is used to acquire a target text sequence to be inferred input by a user; An inference unit, configured to input the target text sequence into a target small model and obtain an original inference result of the target text sequence output by the target small model; A first sending unit is configured to send the target text sequence and the original inference result to a server, wherein the server is configured to verify the original inference result based on a large model, determine an inference result of the target text sequence, and send the inference result to a client; The first receiving unit is configured to receive the inference result sent by the server.
[0013] In a fourth aspect, a large-scale model collaborative reasoning optimization device includes: A second receiving unit is configured to receive a target text sequence and an original inference result sent by a client, where the original inference result is obtained by inferring the target text sequence by a target small model of the client; A verification unit, configured to verify the original inference result based on the large model and determine the inference result of the target text sequence; The second sending unit is configured to send the inference result to the client.
[0014] The large and small model collaborative reasoning optimization method and device provided by the present invention perform preliminary reasoning on the target text sequence through the target small model on the client to obtain the original reasoning result, and verify the original reasoning result through the large model on the server to determine the final reasoning result, thereby realizing collaborative reasoning of large and small models across heterogeneous devices, improving the resource utilization of computing devices, reducing reasoning latency, and improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 This is one of the flow charts of the large and small model collaborative reasoning optimization method provided by an embodiment of the present invention.
[0017] Figure 2 This is the second flow chart of the large and small model collaborative reasoning optimization method provided by an embodiment of the present invention.
[0018] Figure 3 This is the third flow chart of the large and small model collaborative reasoning optimization method provided by an embodiment of the present invention.
[0019] Figure 4 This is one of the structural diagrams of the large and small model collaborative reasoning optimization device provided by an embodiment of the present invention.
[0020] Figure 5 This is the second structural diagram of the large and small model collaborative reasoning optimization device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0022] The terms "first," "second," and the like, as used herein, are used to distinguish similar objects, and are not intended to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable, where appropriate, so that embodiments of the present invention can be implemented in an order other than that illustrated or described herein. Furthermore, the objects distinguished by "first" and "second" generally refer to a class of objects and do not limit the number of objects. For example, the first object can be one or more. Furthermore, the term "and / or" as used herein refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0023] Traditional large-model inference methods typically rely on the computing power of a single device. Even with distributed deployment, they lack optimization strategies tailored to the model's characteristics. This approach can easily lead to poor utilization of computing device resources. For example, graphics memory resources may remain idle, while devices with insufficient computing power become performance bottlenecks. Furthermore, the complexity of collaborative computing between heterogeneous devices (such as GPUs and NPUs) further complicates achieving efficient inference.
[0024] However, large and small models often have different computational requirements. Large models typically have higher computational complexity and memory requirements, but require only a small number of computational steps; small models, on the other hand, have lower computational complexity and memory requirements. Existing technologies often fail to fully consider the computational characteristics of heterogeneous devices, resulting in overloaded devices or idle resources.
[0025] To this end, the present invention provides a method and device for optimizing collaborative reasoning between large and small models. This method uses a client-side target small model to perform preliminary reasoning on a target text sequence, obtaining an initial inference result. This initial inference result is then verified using a server-side large model to determine the final inference result, thereby enabling collaborative reasoning between large and small models. This method can improve computing device resource utilization, reduce inference latency for large models, and enhance user experience.
[0026] The present invention provides a collaborative reasoning optimization system for large and small models. The system includes a server and a client. The server can be a platform based on different operating systems and web server software, and the client can be various terminal devices, such as computers and mobile phones. The server is equipped with a computing device with slower computing speed and larger video memory, such as a domestically produced NPU chip; the client is equipped with a computing device with faster computing speed and smaller video memory, such as a high-performance GPU. The server's computing device is deployed with a large model, and the client's computing device is deployed with multiple small models.
[0027] It's understandable that since large models occupy more video memory but require fewer calculations, deploying them on slower computing devices with larger video memory (such as NPUs) can fully utilize the device's video memory resources. Since small models occupy less video memory but require multiple inference calculations, deploying them on faster computing devices with smaller video memory (such as GPUs) can reduce inference latency.
[0028] Figure 1 This is a flow chart of a large and small model collaborative reasoning optimization method provided by an embodiment of the present invention. Figure 1 As shown, a large-scale model collaborative reasoning optimization method is provided, which is applied to the client of the above system and includes the following steps: step 110, step 120, step 130 and step 140. The steps of the method flow are only a possible implementation of the present invention.
[0029] Step 110: Obtain the target text sequence to be inferred input by the user.
[0030] Optionally, the voice input by the user in the current time period is obtained and voice recognition is performed on it to obtain a text sequence; or the text sequence input by the user in the current time period is directly obtained; the text sequence is preprocessed by text cleaning, word segmentation, encoding, etc. to obtain a target text sequence.
[0031] Optionally, the target text sequence to be inferred input by the user is obtained through web applications, desktop applications, mobile applications, etc.
[0032] Step 120: Input the target text sequence into the target small model to obtain the original inference result of the target text sequence output by the target small model.
[0033] Optionally, the target small model can be an existing text classification model, sentiment analysis model, question-answering model, or machine translation model.
[0034] For example, in a question-answering scenario, the target text sequence includes the question and the corresponding context, and the original inference result is the original answer.
[0035] Step 130: Send the target text sequence and the original reasoning result to the server. The server is used to verify the original reasoning result based on the large model, determine the reasoning result of the target text sequence, and send the reasoning result to the client.
[0036] Optionally, the client sends the original inference result to the server via an HTTP request.
[0037] In some embodiments, sending the target text sequence and the original inference result to the server includes: Encode and encapsulate the target text sequence and original inference results to construct an HTTP request message; Send the HTTP request message to the server.
[0038] Optionally, the server is used to encode and encapsulate the inference result, construct an HTTP response message, and send the HTTP response message to the server.
[0039] It's important to note that using HTTP request messages for data transmission offers excellent compatibility. Since most network devices, operating systems, and programming languages support the HTTP protocol, both the client and server can easily process HTTP request messages. This greatly facilitates system integration and deployment, enabling seamless collaboration between different components and facilitating scalability, reducing the complexity of system development and maintenance.
[0040] For example, if you want to add a new data field or change the data processing logic, you only need to make corresponding modifications in the data encoding and encapsulation stage and update the parsing rules on the server side, without making large-scale adjustments to the communication architecture of the entire system.
[0041] Step 140: Receive the inference result sent by the server.
[0042] Optionally, the client receives an HTTP response message sent by the server, parses the HTTP response message, and obtains an inference result.
[0043] It can be understood that by encoding and encapsulating the data to be transmitted, the integrity of the data during transmission can be ensured, which helps to enhance data security; transmitting data between the client and the server through HTTP request messages and HTTP response messages can improve the reliability, flexibility and real-time performance of data transmission, thereby enhancing the user experience.
[0044] In an embodiment of the present invention, preliminary reasoning is performed on the target text sequence through the target small model on the client to obtain the original reasoning result, and the original reasoning result is verified by the large model on the server to determine the final reasoning result, thereby realizing collaborative reasoning of large and small models across heterogeneous devices, improving the resource utilization of computing devices, reducing reasoning latency, and improving user experience.
[0045] In some embodiments, the original inference result includes a plurality of candidate word-grams and a first predicted probability of each candidate word-gram; Correspondingly, the target text sequence is input into the target small model, and the original inference result of the target text sequence output by the target small model is obtained, including: Based on the target small model, multiple autoregressions are performed according to the target text sequence to generate multiple candidate word units and determine the first prediction probability of each candidate word unit; Verify the original inference results based on the large model and determine the inference results of the target text sequence, including: The target text sequence and multiple candidate word units are input into the large model, and the large model verifies the multiple candidate word units in parallel to obtain the second prediction probability of each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, a greedy method and / or a kernel sampling method is adopted to sample multiple candidate word units to obtain an inference result.
[0046] Optionally, first use the target small model to perform N autoregressions to quickly generate N consecutive candidate word units, where N is a natural number greater than 1; then use the large model to perform one autoregression to verify the N candidate word units in parallel; finally, sample the N candidate word units based on the verification results to obtain the final inference result.
[0047] Optionally, the target text sequence is , the target small model is , use HF Transformers library to load the target small model; the target small model generates a word element each time through autoregression and its predicted probability , after N times of autoregression, N word units and their predicted probabilities are generated : ; in, Indicates that the target small model is targeted The predicted probability of each word unit output after the first autoregression, Indicates that the target small model is targeted The output after the Nth autoregression, .
[0048] Optionally, the large model is , use HF Transformers library to load the large model; the input of the large model is sequence And the inference results of the target small model , ; The output of the large model is the predicted probability of the large model : ; in, Indicates that the large model is targeted The predicted probability of each word is output after an autoregression.
[0049] Optionally, a greedy method is used to sample multiple candidate word units to determine whether the first prediction probability and the second prediction probability of each candidate word unit are the same. If so, the candidate word unit is directly added to the final output sequence; otherwise, the candidate word unit is directly rejected.
[0050] Optionally, the greedy method accepts candidate tokens that satisfy the following conditions: ; That is, the i-th word inferred by the large model Equal to the probability of the original inference result of the target small model The largest candidate word.
[0051] Optionally, a kernel sampling method is used to sample multiple candidate word units.
[0052] Optionally, kernel sampling accepts tokens that satisfy the following conditions: ; Here, r is a random number between 0 and 1.
[0053] In some embodiments, before inputting the target text sequence into the target mini-model, the method further includes: Obtain the attribute information and resource utilization information of multiple small models deployed on the client; Based on the attribute information and resource utilization information of multiple small models, the target text sequence is matched with the multiple small models, and the target small model is determined from the multiple small models according to the matching results.
[0054] In some embodiments, before inputting the target text sequence into the target mini-model, the method further includes: Obtain the attribute information and resource utilization information of multiple small models deployed on the client; Based on the attribute information and resource utilization information of multiple small models, the target text sequence is matched with the multiple small models, and the target small model is determined from the multiple small models according to the matching results.
[0055] Attribute information includes but is not limited to: model type (such as text generation models based on the Transformer architecture, semantic understanding models, etc.), model size (number of parameters, number of layers, etc.), the types of tasks that the model is good at handling (such as long text generation, sentiment analysis, question answering, etc.), and the source and field of the model's training data (such as news, science and technology, literature, etc.).
[0056] Optionally, the attribute information of the small model is obtained from the metadata file or related configuration file of the small model.
[0057] The resource utilization information includes, but is not limited to, CPU usage, video memory usage, memory usage, and network bandwidth usage.
[0058] Optionally, resource utilization information of multiple small models can be obtained through the performance monitoring tool of the operating system, the resource monitoring function of the container platform, or specialized resource management software.
[0059] Optionally, the target text sequence is analyzed to determine key information such as the length, topic, and expected processing task of the text.
[0060] Optionally, key information of the target text sequence is matched with attribute information of multiple small models to obtain an initial matching result, and multiple candidate small models are determined from the multiple small models.
[0061] Optionally, a target small model is determined from the multiple candidate small models based on the initial matching result and resource utilization information of the multiple candidate small models.
[0062] For example, among multiple small models, if there is a small model that completely matches the topic of the target text sequence in terms of attributes and its current resource utilization is low, it is determined as the target small model.
[0063] It can be understood that by matching the target text sequence with multiple small models based on the attribute information and resource utilization information of multiple small models, and determining the target small model from multiple small models according to the matching results, the accuracy, flexibility and real-time performance of model matching are improved, resource allocation can be optimized, computing efficiency can be improved, and the stability and reliability of the system can be enhanced.
[0064] Figure 2 The second flow chart of the large and small model collaborative reasoning optimization method provided by the embodiment of the present invention. Figure 2 As shown, a large and small model collaborative reasoning optimization method is provided, which is applied to the above system and includes the following steps: Step 210: Receive the target text sequence and the original inference result sent by the client. The original inference result is obtained by inferring the target text sequence by the client's target small model.
[0065] Among them, the client is used to obtain the voice input by the user in the current time period, perform voice recognition on it, and obtain a text sequence; or directly obtain the text sequence input by the user in the current time period; perform text cleaning, word segmentation, encoding and other preprocessing on the text sequence to obtain the target text sequence.
[0066] In some embodiments, multiple small models are deployed on the client, and the target small model is determined by the client after matching the target text sequence with the multiple small models based on the attribute information and resource utilization information of the multiple small models.
[0067] Optionally, the client is also used to encode and encapsulate the target text sequence and the original inference result, construct an HTTP request message, and send the HTTP request message to the server.
[0068] Optionally, the server receives an HTTP request message sent by the client, parses the HTTP request message, and obtains a target text sequence and an original inference result.
[0069] Step 220: Verify the original reasoning result based on the large model to determine the reasoning result of the target text sequence.
[0070] Optionally, the original inference result includes multiple candidate word-grams.
[0071] Optionally, multiple candidate word-grams are verified based on the large model, and multiple target word-grams are determined from the multiple candidate word-grams to obtain a final inference result.
[0072] Step 230: Send the inference result to the client.
[0073] In some embodiments, sending the inference result to the client includes: Encode and encapsulate the inference results and construct an HTTP response message; Send the HTTP response message to the client.
[0074] Optionally, the client is further configured to receive an HTTP response message sent by the server, parse the HTTP response message, and obtain an inference result.
[0075] In an embodiment of the present invention, preliminary reasoning is performed on the target text sequence through the target small model on the client to obtain the original reasoning result, and the original reasoning result is verified by the large model on the server to determine the final reasoning result, thereby realizing collaborative reasoning of large and small models across heterogeneous devices, improving the resource utilization of computing devices, reducing reasoning latency, and improving user experience.
[0076] In some embodiments, the original inference result includes a plurality of candidate word-grams and a first predicted probability of each candidate word-gram; Correspondingly, the original inference results are verified based on the large model to determine the inference results of the target text sequence, including: The target text sequence and multiple candidate word units are input into the large model, and the large model verifies the multiple candidate word units in parallel to obtain the second prediction probability of each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, a greedy method and / or a kernel sampling method is adopted to sample multiple candidate word units to obtain an inference result.
[0077] It should be noted that small models focus on local linguistic features or pattern matching when generating raw inference results, while large models can re-evaluate candidate tokens from a more comprehensive perspective, combining contextual semantics, grammatical structure, and a wider range of knowledge.
[0078] It is understandable that the original reasoning result contains multiple candidate word units and their respective first prediction probabilities. These probabilities are calculated based on a small model using a local algorithm and may contain certain errors. By inputting the target text sequence and multiple candidate word units into the large model for parallel verification, the second prediction probability of each candidate word unit is obtained, which is equivalent to further confirming the original reasoning result, thereby reducing the error propagation that may occur in the original reasoning result and improving the accuracy of the final reasoning result. By adopting the greedy method and / or the kernel sampling method to sample multiple candidate word units, the reasoning efficiency is improved, the reasoning results are made more reliable, and the interference of low-quality word units is reduced.
[0079] Figure 3 The third flow chart of the large and small model collaborative reasoning optimization method provided by the embodiment of the present invention. Figure 3As shown, a large and small model collaborative reasoning optimization method is provided, which is applied to the above system and includes the following steps: Step 1: The client loads the small model and uses it to infer the target text sequence to obtain the original inference result; Step 2: The client sends the original inference result to the server through HTTP request; Step 3: The server loads the large model, verifies the original inference results through the large model, and determines the final inference results; Step 4: The server sends the inference results to the client.
[0080] Optionally, the client obtains a user request, parses the user request, and obtains a target text sequence.
[0081] Optionally, the server receives the HTTP request sent by the client, parses the HTTP request, and obtains the original inference result.
[0082] Optionally, the client receives an HTTP response sent by the server, parses the HTTP response, and obtains an inference result.
[0083] Optionally, the client uses the HF Transformers library to load the small model, and the server uses the HF Transformers library to load the large model.
[0084] Among them, the HF Transformers library stores a large number of pre-trained models and open source datasets.
[0085] The following describes the large and small model collaborative reasoning optimization device provided by an embodiment of the present invention. The large and small model collaborative reasoning optimization device described below and the large and small model collaborative reasoning optimization method described above can refer to each other.
[0086] Figure 4 One of the structural diagrams of the large and small model collaborative reasoning optimization device provided by the embodiment of the present invention is as follows Figure 4 As shown, the large and small model collaborative reasoning optimization device 400 includes: A first acquisition unit 410 is configured to acquire a target text sequence to be inferred input by a user; The inference unit 420 is used to input the target text sequence into the target small model and obtain the original inference result of the target text sequence output by the target small model; A first sending unit 430 is configured to send the target text sequence and the original inference result to the server. The server is configured to verify the original inference result based on the large model, determine the inference result of the target text sequence, and send the inference result to the client. The first receiving unit 440 is configured to receive the inference result sent by the server.
[0087] Optionally, the original inference result includes a plurality of candidate word-grams and a first prediction probability of each candidate word-gram; Correspondingly, the target text sequence is input into the target small model, and the original inference result of the target text sequence output by the target small model is obtained, including: Based on the target small model, multiple autoregressions are performed according to the target text sequence to generate multiple candidate word units and determine the first prediction probability of each candidate word unit; Verify the original inference results based on the large model and determine the inference results of the target text sequence, including: The target text sequence and multiple candidate word units are input into the large model, and the large model verifies the multiple candidate word units in parallel to obtain the second prediction probability of each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, a greedy method and / or a kernel sampling method is adopted to sample multiple candidate word units to obtain an inference result.
[0088] Optionally, the target text sequence and the original inference result are sent to the server, including: Encode and encapsulate the target text sequence and original inference results to construct an HTTP request message; Send the HTTP request message to the server.
[0089] Optionally, the large and small model collaborative reasoning optimization device 400 further includes: A second acquisition unit is used to acquire attribute information and resource utilization information of multiple small models deployed on the client; The matching unit is used to match the target text sequence with the multiple small models based on the attribute information and resource utilization information of the multiple small models, and determine the target small model from the multiple small models according to the matching results.
[0090] Figure 5 The second structural diagram of the large and small model collaborative reasoning optimization device provided by the embodiment of the present invention is as follows: Figure 5 As shown, the large-small model collaborative reasoning optimization device 500 includes: The second receiving unit 510 is configured to receive a target text sequence and an original inference result sent by a client, where the original inference result is obtained by inferring the target text sequence by a target small model of the client; A verification unit 520 is used to verify the original inference result based on the large model and determine the inference result of the target text sequence; The second sending unit 530 is configured to send the inference result to the client.
[0091] Optionally, the original inference result includes a plurality of candidate word-grams and a first prediction probability of each candidate word-gram; Correspondingly, the original inference results are verified based on the large model to determine the inference results of the target text sequence, including: The target text sequence and multiple candidate word units are input into the large model, and the large model verifies the multiple candidate word units in parallel to obtain the second prediction probability of each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, a greedy method and / or a kernel sampling method is adopted to sample multiple candidate word units to obtain an inference result.
[0092] Optionally, send the inference results to the client, including: Encode and encapsulate the inference results and construct an HTTP response message; Send the HTTP response message to the client.
[0093] Optionally, multiple small models are deployed on the client, and the target small model is determined by the client after matching the target text sequence with the multiple small models based on the attribute information and resource utilization information of the multiple small models.
[0094] It should be noted here that the large and small model collaborative reasoning optimization device provided by the embodiment of the present invention can implement all the method steps implemented by the above-mentioned large and small model collaborative reasoning optimization method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0096] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A collaborative reasoning optimization method for large and small models, characterized in that: Applied to the client, including: Obtain the target text sequence to be inferred from the user input; Inputting the target text sequence into the target small model to obtain the original inference result of the target text sequence output by the target small model; Sending the target text sequence and the original inference result to a server, wherein the server is used to verify the original inference result based on a large model, determine the inference result of the target text sequence, and send the inference result to the client; Receive the inference result sent by the server.
2. The large and small model collaborative reasoning optimization method according to claim 1 is characterized in that: The original inference result includes a plurality of candidate word-grams and a first prediction probability of each candidate word-gram; Correspondingly, inputting the target text sequence into the target small model to obtain the original inference result of the target text sequence output by the target small model includes: Based on the target small model, multiple autoregressions are performed according to the target text sequence to generate multiple candidate word units, and a first prediction probability of each candidate word unit is determined; The verifying the original inference result based on the large model to determine the inference result of the target text sequence includes: Inputting the target text sequence and multiple candidate word units into a large model, and having the large model perform parallel verification on the multiple candidate word units to obtain a second prediction probability for each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, the multiple candidate word-units are sampled using a greedy method and / or a kernel sampling method to obtain the inference result.
3. The large and small model collaborative reasoning optimization method according to claim 1 is characterized in that: The sending of the target text sequence and the original inference result to the server includes: Encoding and encapsulating the target text sequence and the original inference result to construct an HTTP request message; Send the HTTP request message to the server.
4. The large and small model collaborative reasoning optimization method according to claim 1 is characterized in that: Before inputting the target text sequence into the target small model, the method further includes: Obtaining attribute information and resource utilization information of multiple small models deployed on the client; Based on the attribute information and resource utilization information of the multiple small models, the target text sequence is matched with the multiple small models, and according to the matching results, the target small model is determined from the multiple small models.
5. A collaborative reasoning optimization method for large and small models, characterized in that: Applied to the server side, including: Receiving a target text sequence and an original inference result sent by a client, where the original inference result is obtained by inferring the target text sequence by a target small model of the client; Verifying the original inference result based on the large model to determine the inference result of the target text sequence; The inference result is sent to the client.
6. The large and small model collaborative reasoning optimization method according to claim 5 is characterized in that: The original inference result includes a plurality of candidate word-grams and a first prediction probability of each candidate word-gram; Correspondingly, verifying the original inference result based on the large model to determine the inference result of the target text sequence includes: Inputting the target text sequence and multiple candidate word units into a large model, and having the large model perform parallel verification on the multiple candidate word units to obtain a second prediction probability for each candidate word unit; Based on the first prediction probability and the second prediction probability of each candidate word, the multiple candidate word-units are sampled using a greedy method and / or a kernel sampling method to obtain the inference result.
7. The large and small model collaborative reasoning optimization method according to claim 5 is characterized in that: The sending the inference result to the client includes: Encoding and encapsulating the inference result to construct an HTTP response message; Send the HTTP response message to the client.
8. The large and small model collaborative reasoning optimization method according to claim 5 is characterized in that: A plurality of small models are deployed on the client, and the target small model is determined by the client after matching the target text sequence with the plurality of small models based on attribute information and resource utilization information of the plurality of small models.
9. A collaborative reasoning optimization device for large and small models, characterized in that: include: A first acquisition unit is used to acquire a target text sequence to be inferred input by a user; An inference unit, configured to input the target text sequence into a target small model and obtain an original inference result of the target text sequence output by the target small model; A first sending unit is configured to send the target text sequence and the original inference result to a server, wherein the server is configured to verify the original inference result based on a large model, determine the inference result of the target text sequence, and send the inference result to a client; The first receiving unit is configured to receive the inference result sent by the server.
10. A collaborative reasoning optimization device for large and small models, characterized in that: include: A second receiving unit is configured to receive a target text sequence and an original inference result sent by a client, where the original inference result is obtained by inferring the target text sequence by a target small model of the client; A verification unit, configured to verify the original inference result based on the large model and determine the inference result of the target text sequence; The second sending unit is configured to send the inference result to the client.