Methods and systems for multi-user, high-concurrency, and high-throughput inference of large language models
By setting discontinuities in the large language model to obtain user information, statistically analyzing the vocabulary and determining the limiting boundaries, and distilling the user sub-model, the resource shortage problem of the large model under concurrent users is solved, achieving rapid response and resource optimization.
Patent Information
- Application Number
- CN202511102215.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Large models face challenges in practical applications, such as a large number of concurrent users, strict requirements for inference latency, and high resource costs, leading to a shortage of hardware resources.
By setting breakpoints within a preset time range, user information is acquired and the vocabulary database is statistically analyzed. Macro and micro constraint boundaries are determined, sub-models for each user are distilled, interaction channels are constructed, interaction processes are recorded in real time, and auxiliary tasks are input into the large language model.
It improved response speed, reduced real-time resource consumption for large models, optimized concurrent processing architecture, and improved the response speed of most tasks.
Smart Images

Figure CN120598065B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model scheduling technology, specifically a method and system for high-concurrency, high-throughput inference of large language models for multiple users. Background Technology
[0002] After large models are deployed to real-world application scenarios, they need to handle tens of thousands of concurrent requests and rapid response requirements, such as chatbots, search and answering, writing assistants, and customer service systems. Therefore, they face many problems, such as a large number of concurrent user requests, strict requirements for inference latency, and high resource costs leading to a shortage of hardware resources such as GPUs / TPUs. These problems ultimately boil down to efficiency issues. Therefore, how to improve the response efficiency of large models is the technical problem that this invention aims to solve. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for high-concurrency, high-throughput inference of large language models for multiple users, so as to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] A method and system for high-concurrency, high-throughput inference of large language models for multiple users, the method comprising:
[0006] Set breakpoints within a preset time range based on a preset time step, determine the collection period based on the breakpoints, and obtain user information within the collection period.
[0007] All user information within each collection period is identified, words are extracted and counted to obtain a word library; the word library contains word items and frequency items, and the word order in the word library is a preset unique order;
[0008] Determine the macro-level constraint boundaries based on the aforementioned vocabulary database;
[0009] Using each user as an index, the words corresponding to the user information are counted to obtain the word change information for each user. Based on the word change information, the micro-limit boundary is determined in the macro-limit boundary.
[0010] For any user, a sub-model is obtained by distillation based on micro-constraint boundaries. An interaction channel is constructed based on the sub-model, the interaction process is recorded in real time, an auxiliary task is output based on the interaction process, the auxiliary task is input into the large language model, and the output of the large language model is fed back to the interaction channel.
[0011] Furthermore, the step of setting discontinuities within a preset time range according to a preset time step, determining the collection period based on the discontinuities, and obtaining user information within the collection period includes:
[0012] Generate model update instructions periodically and record the time of instruction generation.
[0013] The time range is determined based on the instruction generation time and the preset backtracking time;
[0014] Selecting time points within a preset time range based on a preset time step is called discontinuity points.
[0015] The adjacent discontinuities are used as the endpoints of the time period to determine the data collection period;
[0016] Based on the pre-acquired permissions, retrieve all query messages with time tags sent by all users within each collection period, sort them according to the time tags, and use the sorted query messages as user information.
[0017] Furthermore, the step of identifying, extracting, and statistically analyzing all user information within each collection period to obtain a word database includes:
[0018] For any given collection period, user information for each user within that period is obtained based on preset permissions;
[0019] User information is input into a preset word recognition model to obtain a word list for each user during the collection period; the word list includes word items and their occurrence count.
[0020] Merge each user's word list to generate a word library; the frequency of each word in the word library is amplified by a factor.
[0021] Furthermore, the step of determining the macro-level constraint boundaries based on the word database includes:
[0022] Read the word database for each collection period in sequence;
[0023] Query the word vector of each word in the word database, determine the first expansion radius based on the frequency of the word, and obtain the word vector set;
[0024] Calculate the union of all word vector sets, and simultaneously determine the number of sets corresponding to each word vector in the union;
[0025] The union is used as the macroscopic constraint boundary, and the macroscopic constraint boundary is modified according to a preset quantity threshold; the modification process is a contraction process.
[0026] Furthermore, the step of using each user as an index to count the words corresponding to user information, obtaining word change information for each user, and determining the micro-constraint boundary within the macro-constraint boundary based on the word change information includes:
[0027] For any user, query the word list for that user in each collection period;
[0028] The word lists of adjacent collection periods are compared sequentially, and newly added words are marked. The number of times each word is added is recorded synchronously. The number of additions decreases gradually as the comparison process progresses.
[0029] The word vectors of the query words are used to determine the second expansion radius based on the number of times they are added, and the word vector set of each word is determined.
[0030] Calculate the union of each word vector set to obtain the micro-constraint boundary;
[0031] Calculate the intersection of the micro-level and macro-level constraint boundaries, and use it as the final micro-level constraint boundary.
[0032] Furthermore, the steps of obtaining a sub-model based on micro-constraint boundary distillation for any user, constructing an interaction channel based on the sub-model, recording the interaction process in real time, outputting an auxiliary task based on the interaction process, inputting the auxiliary task into the large language model, and feeding back the output of the large language model to the interaction channel include:
[0033] For any user, the micro-level constraint boundary is input as a constraint condition into the large language model, and the sub-model is obtained by distillation;
[0034] Upon receiving a user's access request, an interaction channel is constructed based on the sub-model;
[0035] The interaction process is recorded in real time. When a recognition failure is detected, the access request is intercepted as an auxiliary task. The criteria for determining recognition failure include failure to recognize and recognition time exceeding a preset time threshold.
[0036] The auxiliary task is input into the large language model, and the output of the large language model is fed back to the interactive channel.
[0037] The present invention also provides a large language model multi-user high-concurrency high-throughput inference system, the system comprising:
[0038] The user information acquisition module is used to set discontinuities within a preset time range according to a preset time step, determine the collection period based on the discontinuities, and acquire user information within the collection period.
[0039] The word and phrase extraction module is used to identify, extract, and count words and phrases from all user information within each collection period to obtain a word and phrase database. The word and phrase database contains word and phrase items and frequency items, and the word and phrase order in the database is a preset unique order.
[0040] A macro-boundary determination module is used to determine macro-restriction boundaries based on the word library;
[0041] The micro-boundary determination module is used to count the words corresponding to user information using each user as an index, obtain word change information for each user, and determine the micro-boundary in the macro-boundary based on the word change information.
[0042] The interaction flow acquisition module is used to obtain a sub-model for any user by distilling based on micro-constraint boundaries, construct an interaction channel based on the sub-model, record the interaction flow in real time, output auxiliary tasks based on the interaction flow, input the auxiliary tasks into the large language model, and feed back the output of the large language model to the interaction channel.
[0043] Furthermore, the user information acquisition module includes:
[0044] The time recording unit is used to generate model update instructions periodically and record the time when the instructions are generated.
[0045] The time range determination unit is used to determine the time range based on the instruction generation time and the preset backtracking time;
[0046] The discontinuity determination unit is used to select time points within a preset time range according to a preset time step, which are called discontinuities.
[0047] The time period generation unit is used to sequentially use adjacent discontinuities as time period endpoints to determine the collection time period;
[0048] The information sorting unit is used to obtain all query information with time tags sent by all users within each collection period based on pre-acquired permissions, sort the query information according to the time tags, and use the sorted query information as user information.
[0049] Furthermore, the word and phrase extraction module includes:
[0050] The information acquisition unit is used to acquire user information of each user within any given collection period based on preset permissions.
[0051] The word list generation unit is used to input user information into a preset word recognition model to obtain a word list for each user during the collection period; the word list includes word items and their occurrence count items;
[0052] The word list merging unit is used to merge the word lists of each user to generate a word library; the frequency of each word in the word library is amplified by a factor.
[0053] The present invention also provides a storage medium storing at least one piece of program code. When the program code is loaded and executed by a processor, it implements the large language model multi-user high-concurrency high-throughput inference method.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] This invention distills a separate sub-model for each user to fulfill most of their interaction needs. The response speed is extremely fast, and the distilled sub-model is generally a local model, which does not occupy the real-time resources of the large model. This effectively converts most of the user's needs to be processed locally, reducing the workload. Except for a slight increase in time for complex tasks, the response speed of most tasks has been improved, and the concurrent processing architecture has been optimized. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention.
[0057] Figure 1 The overall flowchart of the multi-user high-concurrency high-throughput inference method for large language models is shown.
[0058] Figure 2 The structure diagram of a large language model multi-user high-concurrency high-throughput inference system is shown. Detailed Implementation
[0059] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0060] Figure 1 This invention provides a flowchart of a method and system for high-concurrency, high-throughput inference of large language models for multiple users. In this embodiment, a method for high-concurrency, high-throughput inference of large language models for multiple users includes:
[0061] Step S100: Set discontinuity points within a preset time range according to the preset time step, determine the collection period based on the discontinuity points, and obtain user information within the collection period;
[0062] The process of determining the time range is based on a period, with one period being one day. The time step is how often the statistics are collected, usually every few minutes. In practical applications, the current time is used as the end time, and the process counts backwards by one day to obtain the initial time. The time range from the initial time to the end time is one time range. Assuming the time step is five minutes, an interruption point is set every five minutes, resulting in multiple time periods, called collection periods. For each collection period, user information is obtained based on the default permissions. It should be noted that this invention is applied to service providers, and the obtained user information is only for better service provision. When users use this invention, permissions are granted by default.
[0063] Step S200: Identify all user information within each collection period, extract and count words to obtain a word database;
[0064] Using each collection period as an index, user information for all users is statistically analyzed, and all user information is identified. The identification process is mainly a text recognition process, extracting and statistically analyzing the identified words to obtain a word library. Each collection period yields a word set, and all word sets are statistically analyzed to obtain a word library. The word library contains word items and frequency items, which are used to characterize the frequency of occurrence of each word. In addition, the word order in the word library is a preset unique order, that is, the words in each word library adopt a similar sorting method, such as according to the order of the first character in the same dictionary.
[0065] Step S300: Determine the macro-level constraint boundaries based on the vocabulary database;
[0066] The resulting vocabulary database covers the entire time frame, and the words within it reflect the main needs of users within that time frame. By labeling each word, the operation of the large language model can be restricted. After analyzing each word in the vocabulary database, an overall set of constraints can be obtained, called the macro-constraint boundary. These constraints are jointly determined by all users within the time frame.
[0067] Regarding the specific restriction process, under the large language model architecture, it can be clearly restricted to interacting within a text library related to a certain word, which is equivalent to limiting the "search scope".
[0068] Step S400: Using each user as an index, count the words corresponding to the user information to obtain the word change information for each user, and determine the micro-limit boundary in the macro-limit boundary based on the word change information;
[0069] Furthermore, using each user as an index, the words corresponding to the user information are counted to obtain the word set of each user in different collection periods. By analyzing the word sets of adjacent periods, the word changes can be obtained, which is called word change information. Based on the word change information, the micro-limiting boundary is determined. The micro-limiting boundary needs to be limited by the macro-limiting boundary. The micro-limiting boundary is the limiting condition determined by the user's performance in different collection periods.
[0070] Step S500: For any user, obtain a sub-model by distillation based on the micro-constraint boundary, construct an interaction channel based on the sub-model, record the interaction process in real time, output an auxiliary task based on the interaction process, input the auxiliary task into the large language model, and feed back the output of the large language model to the interaction channel.
[0071] Each user corresponds to a micro-constraint boundary. Sub-models are obtained by distillation based on the micro-constraint boundary. The sub-models are equivalent to small models distilled from the large language model, and their response speed is extremely fast. An interaction channel is constructed based on the sub-models. At the same time, the interaction process is recorded in real time. If a problem occurs in the interaction process, the large language model is still used to handle the problem. The problem is the auxiliary task mentioned above. The auxiliary task is input into the large language model, and the output of the large language model is fed back to the interaction channel.
[0072] The principle behind the above process is to distill a separate sub-model for each user to fulfill most of their interaction needs. This results in a fast response time, and the distilled sub-model is generally a local model that does not consume the real-time resources of the large model. This effectively shifts most of the user's needs to be processed locally, reducing the workload. Except for a slight increase in time for complex tasks (due to the pre-identification process of the sub-model, the large model is only triggered if the identification fails), the response speed of most tasks is improved.
[0073] It is worth mentioning that each user's sub-model is generally updated once a period of time, such as once a day based on the time period mentioned above.
[0074] Regarding step S100, the step of setting discontinuities within a preset time range according to a preset time step, determining the collection period based on the discontinuities, and obtaining user information within the collection period includes:
[0075] Generate model update instructions periodically and record the time of instruction generation.
[0076] The time range is determined based on the instruction generation time and the preset backtracking time;
[0077] Selecting time points within a preset time range based on a preset time step is called discontinuity points.
[0078] The adjacent discontinuities are used as the endpoints of the time period to determine the data collection period;
[0079] Based on the pre-acquired permissions, retrieve all query messages with time tags sent by all users within each collection period, sort them according to the time tags, and use the sorted query messages as user information.
[0080] In one example of the technical solution of this invention, the process of collecting user information is described. Model update instructions are generated periodically, the time of instruction generation is recorded, the time range is determined based on the instruction generation time and the preset backtracking time, adjacent discontinuities are used as time period endpoints to determine the collection period, and all query information with time tags sent by users in each collection period is obtained based on the pre-acquired permissions. The query information with the order is the user information.
[0081] Regarding step S200, the step of identifying, extracting and statistically analyzing words and phrases to obtain a word library for all user information within each collection period includes:
[0082] For any given collection period, user information for each user within that period is obtained based on preset permissions;
[0083] User information is input into a preset word recognition model to obtain a word list for each user during the collection period; the word list includes word items and their occurrence count.
[0084] Merge each user's word list to generate a word library.
[0085] In one example of the technical solution of this invention, the process of generating the word library is described. For any given collection period, user information for each user within that period is obtained based on preset permissions. This user information is then input into a preset word recognition model to obtain a word list for each user during that collection period. Each user corresponds to a word list. The word lists of each user are merged to generate the word library. It should be noted that the merging process described above differs slightly from existing merging processes. It exhibits an amplification effect, meaning that the frequency of each word in the word library is amplified by a factor. The process for determining this amplification factor is as follows:
[0086] In the formula, This is the magnification factor. This is a preset constant used to adjust the range of values for the amplification factor. This represents the total number of words in the vocabulary list that contain this word. Indicates the first The number of times the word appears in a list of words containing that word.
[0087] Regarding step S300, the step of determining the macroscopic constraint boundary based on the word library includes:
[0088] Read the word database for each collection period in sequence;
[0089] Query the word vector of each word in the word database, determine the first expansion radius based on the frequency of the word, and obtain the word vector set;
[0090] Calculate the union of all word vector sets, and simultaneously determine the number of sets corresponding to each word vector in the union;
[0091] The union is used as the macroscopic constraint boundary, and the macroscopic constraint boundary is modified according to a preset quantity threshold; the modification process is a contraction process.
[0092] Each data collection period corresponds to a word library. The word vector of each word in the word library is queried. The conversion process of word vectors can use existing conversion models, and many companies have similar models. In addition, the closer the word vectors are, the closer their meanings are generally. This is equivalent to a "semantic space" where similar words are "compressed" together and dissimilar words are "spaced further apart". The first expansion radius is determined based on the frequency of the word. Using the word vector of the word as a benchmark, a range of word vectors can be obtained. All word vectors within the range of word vectors together form a word vector set. The union of all word vector sets is calculated to obtain a larger word vector set. This means that all user interaction information within the data collection period is within the range of this word vector set, which can be used as a macroscopic boundary. The first expansion radius is proportional to the frequency.
[0093] Based on this, the present invention also introduces a correction process. Therefore, if the macroscopic constraint boundary is too large, the resulting sub-model will also be very large. The present invention aims to obtain a sub-model with slightly better performance, which is only used to complete most of the interaction requirements and does not require excessively high performance. Therefore, a correction process is introduced to shrink the macroscopic constraint boundary. The specific shrinking process is as follows:
[0094] The number of sets corresponding to each word vector is determined and concentrated, that is, how many sets of word vectors the word vector appears in. Word vectors with a number less than a preset threshold are deleted (word vectors that only appear in one or two users are determined to be unimportant words in this invention), thereby reducing and concentrating the number of word vectors and narrowing the macro-restriction boundary.
[0095] Regarding step S400, the step of using each user as an index to count the words corresponding to user information, obtaining word change information for each user, and determining the micro-constraint boundary in the macro-constraint boundary based on the word change information includes:
[0096] For any user, query the word list for that user in each collection period;
[0097] The word lists of adjacent collection periods are compared sequentially, and newly added words are marked. The number of times each word is added is recorded synchronously. The number of additions decreases gradually as the comparison process progresses.
[0098] The word vectors of the query words are used to determine the second expansion radius based on the number of times they are added, and the word vector set of each word is determined.
[0099] Calculate the union of each word vector set to obtain the micro-constraint boundary;
[0100] Calculate the intersection of the micro-level and macro-level constraint boundaries, and use it as the final micro-level constraint boundary.
[0101] In one example of the technical solution of this invention, for any user, the word list of that user in each collection period is queried, and the word lists of adjacent collection periods are compared sequentially. Added words are marked, and the number of times each word is added is recorded synchronously. The number of additions decreases at a constant rate as the comparison process progresses. For example, the number of additions decreases by 0.1 for each comparison. The word vectors of the words are queried, and the second expansion radius is determined based on the number of additions. The word vector set of each word is determined. This process yields the word vector set of each user. The union of each word vector set is calculated to obtain the micro-constraint boundary. The intersection of the micro-constraint boundary and the macro-constraint boundary is calculated and constrained within the macro-constraint boundary to obtain the final micro-constraint boundary.
[0102] The second expansion radius is proportional to the number of increases. It is worth mentioning that the first expansion radius is also proportional to the frequency. Both are proportional, and the growth rate (derivative) is a preset value. Generally, the derivative of the second expansion radius with respect to the number of increases is greater than the derivative of the first expansion radius with respect to the frequency.
[0103] Regarding step S500, the steps of obtaining a sub-model based on micro-constraint boundary distillation for any user, constructing an interaction channel based on the sub-model, recording the interaction process in real time, outputting an auxiliary task based on the interaction process, and inputting the auxiliary task into the large language model, feeding back the output of the large language model to the interaction channel include:
[0104] For any user, the micro-level constraint boundary is input as a constraint condition into the large language model, and the sub-model is obtained by distillation;
[0105] Upon receiving a user's access request, an interaction channel is constructed based on the sub-model;
[0106] The interaction process is recorded in real time. When a recognition failure is detected, the access request is intercepted as an auxiliary task. The criteria for determining recognition failure include failure to recognize and recognition time exceeding a preset time threshold.
[0107] The auxiliary task is input into the large language model, and the output of the large language model is fed back to the interactive channel.
[0108] In practical applications, for any user, the micro-constraint boundary is input into the large language model as a constraint condition, and a sub-model is obtained by distillation. After receiving the user's access request, an interaction channel is constructed based on the sub-model, and the interaction process is recorded in real time. When the user cannot be identified or the identification time exceeds the preset time threshold, the content of the access request is intercepted and input into the large language model as an auxiliary task. The output of the large language model is then fed back to the interaction channel.
[0109] Figure 2A structural diagram of a multi-user high-concurrency, high-throughput inference system for a large language model is shown. In a preferred embodiment of the technical solution of the present invention, a multi-user high-concurrency, high-throughput inference system for a large language model is also provided, the system 10 comprising:
[0110] User information acquisition module 11 is used to set discontinuity points within a preset time range according to a preset time step, determine the collection period according to the discontinuity points, and acquire user information within the collection period.
[0111] The word and phrase extraction module 12 is used to identify, extract, and count words and phrases from all user information within each collection period to obtain a word and phrase database; the word and phrase database contains word and phrase items and frequency items, and the word and phrase order in the word and phrase database is a preset unique order;
[0112] Macro boundary determination module 13 is used to determine macro limit boundaries based on the word library;
[0113] The micro-boundary determination module 14 is used to count the words corresponding to user information using each user as an index, obtain word change information for each user, and determine the micro-boundary in the macro-boundary based on the word change information.
[0114] The interaction process acquisition module 15 is used to obtain a sub-model for any user by distilling based on the micro-constraint boundary, construct an interaction channel based on the sub-model, record the interaction process in real time, output auxiliary tasks based on the interaction process, input the auxiliary tasks into the large language model, and feed back the output of the large language model to the interaction channel.
[0115] Furthermore, the user information acquisition module 11 includes:
[0116] The time recording unit is used to generate model update instructions periodically and record the time when the instructions are generated.
[0117] The time range determination unit is used to determine the time range based on the instruction generation time and the preset backtracking time;
[0118] The discontinuity determination unit is used to select time points within a preset time range according to a preset time step, which are called discontinuities.
[0119] The time period generation unit is used to sequentially use adjacent discontinuities as time period endpoints to determine the collection time period;
[0120] The information sorting unit is used to obtain all query information with time tags sent by all users within each collection period based on pre-acquired permissions, sort the query information according to the time tags, and use the sorted query information as user information.
[0121] Specifically, the word and phrase extraction module 12 includes:
[0122] The information acquisition unit is used to acquire user information of each user within any given collection period based on preset permissions.
[0123] The word list generation unit is used to input user information into a preset word recognition model to obtain a word list for each user during the collection period; the word list includes word items and their occurrence count items;
[0124] The word list merging unit is used to merge the word lists of each user to generate a word library; the frequency of each word in the word library is amplified by a factor.
[0125] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for high-concurrency, high-throughput inference of large language models by multiple users, characterized in that: The method includes: Set breakpoints within a preset time range based on a preset time step, determine the collection period based on the breakpoints, and obtain user information within the collection period. All user information within each collection period is identified, words are extracted and counted to obtain a word library; the word library contains word items and frequency items, and the word order in the word library is a preset unique order; Determine the macro-level constraint boundaries based on the aforementioned vocabulary database; Using each user as an index, the words corresponding to the user information are counted to obtain the word change information for each user. Based on the word change information, the micro-limit boundary is determined in the macro-limit boundary. For any user, a sub-model is obtained by distillation based on the micro-constraint boundary. An interaction channel is constructed based on the sub-model, the interaction process is recorded in real time, an auxiliary task is output based on the interaction process, the auxiliary task is input into the large language model, and the output of the large language model is fed back to the interaction channel. The steps of identifying, extracting, and statistically analyzing all user information within each collection period to obtain a word database include: For any given collection period, user information for each user within that period is obtained based on preset permissions; User information is input into a preset word recognition model to obtain a word list for each user during the collection period; the word list includes word items and their occurrence count. Merge each user's word list to generate a word library; the frequency of each word in the word library is amplified by a factor. The step of determining the macro-level constraint boundary based on the vocabulary database includes: Read the word database for each collection period sequentially; Query the word vector of each word in the word database, determine the first expansion radius based on the frequency of the word, and obtain the word vector set; Calculate the union of all word vector sets, and simultaneously determine the number of sets corresponding to each word vector in the union; The union is used as the macroscopic constraint boundary, and the macroscopic constraint boundary is modified according to a preset quantity threshold; the modification process is a shrinking process. The step of using each user as an index to count the words corresponding to user information, obtaining word change information for each user, and determining the micro-constraint boundary within the macro-constraint boundary based on the word change information includes: For any user, query the word list for that user in each collection period; The word lists of adjacent collection periods are compared sequentially, and newly added words are marked. The number of times each word is added is recorded synchronously. The number of additions decreases gradually as the comparison process progresses. The word vectors of the query words are used to determine the second expansion radius based on the number of times they are added, and the word vector set of each word is determined. Calculate the union of each word vector set to obtain the micro-constraint boundary; Calculate the intersection of the micro-level and macro-level constraint boundaries, and use it as the final micro-level constraint boundary.
2. The method for multi-user high-concurrency high-throughput inference of a large language model according to claim 1, characterized in that, The steps of setting discontinuities within a preset time range according to a preset time step, determining the collection period based on the discontinuities, and obtaining user information within the collection period include: Generate model update instructions periodically and record the time of instruction generation. The time range is determined based on the instruction generation time and the preset backtracking time; Selecting time points within a preset time range based on a preset time step is called discontinuity points. The adjacent discontinuities are used as the endpoints of the time period to determine the data collection period; Based on the pre-acquired permissions, retrieve all query messages with time tags sent by all users within each collection period, sort them according to the time tags, and use the sorted query messages as user information.
3. The method for multi-user high-concurrency high-throughput inference of a large language model according to claim 1, characterized in that, The steps of obtaining a sub-model for any user through micro-constraint boundary distillation, constructing an interaction channel based on the sub-model, recording the interaction process in real time, outputting an auxiliary task based on the interaction process, and inputting the auxiliary task into the large language model, feeding back the output of the large language model to the interaction channel include: For any user, the micro-level constraint boundary is input as a constraint condition into the large language model, and the sub-model is obtained by distillation; Upon receiving a user's access request, an interaction channel is constructed based on the sub-model; The interaction process is recorded in real time. When a recognition failure is detected, the access request is intercepted as an auxiliary task. The criteria for determining recognition failure include failure to recognize and recognition time exceeding a preset time threshold. The auxiliary task is input into the large language model, and the output of the large language model is fed back to the interactive channel.
4. A large language model multi-user high-concurrency high-throughput inference system, characterized in that, The system includes: The user information acquisition module is used to set discontinuities within a preset time range according to a preset time step, determine the collection period based on the discontinuities, and acquire user information within the collection period. The word and phrase extraction module is used to identify, extract, and count words and phrases from all user information within each collection period to obtain a word and phrase database. The word and phrase database contains word and phrase items and frequency items, and the word and phrase order in the database is a preset unique order. A macro-boundary determination module is used to determine macro-restriction boundaries based on the word library; The micro-boundary determination module is used to count the words corresponding to user information using each user as an index, obtain word change information for each user, and determine the micro-boundary in the macro-boundary based on the word change information. The interaction process acquisition module is used to obtain a sub-model for any user by distilling based on the micro-constraint boundary, construct an interaction channel based on the sub-model, record the interaction process in real time, output auxiliary tasks based on the interaction process, input the auxiliary tasks into the large language model, and feed back the output of the large language model to the interaction channel. The word and phrase extraction module includes: The information acquisition unit is used to acquire user information of each user within any given collection period based on preset permissions. The word list generation unit is used to input user information into a preset word recognition model to obtain a word list for each user during the collection period; the word list includes word items and their occurrence count items; The word list merging unit is used to merge the word lists of each user to generate a word library; the frequency of each word in the word library is amplified by a factor. The content of determining the macro-level constraint boundary based on the vocabulary database includes: Read the word database for each collection period in sequence; Query the word vector of each word in the word database, determine the first expansion radius based on the frequency of the word, and obtain the word vector set; Calculate the union of all word vector sets, and simultaneously determine the number of sets corresponding to each word vector in the union; The union is used as the macroscopic constraint boundary, and the macroscopic constraint boundary is modified according to a preset quantity threshold; the modification process is a shrinking process. The process of using each user as an index to statistically analyze the words corresponding to user information, obtaining word change information for each user, and determining the micro-level constraint boundary within the macro-level constraint boundary based on the word change information includes: For any user, query the word list for that user in each collection period; The word lists of adjacent collection periods are compared sequentially, and newly added words are marked. The number of times each word is added is recorded synchronously. The number of additions decreases gradually as the comparison process progresses. The word vectors of the query words are used to determine the second expansion radius based on the number of times they are added, and the word vector set of each word is determined. Calculate the union of each word vector set to obtain the micro-constraint boundary; Calculate the intersection of the micro-level and macro-level constraint boundaries, and use it as the final micro-level constraint boundary.
5. The large language model multi-user high-concurrency high-throughput inference system according to claim 4, characterized in that, The user information acquisition module includes: The time recording unit is used to generate model update instructions periodically and record the time when the instructions are generated. The time range determination unit is used to determine the time range based on the instruction generation time and the preset backtracking time; The discontinuity determination unit is used to select time points within a preset time range according to a preset time step, which are called discontinuities. The time period generation unit is used to sequentially use adjacent discontinuities as time period endpoints to determine the collection time period; The information sorting unit is used to obtain all query information with time tags sent by all users within each collection period based on pre-acquired permissions, sort the query information according to the time tags, and use the sorted query information as user information.
6. A storage medium, characterized in that, The storage medium stores at least one piece of program code, which, when loaded and executed by the processor, implements the large language model multi-user high-concurrency high-throughput inference method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Efficient knowledge distillation method for digital twin hydraulic engineering large language model
CN119918617A
Model reasoning method and device suitable for question and answer scene, equipment and medium
CN120387522A