A domain large model-based knowledge base system

By optimizing the user allocation method, and adopting either profile-based allocation or load-based allocation according to the stable state and load status of user profiles, the problem of uneven server load in the knowledge base system was solved, improving operating efficiency and processing speed.

CN120596252BActive Publication Date: 2026-02-27BEIJING XINRUIXIANGTONG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510679206.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2026-02-27
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

In existing knowledge base systems, a single server allocation method is insufficient to meet user needs, resulting in unbalanced load and affecting operational efficiency.

Method used

The data analysis unit determines the stable state of user profiles based on the continuous proportion of user categories and the number of requests. Combined with the early warning ratio coefficient and load status, it optimizes server resource allocation by adopting profile allocation or load allocation methods.

Benefits of technology

It improves the operational efficiency of the knowledge base system by reducing server load and increasing processing speed and efficiency through precise user allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596252B_ABST
    Figure CN120596252B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of knowledge bases, in particular to a knowledge base system based on a domain large model, which comprises a data processing unit, which carries out data caching and data processing; a data analysis unit, which determines a user portrait stable state according to a user category continuous proportion and a user request frequency, and determines a user allocation mode according to the user portrait stable state; a portrait allocation unit, which determines a server state according to a pre-warning proportion coefficient and a first-class user dispersion, and determines a portrait allocation mode according to the server state; and a load allocation unit, which determines a load state according to a load balancing coefficient and a load balance coefficient, and determines a load allocation mode according to the load state. The application optimizes the user allocation mode, improves the allocation accuracy of the server, and effectively improves the overall operation efficiency of the knowledge base system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of knowledge base, and particularly relates to a knowledge base system based on a domain large model. BACKGROUND

[0002] The knowledge base system based on the domain large model is an intelligent system combining a large language model (LLM) and specific domain knowledge, and is widely applied to the fields of finance, medicine, enterprise knowledge management and the like. At present, with the continuous increase in the number of users in the application of the knowledge base system, a single server distribution mode is difficult to meet the use demand, the server load is increasingly improved, and the conditions of uneven load distribution or load overload are prone to occur, thereby affecting the operation efficiency of the knowledge base. Therefore, how to improve the server distribution efficiency is a technical problem to be solved by the person skilled in the art.

[0003] A knowledge base-based question and answer processing system is disclosed in Chinese Patent Publication No. CN118093843B. The technical solution constructs a question and answer knowledge base by setting a basic data unit, a data analysis module analyzes and obtains a preset question training set and a preset answer training set under multiple consultations of a consulting user in big data, and a preset service unit splits an enterprise question and answer model to obtain a fine-tuning question and answer small model under multiple consultations and a small question and answer knowledge base corresponding thereto. In this way, the traditional preset reply answer is not comprehensive, the fine-tuning model trained is used to reply to the consulting user with few question and answer service times, the accuracy and backtracking rate of the reply answer are improved, and the consulting user with more consultation times is replied to by the question and answer model. It can be seen that the above technical solution has the following problems: different models are selected for the user to reply according to the number of question and answer services to reduce the load of the server carrying the question and answer large model of the enterprise. However, the accuracy of the fine-tuning question and answer small model needs to be supported by a large amount of data training, and when the score of the answer of the consulting question by the fine-tuning question and answer small model is less than a preset score threshold, the consulting question needs to be transmitted to the question and answer service unit again, which reduces the processing efficiency of the consulting question of the user. SUMMARY

[0004] Therefore, the present application provides a knowledge base system based on a domain large model to overcome the problem that the single user distribution mode of the prior art knowledge base system is difficult to meet the actual processing efficiency requirement.

[0005] To achieve the above-mentioned purpose, the present application provides a knowledge base system based on a domain large model, comprising:

[0006] A data processing unit comprising a plurality of servers for data caching and data processing;

[0007] a data analysis unit connected with the data processing unit, configured to determine a user portrait stable state according to a user category continuous proportion and a user request number, and determine a user allocation mode as portrait allocation or load allocation according to the user portrait stable state;

[0008] a portrait allocation unit connected with the data processing unit and the data analysis unit, configured to determine a server state according to an early warning proportion coefficient and a one-category user dispersion degree, and determine a portrait allocation mode as sequential allocation or user allocation according to a cache correlation degree according to the server state;

[0009] a load allocation unit connected with the data processing unit and the data analysis unit, configured to determine a load state according to a load balancing coefficient and a load margin coefficient, and determine a load allocation mode as gradient allocation or user allocation according to an estimated processing time length according to the load state.

[0010] Further, the data analysis unit determines the user allocation mode according to the user portrait stable state in response to a preset allocation condition;

[0011] If the user portrait stable state is that the user category continuous proportion is greater than a preset user category continuous proportion and the user request number is greater than a preset user request number, the user allocation mode is portrait allocation.

[0012] If the user portrait stable state is that the user category continuous proportion is less than or equal to the preset user category continuous proportion or the user request number is less than or equal to the preset user request number, the user allocation mode is load allocation.

[0013] The preset allocation condition is a new connection request.

[0014] Further, the data analysis unit determines the user category according to a text recognition difficulty coefficient and an operation frequency coefficient.

[0015] If the text recognition difficulty coefficient is greater than a preset text recognition difficulty coefficient or the operation frequency coefficient is greater than a preset operation frequency coefficient, the data analysis unit determines that the user category is a one-category user.

[0016] If the text recognition difficulty coefficient is less than or equal to the preset text recognition difficulty coefficient and the operation frequency coefficient is less than or equal to the preset operation frequency coefficient, the data analysis unit determines that the user category is a two-category user.

[0017] The text recognition difficulty coefficient is determined according to a domain keyword distribution coefficient and a picture-text combination coefficient.

[0018] Further, the portrait allocation unit determines a number of early warning proportion coefficients according to the proportion coefficients corresponding to each server in response to a portrait allocation condition, and determines the portrait allocation mode according to the server state.

[0019] If the number of early warning proportionality coefficients is less than the preset number of early warning proportionality coefficients or the dispersion of the first type of users is less than the preset dispersion of the first type of users, it is determined that the portrait allocation mode is sequential allocation.

[0020] If the number of early warning proportionality coefficients is greater than or equal to the preset number of early warning proportionality coefficients and the dispersion of the first type of users is greater than or equal to the preset dispersion of the first type of users, it is determined that the portrait allocation mode is user allocation according to cache relevance.

[0021] The portrait allocation condition is that the data analysis unit determines the user allocation mode to be portrait allocation.

[0022] Further, the portrait allocation unit, in response to a first allocation condition, obtains the user category of the user to be allocated and the compensation coefficients corresponding to each server, and records the server with the largest compensation coefficient as the allocation server corresponding to the user to be allocated.

[0023] The first allocation condition is that the portrait allocation unit determines the portrait allocation mode to be sequential allocation, and the compensation coefficient has a positive correlation with the number of related user categories.

[0024] Further, the portrait allocation unit, in response to a second allocation condition, detects the cache relevance between the request task corresponding to the user to be allocated and each server, and records the server with the largest cache relevance as the allocation server corresponding to the user to be allocated.

[0025] The second allocation condition is that the portrait allocation unit determines the portrait allocation mode to be user allocation according to cache relevance.

[0026] Further, the portrait allocation unit is provided with a cache relevance calculation strategy, which includes determining a keyword quantity ratio according to the number of general keywords and the number of exclusive keywords in the request task corresponding to the user to be allocated, and determining a cache relevance confirmation mode according to the keyword quantity ratio.

[0027] If the keyword quantity ratio is greater than a preset keyword quantity ratio, the cache relevance is determined according to a user preference coefficient.

[0028] If the keyword quantity ratio is less than or equal to the preset keyword quantity ratio, the cache relevance is determined according to an exclusive keyword similarity.

[0029] Further, the portrait allocation unit determines the exclusive keyword and the general keyword according to a combined use uniformity.

[0030] If the combined use uniformity is less than or equal to a preset combined use uniformity, the keyword is a general keyword.

[0031] If the combined use uniformity is greater than the preset combined use uniformity, the keyword is an exclusive keyword.

[0032] Further, the load distribution unit determines the load distribution mode according to the load state in response to the load distribution condition;

[0033] If the load state is that the load balance coefficient is greater than the preset load balance coefficient, or the load balance coefficient is less than or equal to the preset load balance coefficient and the load margin coefficient is greater than the preset load margin coefficient, the load distribution mode is gradient distribution.

[0034] If the load state is that the load balance coefficient is less than or equal to the preset load balance coefficient and the load margin coefficient is less than or equal to the preset load margin coefficient, the load distribution mode is user distribution according to the estimated processing time.

[0035] Further, the load distribution unit determines the estimated processing time according to the user repetition frequency corresponding to each margin server in response to the third distribution condition, and records the margin server with the minimum estimated processing time as the distribution server corresponding to the user to be distributed.

[0036] The third distribution condition is that the load distribution unit determines the load distribution mode as user distribution according to the estimated processing time.

[0037] Compared with the prior art, the beneficial effects of the present application are that the user portrait stable state is determined according to the user category continuous proportion and the user request number in the technical scheme of the present application, the user portrait stable state reflects the stability and reliability of the user portrait, and different user distribution modes are selected correspondingly, so that the user distribution mode can effectively meet the actual user situation, avoid the difficulty of effective distribution of a single user distribution mode, and further improve the operation efficiency of the knowledge base system.

[0038] Further, in the technical scheme of the present application, the user category is determined according to the text recognition difficulty coefficient and the operation frequency coefficient, the operation characteristics of the user are determined according to the text recognition difficulty coefficient and the operation frequency coefficient, the user category is determined, and the accuracy of the user distribution mode of the portrait distribution is improved, the balance degree of the user category corresponding to each server is improved, and the processing efficiency of each server in the knowledge base system is improved.

[0039] Further, in the technical scheme of the present application, the user is distributed according to the cache correlation under the second distribution condition, the server with greater correlation between the cache data and the user to be distributed is selected for distribution according to the cache correlation, so as to improve the processing efficiency of the subsequent problems of the user to be distributed, further improve the processing speed of the server, and reduce the load of the server.

[0040] Further, in the technical scheme of the present application, the keyword quantity ratio is determined according to the number of general keywords and exclusive keywords in the request task corresponding to the user to be allocated, and the cache relevance confirmation mode is determined according to the keyword quantity ratio, wherein the generality of the keywords is uniformly reflected through the combined use of the keywords, and the keywords are correspondingly divided into exclusive keywords and general keywords, so that the selected cache relevance confirmation mode is more in line with the actual scene, avoiding the problem of large determination error of a single cache relevance confirmation mode, thereby improving the user allocation accuracy.

[0041] Further, in the technical scheme of the present application, the estimated processing time is determined according to the user repetition frequency corresponding to each residual server, so that the determination accuracy of the estimated processing time is more accurate, thereby improving the server allocation accuracy and further improving the user allocation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0042] Fig. 1 The unit connection graph of the knowledge base system based on the field large model of the present application;

[0043] Fig. 2 The flowchart of the present application for determining the user allocation mode according to the user portrait stable state;

[0044] Fig. 3 The flowchart of the present application for determining the user category according to the text recognition difficulty coefficient and the operation frequency coefficient;

[0045] Fig. 4 The flowchart of the present application for determining the load allocation mode according to the load state. DETAILED DESCRIPTION

[0046] In order to make the purpose and advantages of the present application more clear and explicit, the present application will be further described below in combination with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the present application.

[0047] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application, and are not intended to limit the protection scope of the present application.

[0048] In addition, it should be further pointed out that, in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection" and "connection" should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication inside two elements. Those skilled in the art can understand the specific meaning of the above terms in the present application according to the specific circumstances.

[0049] Referring to Figs. 1 to 4 As shown in the drawings, the application provides a knowledge base system based on a domain large model, characterized in that it comprises:

[0050] a data processing unit comprising a plurality of servers for data caching and data processing;

[0051] a data analysis unit connected to the data processing unit for determining a user portrait stability state according to a user category continuous proportion and a user request frequency, and determining a user allocation mode as portrait allocation or load allocation according to the user portrait stability state;

[0052] a portrait allocation unit connected to the data processing unit and the data analysis unit for determining a server state according to a pre-warning proportion coefficient and a one-class user dispersion, and determining a portrait allocation mode as sequential allocation or user allocation according to a cache correlation degree according to the server state;

[0053] a load allocation unit connected to the data processing unit and the data analysis unit for determining a load state according to a load balancing coefficient and a load margin coefficient, and determining a load allocation mode as gradient allocation or user allocation according to an estimated processing time length according to the load state.

[0054] The application is applied to domain knowledge question answering, and the knowledge base system data processing unit of the application is connected to a user end, and a user in the user end can upload a data request, the data request is a user uploaded request task, the request task includes a request text and a corresponding image, the knowledge base system allocates the request text and the corresponding image to a server for processing and generates a reply result corresponding to the request text and the corresponding image to return to the user end, how the server applies a domain large model to generate a reply result for the request text and the corresponding image proposed by the user is a content mastered by those skilled in the art, and is not repeated here, the new connection request exists a user uploaded data request and the user is not allocated to the server, and the user is recorded as a user to be allocated. The application also has a history record, and a single history record at least includes a user category continuous proportion, a user request frequency, a domain keyword distribution density, a domain keyword distribution coefficient, an image-text combination coefficient, an operation frequency coefficient, a pre-warning proportion coefficient, a one-class user dispersion, a general keyword quantity, an exclusive keyword quantity and a combination use uniformity degree, and the history record is correspondingly provided with a qualified mark, and the qualified mark records whether the history record meets the needs of the management personnel, wherein whether a history record meets the needs of the management personnel can be determined according to whether the server crashes in a plurality of monitoring periods, wherein it can be understood that whether the history record meets the needs of the management personnel is determined according to the server performance indicators (such as the number of crashes, user evaluation statistics) set by the management personnel, which is a content mastered by those skilled in the art, and is not limited here.

[0055] Specifically, the data analysis unit determines the user allocation mode according to the user portrait stable state in response to a preset allocation condition;

[0056] If the user portrait stable state is that the user category continuous proportion is greater than a preset user category continuous proportion and the user request number is greater than a preset user request number, the user allocation mode is portrait allocation;

[0057] If the user portrait stable state is that the user category continuous proportion is less than or equal to the preset user category continuous proportion or the user request number is less than or equal to the preset user request number, the user allocation mode is load allocation;

[0058] The preset allocation condition is a new connection request.

[0059] The user categories corresponding to each historical request of the user to be allocated are obtained, the user categories corresponding to the historical requests are sorted in the order from early to late to obtain a request sequence, the request sequence is divided into a plurality of continuous paragraphs, the user categories in any continuous paragraph are all the same, and any adjacent user category of the continuous paragraphs is different from the user categories in the continuous paragraphs, and the user category continuous proportion = the number of user categories of the continuous paragraph with the maximum number of corresponding user categories / the total number of user categories corresponding to the historical requests;

[0060] The number of historical requests corresponding to the user to be allocated is recorded as the user request number;

[0061] The historical requests are each data request performed by the user to be allocated before the current moment;

[0062] The preset user category continuous proportion and the preset user request number can be set by the management personnel according to the actual scene. It can be understood that the greater the demand of the management personnel for the determination accuracy of the user allocation mode, the greater the preset user category continuous proportion and the preset request number. A value taking method is provided, the user category continuous proportion and the user request number corresponding to the historical record meeting the demand of the management personnel are extracted, the abnormal values corresponding to the user category continuous proportion and the user request number are removed respectively, and the average values of the user category continuous proportion and the user request number after removing the abnormal values are recorded as the preset user category continuous proportion and the preset user request number respectively. The method for removing abnormal values includes but is not limited to 3σ criterion method or IQR method. A value taking method of the preset user category continuous proportion and the preset user request number is provided, the preset user category continuous proportion is 80%, and the preset user request number is 15 times.

[0063] Specifically, the data analysis unit determines the user category according to the text recognition difficulty coefficient and the operation frequency coefficient;

[0064] If the text recognition difficulty coefficient is greater than the preset text recognition difficulty coefficient or the operation frequency coefficient is greater than the preset operation frequency coefficient, the data analysis unit determines the user category as a first type of user;

[0065] If the text recognition difficulty coefficient is less than or equal to the preset text recognition difficulty coefficient and the operation frequency coefficient is less than or equal to the preset operation frequency coefficient, the data analysis unit determines the user category as a second type of user;

[0066] The text recognition difficulty coefficient is determined according to a field keyword distribution coefficient and an image-text combination coefficient.

[0067] The field keyword distribution coefficient is an average value of sub-distribution coefficients of request texts of each historical request corresponding to the user to be allocated, and for a single request text, the corresponding sub-distribution coefficient = field keyword distribution density * α1 + field keyword quantity * α2.

[0068] The field keyword quantity is the total number of field keywords appearing in the request text.

[0069] The confirmation method of the field keyword distribution density is to detect the number of characters A between the first appearing field keyword and the last appearing field keyword in the request text, and the field keyword distribution density = total number of characters in the request text / A.

[0070] Wherein, α1 is a first weight coefficient, and α2 is a second weight coefficient, the values of the first weight coefficient and the second weight coefficient can be determined according to historical records, it can be understood that the management personnel can determine the contribution degree of the field keyword distribution density to the sub-recognition difficulty according to the historical records, and the corresponding values are taken, the greater the contribution degree, the greater the value of the corresponding weight coefficient α1, and the value of α2 is the same, a value of the weight coefficient is provided, α1 = 0.6, α2 = 0.4.

[0071] The confirmation method of the image-text combination coefficient is to detect the number of corresponding images of the request texts of each historical request corresponding to each user to be allocated, and the average value of the number of images is recorded as the image-text combination coefficient.

[0072] The text recognition difficulty coefficient = field keyword distribution coefficient / preset field keyword distribution coefficient + image-text combination coefficient / preset image-text combination coefficient.

[0073] The preset domain keyword distribution coefficient and preset image-text combination coefficient can be set by the administrator according to actual needs. It is understood that the greater the administrator's acceptance of the difficulty of text recognition, the greater the values ​​of the preset domain keyword distribution coefficient and preset image-text combination coefficient. One method is to extract the corresponding domain keyword distribution coefficient and image-text combination coefficient from the historical records that meet the administrator's needs, remove outliers, and record the average values ​​of the domain keyword distribution coefficient and image-text combination coefficient after removing outliers as the preset domain keyword distribution coefficient and preset image-text combination coefficient, respectively.

[0074] The operation frequency coefficient is determined by detecting the user's continuous operation behavior during the user's historical operation process. Continuous operation behavior is defined as the user making at least two data requests with the interval between adjacent data requests being less than a preset interval. The maximum number of data requests corresponding to continuous operation behavior is recorded as the operation frequency coefficient. The preset interval can be set by the administrator according to actual needs. It can be understood that the smaller the interval, the more frequent the user operation. Therefore, the greater the administrator's acceptance of the user's operation frequency, the larger the preset interval. One preset interval value is provided: preset interval = 5 minutes.

[0075] The preset text recognition difficulty coefficient and preset operation frequency coefficient can be set by administrators according to the actual scenario. It is understood that the higher the administrator's acceptance of the user's task processing difficulty, the lower the preset text recognition difficulty coefficient and preset operation frequency coefficient will be. A method for setting these values ​​is provided, which extracts the corresponding text recognition difficulty coefficient and operation frequency coefficient from the historical records that meet the administrator's needs, removes outliers from the text recognition difficulty coefficient and operation frequency coefficient respectively, and records the average values ​​of the text recognition difficulty coefficient and operation frequency coefficient after removing outliers as the preset text recognition difficulty coefficient and preset operation frequency coefficient respectively. The methods for removing outliers include, but are not limited to, the 3σ criterion method or the IQR method. The method for setting the preset text recognition difficulty coefficient and preset operation frequency coefficient is provided.

[0076] Specifically, the image allocation unit responds to the image allocation conditions, determines the number of warning ratio coefficients based on the ratio coefficients corresponding to each server, and determines the image allocation method based on the server status.

[0077] If the number of warning ratio coefficients is less than the preset number of warning ratio coefficients or the dispersion of a type of user is less than the preset dispersion of a type of user, then the profile allocation method is determined to be sequential allocation.

[0078] If the number of early warning proportionality coefficients is greater than or equal to the preset number of early warning proportionality coefficients and the one-class user dispersion is greater than or equal to the preset one-class user dispersion, it is determined that the portrait allocation mode is user allocation according to the cache relevance.

[0079] The portrait allocation condition is determined by the data analysis unit to determine the user allocation mode as portrait allocation.

[0080] For a single server, the corresponding proportionality coefficient = the number of one-class users of the server / the number of two-class users of the server, and if the proportionality coefficient is greater than the preset proportionality coefficient, the proportionality coefficient is recorded as an early warning proportionality coefficient.

[0081] The value of the preset proportionality coefficient can be set by the administrator according to the actual scene. It can be understood that the greater the demand of the administrator for the processing efficiency of the server, the smaller the preset proportionality coefficient. A value of the preset proportionality coefficient is provided, and the preset proportionality coefficient = 1.5.

[0082] The confirmation method of the one-class user dispersion is to detect the number of one-class users corresponding to each server, and the one-class user dispersion is recorded as S. The calculation method of S is as follows:

[0083]

[0084] Wherein, i = 1, 2, 3, …, Mz, the Mz is the number of servers, Mi is the number of one-class users corresponding to the i-th server, and M0 is the average value of the number of one-class users corresponding to all servers. Wherein, the value order of the server corresponding to i is randomly set, and the order of i does not affect the calculation result.

[0085] The values of the preset number of early warning proportionality coefficients and the preset one-class user dispersion can be set by the administrator according to the actual scene. It can be understood that the current load balancing degree of the server is reflected through the number of early warning proportionality coefficients and the one-class user dispersion. The greater the demand of the administrator for the load processing efficiency and balance of the server, the smaller the preset number of early warning proportionality coefficients and the value of the preset one-class user dispersion. A value is provided. The corresponding number of early warning proportionality coefficients and one-class user dispersion in the historical record meeting the user demand are extracted, and the abnormal values corresponding to the number of early warning proportionality coefficients and the one-class user dispersion are removed respectively. The average values of the number of early warning proportionality coefficients and the one-class user dispersion after removing the abnormal values are recorded as the preset number of early warning proportionality coefficients and the preset one-class user dispersion respectively. A value of the preset number of early warning proportionality coefficients and the preset one-class user dispersion is provided.

[0086] Specifically, the image allocation unit acquires the user category of the user to be allocated and the compensation coefficient corresponding to each server in response to the first allocation condition, and records the server with the largest compensation coefficient as the allocation server corresponding to the user to be allocated.

[0087] The first allocation condition is that the image allocation unit determines the image allocation mode as sequential allocation, and the compensation coefficient has a positive correlation with the number of related user categories.

[0088] When the user category is a first user category, the number of related user categories is the number of second user categories, that is, the compensation coefficient corresponding to the user to be allocated and a server is the number of second user categories in the server.

[0089] When the user category is a second user category, the number of related user categories is the number of first user categories, that is, the compensation coefficient corresponding to the user to be allocated and a server is the number of first user categories in the server.

[0090] Specifically, the image allocation unit detects the cache correlation between the request task corresponding to the user to be allocated and each server in response to the second allocation condition, and records the server with the largest cache correlation as the allocation server corresponding to the user to be allocated.

[0091] The second allocation condition is that the image allocation unit determines the image allocation mode as user allocation according to the cache correlation.

[0092] Specifically, the image allocation unit is provided with a cache correlation calculation strategy, which includes determining a keyword quantity ratio according to the number of general keywords and the number of exclusive keywords in the request task corresponding to the user to be allocated, and determining a cache correlation confirmation mode according to the keyword quantity ratio.

[0093] If the keyword quantity ratio is greater than a preset keyword quantity ratio, the cache correlation is determined according to a user preference coefficient.

[0094] If the keyword quantity ratio is less than or equal to the preset keyword quantity ratio, the cache correlation is determined according to an exclusive keyword similarity.

[0095] The keyword quantity ratio = general keyword quantity / exclusive keyword quantity. It can be understood that the higher the requirement of the administrator on the matching degree of the request task corresponding to the user to be allocated and the cache correlation, the smaller the preset keyword quantity ratio. A value mode of the preset keyword quantity ratio is provided, the general keyword quantity and the exclusive keyword quantity corresponding to the historical record satisfying the requirement of the administrator are acquired, the abnormal values of the general keyword quantity and the exclusive keyword quantity are removed respectively, and the average values of the general keyword quantity and the exclusive keyword quantity after removing the abnormal values are recorded as the preset general keyword quantity and the preset exclusive keyword quantity respectively. Therefore, the preset keyword quantity ratio = preset general keyword quantity / preset exclusive keyword quantity.

[0096] When determining the cache relevance according to the user preference coefficient, the user preference coefficient is recorded as the cache relevance, and for a server, the confirmation manner of the corresponding user preference coefficient is to obtain the historical browsing attention degree of each type of data of the to-be-assigned user, if only the historical browsing attention degree of one type of data is greater than the preset historical browsing attention degree, then the data type with the historical browsing attention degree greater than the preset historical browsing attention degree is recorded as the preferred type, and the cache data amount of the preferred type in the server is recorded as the preference coefficient corresponding to the server.

[0097] If the historical browsing attention degrees of two or more types of data are greater than the preset historical browsing attention degree, then the difference between the maximum value and the minimum value of the data amount of the cache data corresponding to each data type in the server is recorded as the preference coefficient corresponding to the server.

[0098] The preset historical browsing attention degree is the average value of the historical browsing attention degrees of three types of data.

[0099] The content in the reply result is divided into code type data, character type data and image type data according to data types, the confirmation manner of the historical browsing attention degree is to count the browsing time of each data type in the reply result when the to-be-assigned user browses the reply result for a certain number of times, the browsing time is recorded as the historical browsing attention degree corresponding to the data type, and the browsing time of data of any data type is the total time of the display for the data of the data type, which is easily understood by those skilled in the art and will not be described here. The number of times of counting the recent browsing of the reply result by the to-be-assigned user can be set according to actual needs. The greater the accuracy requirement of the historical browsing attention degree by the management personnel, the greater the number.

[0100] When determining the cache relevance according to the exclusive keyword similarity, the exclusive keyword similarity is recorded as the cache relevance, and for a server, the confirmation manner of the corresponding exclusive keyword similarity is to detect the number of exclusive keywords in each cache data corresponding to the server that are the same as the exclusive keywords in the request text, and the average value of the number of exclusive keywords in each cache data that are the same as the exclusive keywords in the request text is recorded as the exclusive keyword similarity.

[0101] Specifically, the image allocation unit determines the exclusive keywords and the general keywords in combination with the use uniformity;

[0102] If the use uniformity is less than or equal to the preset use uniformity, the keyword is a general keyword;

[0103] If the use uniformity is greater than the preset use uniformity, the keyword is an exclusive keyword.

[0104] When determining whether a single field keyword is an exclusive keyword or a general keyword, the field keyword is recorded as a target field keyword, and the corresponding confirmation method of the combined use degree is to obtain a plurality of data samples, the data samples are request texts of historical completed request tasks and the target field keyword exists in the request texts, detect the number of combined keywords corresponding to the target field keyword in all data samples, and the combined use degree corresponding to the target field keyword is 1 / combined keyword number, wherein the combined keyword is a field keyword that appears in the data sample at the same time as the target field keyword more than a preset co-occurrence number of times;

[0105] The value of the preset combined use degree can be set by the administrator according to the actual scene. It can be understood that the greater the precision requirement of the administrator for the cache relevance determined according to the exclusive keyword similarity, the greater the value of the preset combined use degree. A value setting method is provided to extract the combined use degree corresponding to the historical record that meets the administrator's requirement, remove the outliers of the combined use degree, and record the average value of the combined use degree after removing the outliers as the preset combined use degree.

[0106] Specifically, the load distribution unit determines a load distribution mode according to a load state in response to a load distribution condition;

[0107] If the load state is that the load margin coefficient is greater than the preset load margin coefficient, or the load margin coefficient is less than or equal to the preset load margin coefficient and the load balance coefficient is greater than the preset load balance coefficient, the load distribution mode is gradient distribution;

[0108] If the load state is that the load balance coefficient is less than or equal to the preset load balance coefficient and the load margin coefficient is less than or equal to the preset load margin coefficient, the load distribution mode is user distribution according to the estimated processing time.

[0109] The load margin coefficient is the average value of the current remaining CPU of each server;

[0110] The load balance coefficient is extracted from the current remaining CPU of each server, and the maximum value and the minimum value are extracted. The absolute value H1 of the difference between the maximum value and the average value of the current remaining CPU of each server, and the absolute value H2 of the difference between the minimum value and the average value of the current remaining CPU of each server are calculated respectively, and the load margin coefficient is 1 / H1+H2;

[0111] Gradient distribution is to directly distribute the to-be-distributed users to the server with the largest current remaining CPU;

[0112] Specifically, the load distribution unit determines, in response to a third distribution condition, a predicted processing time length according to a user repeat frequency of each remaining server corresponding to the user, and records a remaining server with the shortest predicted processing time length as a distribution server corresponding to the user to be distributed.

[0113] The third distribution condition is that the load distribution unit determines the load distribution manner as user distribution according to the predicted processing time length.

[0114] The remaining server is a server with a current remaining CPU greater than a load remaining coefficient, the user repeat frequency is a number of users corresponding to a data request in a latest monitoring period, the number of users is greater than a preset repeat number, and the preset repeat number and the monitoring period are values that can be set by an administrator according to an actual scene. It can be understood that the greater the accuracy requirement of the administrator for the user repeat frequency, the greater the length of a single monitoring period, and the greater the requirement of the administrator for the CPU processing speed, the smaller the value of the preset repeat number. The monitoring period is cyclically performed.

[0115] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will fall within the protection scope of the present application.

[0116] The above description is only the preferred embodiments of the present application and is not intended to limit the present application; for those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A domain large model-based knowledge base system, characterized by, The application comprises: a data processing unit comprising several servers for data caching and data processing; a data analysis unit connected to the data processing unit for determining a user portrait stability state according to a user category continuous proportion and a user request number, and determining a user distribution mode as portrait distribution or load distribution according to the user portrait stability state; a portrait distribution unit connected to the data processing unit and the data analysis unit for determining a server state according to a pre-warning proportion coefficient and a one-category user dispersion, and determining a portrait distribution mode as sequential distribution or user distribution according to a cache correlation degree according to the server state; a load distribution unit connected to the data processing unit and the data analysis unit for determining a load state according to a load balancing coefficient and a load margin coefficient, and determining a load distribution mode as gradient distribution or user distribution according to an estimated processing time length according to the load state; the data analysis unit determines a user category according to a text recognition difficulty coefficient and an operation frequency coefficient; if the text recognition difficulty coefficient is greater than a preset text recognition difficulty coefficient or the operation frequency coefficient is greater than a preset operation frequency coefficient, the data analysis unit determines that the user category is a one-category user; if the text recognition difficulty coefficient is less than or equal to the preset text recognition difficulty coefficient and the operation frequency coefficient is less than or equal to the preset operation frequency coefficient, the data analysis unit determines that the user category is a two-category user; the text recognition difficulty coefficient is determined according to a field keyword distribution coefficient and a picture-text combination coefficient; for a single server, the corresponding proportion coefficient = the number of one-category users of the server / the number of two-category users of the server, and if the proportion coefficient is greater than a preset proportion coefficient, the proportion coefficient is recorded as a pre-warning proportion coefficient; A confirmation mode of the user dispersion degree is to detect the number of the type of users corresponding to each server, and the type of user dispersion degree is denoted as S, and the calculation mode of S is: ; wherein i = 1, 2, 3, …, Mz, the Mz is the number of servers, Mi is the number of one-category users corresponding to the i-th server, and M0 is the average value of the number of one-category users corresponding to all servers; the portrait distribution unit is provided with a cache correlation degree calculation strategy, which comprises determining a keyword quantity ratio according to the ratio of general keywords and exclusive keywords in the request task corresponding to the user to be distributed, and determining a cache correlation degree confirmation mode according to the keyword quantity ratio; if the keyword quantity ratio is greater than a preset keyword quantity ratio, the cache correlation degree is determined according to a user preference coefficient; if the keyword quantity ratio is less than or equal to the preset keyword quantity ratio, the cache correlation degree is determined according to an exclusive keyword similarity.

2. The domain large model based knowledge base system according to claim 1, wherein, the data analysis unit determines a user distribution mode according to a user portrait stability state in response to a preset distribution condition; if the user portrait stability state is that the user category continuous proportion is greater than a preset user category continuous proportion and the user request number is greater than a preset user request number, the user distribution mode is portrait distribution; if the user portrait stability state is that the user category continuous proportion is less than or equal to the preset user category continuous proportion or the user request number is less than or equal to the preset user request number, the user distribution mode is load distribution; wherein the preset distribution condition is a new connection request.

3. The domain large model based knowledge base system according to claim 2, wherein, The image distribution unit determines a warning proportionality coefficient quantity according to the proportionality coefficient corresponding to each server in response to an image distribution condition, and determines an image distribution mode according to a server state; If the warning proportionality coefficient quantity is less than a preset warning proportionality coefficient quantity or the first user dispersion is less than a preset first user dispersion, the image distribution mode is determined to be sequential distribution; If the warning proportionality coefficient quantity is greater than or equal to the preset warning proportionality coefficient quantity and the first user dispersion is greater than or equal to the preset first user dispersion, the image distribution mode is determined to be user distribution according to a cache correlation; The image distribution condition is that the data analysis unit determines the user distribution mode to be image distribution.

4. The domain large model based knowledge base system according to claim 3, characterized in that, The image distribution unit obtains a user category of a user to be distributed and a compensation coefficient corresponding to each server in response to a first distribution condition, and records a server with the largest compensation coefficient as a distribution server corresponding to the user to be distributed; The first distribution condition is that the image distribution unit determines the image distribution mode to be sequential distribution, and the compensation coefficient and the related user category number are in a positive correlation. 5.The domain large model based knowledge base system according to claim 3, characterized in that, The image distribution unit detects a cache correlation between a request task corresponding to the user to be distributed and each server in response to a second distribution condition, and records a server with the largest cache correlation as the distribution server corresponding to the user to be distributed; The second distribution condition is that the image distribution unit determines the image distribution mode to be user distribution according to the cache correlation. 6.The domain large model based knowledge base system according to claim 5, characterized in that, The image distribution unit determines the exclusive keyword and the general keyword according to a combination use total; If the combination use total is less than or equal to a preset combination use total, the keyword is a general keyword; If the combination use total is greater than the preset combination use total, the keyword is an exclusive keyword. 7.The domain large model based knowledge base system according to claim 2, characterized in that, The load distribution unit determines a load distribution mode according to a load state in response to a load distribution condition; If the load state is that a load margin coefficient is greater than a preset load margin coefficient, or the load margin coefficient is less than or equal to the preset load margin coefficient and a load balance coefficient is greater than a preset load balance coefficient, the load distribution mode is gradient distribution; If the load state is that the load balance coefficient is less than or equal to the preset load balance coefficient and the load margin coefficient is less than or equal to the preset load margin coefficient, the load distribution mode is user distribution according to an estimated processing time. 8.The domain large model based knowledge base system according to claim 7, characterized in that, The load distribution unit determines an estimated processing time according to a user repetition frequency corresponding to each margin server in response to a third distribution condition, and records a margin server with the smallest estimated processing time as the distribution server corresponding to the user to be distributed; The third distribution condition is that the load distribution unit determines the load distribution mode to be user distribution according to the estimated processing time.

Citation Information

Patent Citations

  • Question answering system based on knowledge base

    CN118093843B

  • Question and answer consulting method and system based on user portraits

    CN117539996A

  • Product recommendation method and device based on intelligent agent and medium

    CN119128277A