Data processing method of large language model based on improved hybrid expert architecture
Through the improved hybrid expert architecture, query information is processed in layers and the output results are integrated, which solves the problems of insufficient accuracy and problem-solving ability of existing large language models in professional segments, and achieves higher accuracy and professionalism.
Patent Information
- Application Number
- CN202511137752.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-14
AI Technical Summary
The existing large language models with hybrid expert architectures have low accuracy in professional niche areas and poor problem-solving capabilities. Especially in complex professional fields such as trade, law, and medicine, they are unable to meet the high requirements of knowledge accuracy, reasoning reliability, and depth of semantic understanding.
An improved hybrid expert architecture is adopted to process the input query information through the semantic understanding module, map it to the knowledge graph matching the professional field, divide it into multi-layer expert models for processing, and route it through the knowledge graph and semantic understanding results. Finally, the output results are fused, and the collaboration of the multi-layer expert model and the knowledge coverage and redundancy are optimized to finally obtain accurate answer results.
It improves the accuracy and professionalism of large language models in professional and niche areas, enhances the ability to answer complex query information, and meets the needs of queries with higher knowledge depth.
Smart Images

Figure CN120725151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data processing method for a large language model based on an improved hybrid expert architecture. Background Art
[0002] In related technologies, large language models (LLMs) employ a Mixture-of-Experts (MoE) architecture. This architecture distributes model parameters across multiple expert models and dynamically selects expert neural networks through a gating routing mechanism for query processing. This allows for more specialized query information and produces more accurate responses.
[0003] However, existing hybrid expert architectures are all single-layer, flat designs, and the expert model's granularity is relatively coarse. This often makes them ineffective for handling more specialized questions, resulting in low accuracy and expertise in their responses. Furthermore, the routing module, primarily based on statistical learning, suffers from poor routing accuracy for more specialized queries.
[0004] In summary, there is currently no effective solution to the problem that large language models with hybrid expert architectures in related technologies have low accuracy and poor problem-solving capabilities in professional niche areas. Summary of the Invention
[0005] The present invention provides a data processing method for a large language model based on an improved hybrid expert architecture, which at least solves the problem in related technologies that the large language model of the hybrid expert architecture has low accuracy and poor problem-solving ability in professional sub-fields.
[0006] According to one aspect of the present invention, a data processing method for a large language model based on an improved hybrid expert architecture is provided, comprising: processing input query information through a semantic understanding module to obtain an understanding result; based on the understanding result and a pre-stored knowledge graph, sending the query information to corresponding multiple expert models, and obtaining corresponding output results from the expert models, wherein the multiple expert models are arranged in multiple layers based on the improved hybrid expert architecture, and the expert models of different layers are obtained by dividing according to different expert dimensions, and at least one expert model in each layer of expert models processes the query information and obtains an output result, and the professional field of the knowledge graph matches the professional field of the query information; the output results of the multiple expert models are fused to obtain an answer result corresponding to the query information.
[0007] As an optional solution, based on the understanding result and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models, and the expert models obtain the corresponding output results, including: mapping the query information to the knowledge graph according to the understanding result, and determining the relevant knowledge nodes of the query information; calculating the fitness of the query information and each of the expert models based on the relevant knowledge nodes; incorporating the expert models that meet the requirements into the expert model set, and screening them to obtain an expert combination; sending the query information to the multiple expert models of the expert combination, and having the multiple expert models of the expert combination output the corresponding output results.
[0008] As an optional solution, the degree of fit between the query information and each of the expert models is calculated based on the relevant knowledge nodes, including: calculating the domain relevance between the query information and the expert model based on the relevant knowledge nodes; determining the similarity between the semantic features in the understanding results of the query information and the semantic profile of the expert model; obtaining the historical scores of the expert model for similar queries to which the query information belongs; and determining the degree of fit based on the domain relevance, the similarity, and the historical scores, combined with corresponding dynamic weights.
[0009] As an optional solution, expert models that meet the requirements of adaptability are included in the expert model set and screened to obtain an expert combination, including: determining multiple candidate subsets from the expert model set through a screening algorithm; calculating the knowledge coverage of the expert models in each candidate subset for the query information, as well as the knowledge redundancy between different expert models; and using an optimal solution algorithm to calculate the optimal candidate subset as the expert combination based on the knowledge coverage and the knowledge redundancy.
[0010] As an optional solution, the query information is sent to multiple expert models of the expert combination, and the multiple expert models of the expert combination output corresponding output results, including: inputting the query information into the expert model in the first layer of the expert combination and outputting a first result, wherein the first layer is the first layer of the improved hybrid expert architecture; sending the first result to the expert model in the next layer and outputting a second result, and inputting the second result to the expert model in the next layer of the expert combination layer by layer to obtain the output results of the multiple expert models of the expert combination in each layer; wherein the improved hybrid expert architecture includes at least: an upper professional dimension layer, and a lower functional dimension layer, the expert model of the professional dimension layer includes at least one of the following: customs experts, tax experts, logistics experts, and financial experts, and the expert model of the functional dimension layer includes at least one of the following: knowledge interpretation experts, rule reasoning experts, risk assessment experts, and decision support experts.
[0011] As an optional solution, the query information is input into the expert model in the first layer of the expert combination, and a first result is output, including: embedding the query information into professional terms through a domain professional term embedding layer to obtain a professional term embedding result; using an adaptive multi-head attention layer to process the professional term embedding result to obtain a context-aware representation; using a knowledge-parameter mapping engine to map the structured professional knowledge in the context-aware representation to the model parameter space to obtain a knowledge-enhanced representation; based on the knowledge-enhanced representation, using an interpretable output layer to obtain a corresponding first result; wherein, the expert model of the next layer receives the upper layer result, and uses the domain professional term embedding layer, the adaptive multi-head attention layer, the knowledge-parameter mapping engine and the interpretable output layer to obtain a corresponding output result.
[0012] As an optional solution, the output results of multiple expert models are fused to obtain an answer result corresponding to the query information, including: performing a credibility assessment on the expert model and the corresponding output result based on the professional coverage of the knowledge nodes of the query information by the expert model's knowledge graph, the accuracy of the expert model's historical output, and the certainty of the output result of the query information to obtain credibility; calculating the weight of each output result based on the credibility, the query relevance between the expertise vector of the expert model and the query vector of the query information, and the complementarity parameters between different expert models; reconciling conflicting output results based on the credibility to obtain an adjusted output result; and determining the answer result based on the adjusted output result and the corresponding weight.
[0013] As an optional solution, based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models. After the expert model obtains the corresponding output results, it includes: calculating the self-assessment score according to the expert model's knowledge coverage of the query information, as well as the reasoning reliability and quantitative uncertainty of the output results; through expert mutual review, determining whether there are errors in the output results of the expert model. If it is determined that there are errors, the corresponding mutual review experts output correction suggestions for error correction and feedback learning; through multiple rounds of mutual review, the output results of multiple expert models that form a consensus are obtained, which are used for subsequent fusion to obtain the answer result.
[0014] As an optional solution, based on the understanding result and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models. Before the expert model obtains the corresponding output result, the method also includes: independently pre-training the expert model through training data matched with the professional field; based on the expert model after pre-training, training the routing module of the improved hybrid expert architecture, wherein the routing module is used to send the query information to the corresponding multiple expert models based on the understanding result and the pre-stored knowledge graph; based on the trained routing module and the expert model after pre-training, joint optimization is performed through the target optimization function, wherein the target optimization function includes: answer result loss function, routing quality loss function, expert diversity loss function, and knowledge consistency loss function.
[0015] According to another aspect of the present invention, there is provided an electronic device comprising: a processor and a memory for storing a program, wherein the program comprises instructions which, when executed by the processor, cause the processor to perform the above method.
[0016] According to one aspect of the present invention, there is provided an electronic device, comprising: a processor, and a memory for storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the above method.
[0017] The data processing method of the large language model based on the improved hybrid expert architecture provided by the embodiment of the present invention adopts a multi-layer expert model obtained by dividing according to different expert dimensions, obtains output results based on query information, and finally fuses the output results to obtain corresponding answer results, which has higher accuracy and professionalism. Moreover, based on the semantic understanding results of the query information combined with the knowledge graph matching the professional field, the query information is routed and sent to the expert model in the improved hybrid expert architecture, which has higher accuracy. It further improves the professionalism and accuracy of answering query information with deeper knowledge depth. This solves the problem that the large language model of the hybrid expert architecture in the related art has low accuracy in professional subdivision fields and poor problem-solving ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without inventive effort.
[0019] Figure 1This is a flowchart of a data processing method for a large language model based on an improved hybrid expert architecture according to an embodiment of the present invention.
[0020] Figure 2 Schematic diagram of an improved hybrid expert architecture according to an embodiment of the present invention.
[0021] Figure 3 It is a schematic diagram of data processing of an improved hybrid expert architecture according to an embodiment of the present invention.
[0022] Figure 4 It is a schematic diagram of the expert model architecture of an embodiment created by the present invention.
[0023] Figure 5 It is a schematic diagram of the routing module architecture of an embodiment created by the present invention.
[0024] Figure 6 It is a schematic diagram of the fusion architecture of the output results of the embodiment created by the present invention.
[0025] Figure 7 Schematic diagram of the self-assessment and mutual assessment architecture of an embodiment of the present invention.
[0026] Figure 8 It is a schematic diagram of the training architecture of an embodiment of the present invention.
[0027] Figure 9 It is a schematic diagram of the knowledge graph architecture of an embodiment created by the present invention.
[0028] Figure 10 It is a schematic diagram of the implementation architecture of the large language model of the improved hybrid expert architecture of the embodiment created by the present invention.
[0029] Figure 11 It is a structural schematic diagram of an electronic device created by the present invention. DETAILED DESCRIPTION
[0030] The following describes embodiments of the present invention in more detail with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0031] With the rapid development of artificial intelligence (AI), large language models have become a mainstream technology in natural language processing. However, as application scenarios continue to expand, general-purpose large language models face numerous challenges in specialized fields, particularly in terms of knowledge depth, expertise, and accuracy.
[0032] The Mixture of Experts (MoE) architecture has been widely used in large language models in recent years. Mainstream MoE architectures include Google's Mixture of Experts Model Sharding (GShard) and the Switch Transformer architecture. GShard is a transformer-based sparse conditional computation system that enhances the model's representational capabilities by using MoEs in feedforward network layers. The Switch Transformer further simplifies the routing design based on GShard, adopting a routing strategy where each query is routed to a single optimal expert model.
[0033] However, the above MoE architecture has the following major disadvantages.
[0034] (1) Flat structure design of single-layer expert architecture: The traditional MoE architecture adopts a flat expert model organizational structure, where all expert models are at the same level. It cannot effectively express the hierarchical knowledge system of professional fields and it is difficult to capture the multi-dimensional professional needs in complex problems.
[0035] (2) Coarse-grained division of expert models: The division of expert models in existing MoE models is usually based on computational efficiency rather than knowledge boundaries, which leads to knowledge overlap or gaps between expert models and reduces the accuracy of professional field applications.
[0036] (3) The routing mechanism lacks semantic awareness: The routing mechanism in the traditional MoE architecture is mainly based on statistical learning methods, which lacks a deep understanding of professional concepts and terminology, making it difficult to accurately distribute queries to the most suitable expert model combination.
[0037] (4) Simple expert model collaboration mechanism: The expert model collaboration in the existing MoE architecture usually adopts a simple weighted average or selects the expert model with the highest score, which lacks the ability of expert model complementarity, error correction and collaborative reasoning in complex scenarios.
[0038] The above shortcomings limit the application of the existing MoE architecture in complex professional fields (such as trade, law, and medicine), making it difficult to meet the high requirements of professional fields for knowledge accuracy, reasoning reliability, and depth of semantic understanding.
[0039] In order to solve the above problems, according to one aspect of the present invention, a data processing method for a large language model based on an improved hybrid expert architecture is provided. Figure 1It is a flowchart of a data processing method for a large language model based on an improved mixture of experts architecture in an embodiment of the present invention, as Figure 1 shown, and the method specifically includes the following steps.
[0040] Step S101, process the input query information through a semantic understanding module to obtain an understanding result.
[0041] Step S102, based on the understanding result and a pre-stored knowledge graph, send the query information to a corresponding plurality of expert models, and obtain corresponding output results by the expert models. Among them, the plurality of expert models are set with multiple layers based on an improved mixture of experts architecture, and the expert models of different layers are obtained according to different expert dimensions. At least one expert model in each layer of expert models processes the query information and obtains an output result. The professional field of the knowledge graph matches the professional field of the query information.
[0042] Step S103, perform a fusion process on the output results of the plurality of expert models to obtain an answer result corresponding to the query information.
[0043] The above data processing method for a large language model based on an improved mixture of experts architecture in this embodiment adopts multiple layers of expert models obtained according to different expert dimensions, obtains an output result based on the query information, and finally performs output result fusion to obtain a corresponding answer result, which has higher accuracy and professionalism. Moreover, based on the semantic understanding result of the query information and combined with a knowledge graph with a matching professional field, the query information is routed and sent to the expert models in the improved mixture of experts architecture, which has higher accuracy.
[0044] The execution subject of the above steps can be a large language model based on an improved mixture of experts architecture, and this large language model can run on a computer, a server, or in the cloud. The above semantic understanding module can be a processing module based on semantic understanding, and finally outputs a structured understanding result of the query information. The understanding result can include the key concepts and requirement types of the query information.
[0045] When the above semantic understanding module processes the query information, it usually includes text preprocessing to remove meaningless words in the text, such as prepositions "and", "of", etc. The preprocessed text is passed through a classification model for intention recognition to determine the purpose of the query information, that is, the requirement type. Then extract the key concepts in the query information, that is, the entities. Based on the extracted key concepts, entity annotation and relationship extraction are performed. The finally obtained entities, entity relationships, and requirement types are all included in the understanding result.
[0046] Map entities and relationships into a knowledge graph for subsequent use. This knowledge graph matches the domain of the query. For example, the domain could be trade, law, healthcare, or commercial law, a sub-domain of law.
[0047] The above-mentioned knowledge graph is pre-created and stored in a fixed location before use. Creation of the knowledge graph involves extracting entities and relationships from authoritative sources such as professional materials and professional books in the corresponding patent field. Knowledge from different sources is then integrated to eliminate redundancies and contradictions. The knowledge graph is then constructed and expanded through rule-based reasoning. Finally, the knowledge is verified for correctness through expert review and automated consistency checks, and stored after passing verification.
[0048] Because the understanding results include the query's requirement type and key concepts, these key concepts can be mapped into a knowledge graph. Using the knowledge graph's nodes and relationships, we can identify the knowledge nodes closest to the query, or the expert models that best cover the knowledge nodes. This means that not all expert models in the improved hybrid expert architecture will process the query.
[0049] Based on the query information's understanding, a search and match is performed to identify a subset of qualified expert models to process the query information. Expert models typically have a built-in knowledge base that represents the expert's expertise. The compatibility of the knowledge base with key concepts can be used to characterize the expert model's compatibility with the query information. Based on this compatibility, the expert model most relevant to the query information can be found to solve the problem, thereby improving the accuracy of the expert model's output for the query information.
[0050] Based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models, and the expert models obtain the corresponding output results. The specific details will be explained later.
[0051] The above-mentioned multiple expert models are set up in multiple layers based on the improved hybrid expert architecture. The expert models at different layers are obtained according to different expert dimensions. For example, Figure 2 The improved hybrid expert architecture shown here has two layers of expert models: a professional dimension layer and a functional dimension layer. These layers of expert models enable vertical collaboration between different layers to improve the knowledge depth of query information processing and comprehension.
[0052] For example, after the first-layer expert model processes the query information, the resulting output is used as the input for the next-layer expert model to obtain the output of the next-layer expert model. This enables the linkage of multi-dimensional experts, allowing for multi-dimensional understanding and processing of query information, solving the problem of understanding and resolving query information at a higher level of knowledge depth.
[0053] At least one expert model in each layer processes the query information and generates an output. This is because different layers are divided according to different professional dimensions. To improve the depth of query processing and understanding, each professional dimension needs to cooperate and collaborate to deepen the processing depth of query information, thereby meeting the needs of query information with higher knowledge depth.
[0054] Since different expert models have different levels and different dimensions, the results directly output by different expert models may not form a consensus. Therefore, this embodiment fuses the output results of multiple expert models to obtain the answer corresponding to the query information.
[0055] It's important to note that before fusing the outputs of expert models, the credibility of the expert models can be assessed using multiple parameters. Based on this credibility, the expert models are weighted, and then the fusion is performed based on the weights. Conflict detection can also be performed before direct fusion to avoid obvious conflicts and logical loopholes between the outputs of different expert models, which could lead to reduced accuracy and professionalism in the final fused answer.
[0056] The details of the above credibility evaluation, weight allocation and conflict detection are described later.
[0057] In summary, based on the semantic understanding of the query information and the knowledge graph matching the professional domain, the query information is routed and sent to the expert model in the improved hybrid expert architecture. Using a multi-layered expert model divided according to different expert dimensions, the model generates output based on the query information, and finally fuses the output results to obtain the corresponding answer. This improves the accuracy and professionalism of query processing, meeting the needs of responding to queries with higher knowledge depth. This addresses the problem of large language models in hybrid expert architectures in related technologies, which have low accuracy and poor problem-solving capabilities in specialized domains.
[0058] As an optional embodiment, based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models, and the expert models obtain corresponding output results, including. Based on the understanding results, the query information is mapped to the knowledge graph, and the relevant knowledge nodes of the query information are determined. Based on the relevant knowledge nodes, the fitness of the query information and each expert model in the improved hybrid expert architecture is calculated. The expert models that meet the requirements of fitness are included in the expert model set and screened to obtain an expert combination. The query information is sent to the multiple expert models of the expert combination, and the multiple expert models of the expert combination output corresponding output results.
[0059] Because the understanding results only contain the key concepts of the entities contained in the query information itself and cannot directly represent the knowledge nodes involved in the query information, the query information can be mapped to the knowledge graph based on the understanding results to determine the relevant knowledge nodes of the query information. These relevant knowledge nodes can then represent the knowledge involved in the query information.
[0060] The fit between the query information and each expert model in the improved hybrid expert architecture is then calculated based on the relevant knowledge nodes to characterize the degree of knowledge match between each expert model and the query information. Clearly, the higher the knowledge match, the more accurate the query information processing and understanding, and the higher the expertise. Therefore, based on the fit, expert models can be selected for query processing. Specifically, expert models that meet the required fit are first included in the expert model set. This expert model set can include all the expert models in the improved expert architecture that can be used for query processing.
[0061] However, if the query information is not sufficiently deep and specialized, the number of expert models in the expert model set will be large, resulting in a large number of outputs. This can lead to new problems during fusion and increase the probability of knowledge conflicts. Furthermore, the processing process will consume more processing resources, take longer, and be less efficient.
[0062] To address this issue, the expert model set can be screened to obtain an optimized expert combination, which serves as the set of expert models used to process the query information. This screening process can consider the knowledge redundancy of different expert models, the knowledge coverage of each expert model for the query information, and the difficulty of collaboration between different expert models. This will be explained in detail later.
[0063] After obtaining the expert combination, the query information can be sent to the multiple expert models in the expert combination, which then generate corresponding outputs. This serves as the data foundation for the final fusion. Furthermore, the expert models in the expert combination are hierarchical, and the query information can be sent to the expert combination directly or layer by layer. As mentioned above, this embodiment uses a layer-by-layer approach to achieve collaboration among experts in multiple dimensions. The specific details will be explained in detail later.
[0064] As an optional embodiment, the compatibility between the query information and each expert model in the improved hybrid expert architecture is calculated based on relevant knowledge nodes, including: calculating the domain relevance between the query information and the expert model based on the relevant knowledge nodes; determining the similarity between the semantic features in the query information understanding result and the semantic profile of the expert model; obtaining the expert model's historical scores for similar queries to which the query information belongs; and determining the compatibility based on the domain relevance, similarity, and historical scores, combined with corresponding dynamic weights.
[0065] The aforementioned calculation of the domain relevance between the query information and the expert model based on relevant knowledge nodes can be performed using a relevance algorithm after extracting the knowledge nodes of the query information and determining the knowledge nodes of the expert model's knowledge base. Specifically, domain relevance can be calculated using the Jaccard similarity function. Alternatively, a vector space model can be used to encode the node set into a vector (such as a one-hot or graph embedding) and calculate cosine similarity. Alternatively, a personalized PageRank algorithm can be run on the knowledge graph to calculate the relevance score of the expert node starting from the query node.
[0066] It should be noted that the expert model itself is also provided with a semantic profile, which is a metadata file used to structurally describe the knowledge domain, capability boundaries and semantic characteristics of the expert model. Its core function is to achieve accurate matching, collaborative optimization and system interpretability of the expert model. Taking into account the semantic characteristics of the query information, this embodiment also refers to the similarity between the semantic features of the query information and the semantic profile of the expert model in the process of calculating the adaptability of the expert model. The semantic profile is a structured semantic representation that usually contains information such as the semantic category, attributes, and contextual dependencies of the entity. The semantic feature is the smallest unit that describes the semantics of the entity and is the basic component of the semantic profile. Its specific similarity calculation methods may include the overlap coefficient method based on feature matching, the similarity algorithm based on vector space and the Euclidean distance algorithm, etc.
[0067] The historical scores of the expert model can be obtained from historical data, which can reflect the historical performance of the expert model in the improved hybrid expert architecture and the language model, and can also characterize the accuracy and professionalism of the expert model in processing query information to a certain extent.
[0068] This embodiment determines the degree of adaptation based on domain relevance, similarity, and historical scores, combined with corresponding dynamic weights, which can be expressed by the following formula.
[0069] S_i = α·sim(Q_emb, E_i_profile) + β·KG_relevance(Q, E_i) + γ·E_i_performance. Sim(Q_emb, E_i_profile) is the similarity between the semantic features of the query and the expert's semantic profile. KG_relevance(Q, E_i) is the domain relevance between the query and the expert model based on the knowledge graph. E_i_performance is the expert's historical rating on similar queries. α, β, and γ are dynamically adjusted weighting coefficients.
[0070] The above method can accurately calculate the adaptability of the expert model and the query information, and can more accurately capture the relevance and hierarchy of professional knowledge, and better meet the needs of the professional field.
[0071] As an optional embodiment, expert models with a degree of adaptability that meets the requirements are included in the expert model set and screened to obtain an expert combination, including: determining multiple candidate subsets from the expert model set through a screening algorithm; calculating the knowledge coverage of the expert models in each candidate subset for the query information, as well as the knowledge redundancy between different expert models; and using an optimal solution algorithm to calculate the optimal candidate subset as the expert combination based on the knowledge coverage and knowledge redundancy.
[0072] When selecting expert combinations from a set of expert models, a screening algorithm is first used to identify multiple candidate subsets. The expert model set involves multiple expert models at different levels, and the collaborative effects of different expert models vary, especially given the large number of possible combinations. A screening algorithm, such as the Top-K preliminary screening algorithm, can be used to identify multiple candidate subsets that meet the requirements. This preliminary screening of potential expert models quickly narrows the candidate pool and improves the efficiency of selecting expert combinations. After the initial screening of expert models, a candidate subset is generated using an exploratory combination algorithm.
[0073] Then, for each candidate subset, the knowledge coverage of the expert model in each candidate subset for the query information and the knowledge redundancy between different expert models are calculated. Based on the knowledge coverage and knowledge redundancy, the optimal candidate subset is calculated as the expert combination using the optimal solution algorithm.
[0074] The knowledge coverage of the query information by the multiple expert models of the candidate subset can be determined by combining the relevant knowledge nodes of the query information with the knowledge base of the multiple expert models of the candidate subset. It should be noted that when querying posture coverage based on relevant knowledge nodes, more relevant knowledge nodes can be extracted from the knowledge graph to ensure the knowledge coverage of the candidate subset for the query information.
[0075] When determining the fit between the query information and the expert model, a higher threshold can be used to filter the knowledge nodes in the knowledge graph, thereby reducing the number of relevant knowledge nodes. This ensures that the number of relevant knowledge nodes in the query information is not significantly different from the magnitude of the expert model's knowledge base, and also ensures a certain degree of accuracy in the fit determination. Based on the knowledge coverage and knowledge redundancy, an optimal solution algorithm is used to calculate the optimal candidate subset as the expert combination. The optimal solution algorithm can be an integer programming algorithm, a Monte Carlo tree search algorithm, or a greedy algorithm. The optimal candidate subset is selected from multiple candidate subsets as the expert combination.
[0076] As described above, the method of inputting query information layer by layer adopted in this embodiment is specifically as follows.
[0077] As an optional embodiment, query information is sent to multiple expert models in an expert combination, and the multiple expert models in the expert combination output corresponding output results, including: inputting the query information to an expert model in the first layer of the expert combination, and outputting a first result, wherein the first layer is the first layer of the improved hybrid expert architecture; sending the first result to an expert model in the next layer, and outputting a second result; and inputting the second result to the expert model in the next layer of the expert combination layer by layer, to obtain the output results of the multiple expert models in each layer of the expert combination.
[0078] In this embodiment, if Figure 2 As shown, the improved hybrid expert architecture for the trade field includes at least: an upper professional dimension layer, and a lower functional dimension layer. The expert model of the professional dimension layer includes at least one of the following: customs experts, tax experts, logistics experts, and financial experts. The expert model of the functional dimension layer includes at least one of the following: knowledge interpretation experts, rule reasoning experts, risk assessment experts, and decision support experts.
[0079] This two-layer expert model refines the granularity of expert divisions and improves the accuracy of expert model applications. Expert collaboration enables complementary and joint reasoning, further enhancing the accuracy and expertise of query processing in the large language model using the improved hybrid expert architecture.
[0080] like Figure 3As shown, as an optional embodiment, the query information is input into the expert model in the first layer of the expert combination, and the first result is output, including. The query information is embedded in professional terms through the domain professional term embedding layer to obtain the professional term embedding result. The professional term embedding result is processed using the adaptive multi-head attention layer to obtain a context-aware representation. For the context-aware representation, the structured professional knowledge in the context-aware representation is mapped to the model parameter space using the knowledge-parameter mapping engine to obtain a knowledge-enhanced representation. Based on the knowledge-enhanced representation, the interpretable output layer is used to obtain the corresponding first result. Among them, the expert model of the next layer receives the upper layer result, and obtains the corresponding output result using the domain professional term embedding layer, the adaptive multi-head attention layer, the knowledge-parameter mapping engine and the interpretable output layer.
[0081] The executor of the aforementioned domain-specific terminology embedding layer can be a domain-customized embedding matrix. Processing steps: Input: Original text (e.g., "FOB Shanghai Incoterms® 2020"). Processing: Using a pre-trained trade term embedding matrix (e.g., trained on corpora such as Incoterms® and UCP600), the technical terms are mapped into high-dimensional vectors. Compound terms (e.g., "FOB Shanghai") are semantically combined and encoded to preserve the integrity of the terms. Output changes: Original word segmentation result ["FOB", "Shanghai"] → vector [0.7, -0.2, ..., 1.4] that incorporates technical semantics, i.e., the aforementioned technical term embedding result.
[0082] The adaptive multi-head attention layer takes the above vector as input and uses domain attention gating to calculate weights for each attention head. For example, tokens related to trade terms (such as "CIF") automatically assign higher weights to the "Shipping Liability" attention head. Tokens related to legal terms (such as "Force Majeure") activate the "Legal Terms Parsing" attention head. Finally, a context-aware representation is generated, which emphasizes the relevance of technical terms (such as the implicit association between "FOB" and "risk transfer").
[0083] The expertise enhancement layer can include the aforementioned knowledge-parameter mapping engine, which injects knowledge into the attentive contextual representation. This involves querying a structured knowledge base (e.g., the Incoterms® rule tree) to convert the clause logic into parameter constraints. Dynamic parameter adjustment reconstructs network parameters based on activated knowledge units (e.g., trade terms module, tariff calculation module). Knowledge fusion fuses the rule logic output (e.g., "Seller's liability ends at the port of shipment under FOB") with the neural network representation via residual connections. This enhances the hidden layer output of a standard Transformer into a representation for the injected rule constraints.
[0084] The interpretable output layer utilizes a traceable decoder to enhance the knowledge representation, performing multi-granular output and logical verification. A lightweight inference engine verifies the output's consistency with the knowledge base. For example, consider the input "Payment by LC (irrevocable letter of credit) in accordance with UCP600 (Uniform Customs and Practice for Documentary Credits)." The expert module processes the input using the following methods: The word embedding layer labels "LC" as a specific letter of credit vector (not a common abbreviation). The attention mechanism associates "UCP600" with the "Letter of Credit Operating Specifications" attention head. The knowledge enhancement layer incorporates the parameter constraints for irrevocable letters of credit from Article 9 of UCP600. The output is annotated with references to specific UCP600 provisions.
[0085] This allows expert models to maintain the data-driven advantages of deep learning while strictly adhering to domain knowledge standards when processing professional texts. This ensures that data flow conforms to the forward propagation mechanism of deep learning while maintaining the rule transparency of expert systems.
[0086] As an optional embodiment, the output results of multiple expert models are fused to obtain an answer result corresponding to the query information, including: Based on the professional coverage of the knowledge nodes of the query information by the expert model's knowledge graph, the accuracy of the expert model's historical output, and the certainty of the output results for the query information, the expert model and the corresponding output result are evaluated for credibility to obtain credibility. Based on the credibility, the query relevance between the expert model's expertise vector and the query vector of the query information, and the complementarity parameters between different expert models, the weight of each output result is calculated. Based on the credibility, the conflicting output results are reconciled to obtain an adjusted output result. The answer result is determined based on the adjusted output result and the corresponding weight.
[0087] like Figure 6As shown in the figure, when fusing the output results of multiple expert models in an expert combination, it is necessary to determine the expert weights to ensure the effectiveness and accuracy of the fusion and to reconcile conflicts to avoid degradation of the fusion results. The weights can be calculated based on the credibility of the expert model. Specifically, based on the professional coverage of the expert model's knowledge graph for the knowledge nodes of the query information, the accuracy of the expert model's historical output, and the certainty of the output results for the query information, the expert model and the corresponding output results are evaluated for credibility to obtain credibility. The weights of each output result are then calculated based on the credibility, the query relevance between the expert model's expertise vector and the query vector of the query information, and the complementarity parameters between different expert models.
[0088] Specifically, it can be calculated using the following formula.
[0089] C_i=ω1·Domain_Coverage_i+ω2·Historical_Accuracy_i+ω3·Certainty_ Score_i, in, C_i is the credibility of expert model i, Domain_Coverage_i For professional coverage, Historical_ Accuracy_i is the historical accuracy, Certainty_Score_i is the certainty of the current output, ω1, ω2, ω3 are their respective weights.
[0090] Dynamic weight allocation: W_i= softmax(λ1·C_i+λ2·Relevance_i+λ3· Complementarity_i), Among them, λ1, λ2, λ3 are learnable weights, softmax() is the probability distribution function, W_i is the dynamic weight of expert model i, Relevance_i To query relevance, Complementarity_i is the expert collaboration characteristic, that is, the complementarity parameter.
[0091] When conflicts are detected, they can be reconciled based on credibility. Specifically, the conflict reconciliation function: When knowledge conflicts are detected between experts, the reconciliation function is used to resolve them. This step effectively addresses knowledge conflicts and complementarities during multi-expert collaboration, significantly improving the accuracy and reliability of the system's processing of complex queries.
[0092] As an optional embodiment, based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models. After the expert models obtain the corresponding output results, a self-assessment score is calculated based on the expert model's knowledge coverage of the query information, as well as the reasoning reliability and quantitative uncertainty of the output results. Through expert peer review, it is determined whether there are errors in the output results of the expert models. If errors are determined, the corresponding peer review experts will output correction suggestions for correction and feedback learning. After multiple rounds of peer review, the output results of multiple expert models that reach a consensus are obtained and then fused to obtain the answer results.
[0093] like Figure 7 As shown in Figure 2, the expert's self-assessment is evaluated using the following metrics: Knowledge coverage: assesses the extent to which the expert covers the knowledge required for the current query. Reasoning reliability: assesses the logical rigor of the reasoning process. Uncertainty quantification: quantifies the uncertainty of the output.
[0094] The expert peer review and error correction process is as follows: Output peer review: The output of expert model A is reviewed by expert model B. Error correction suggestion generation: When an error is discovered, specific correction suggestions are generated. Error correction is initiated when there is a clear conflict between the judgment of expert model B and the output of expert model A. Consensus formation: Expert consensus is reached through multiple rounds of peer review and discussion. The maximum number of peer review rounds is three. Feedback learning is also possible: feedback learning based on the error correction results is conducted to continuously improve expert capabilities.
[0095] This self-assessment and error correction mechanism greatly enhances the model's self-verification capabilities and improves the reliability and credibility of the system. It is particularly suitable for application scenarios such as trade, which are highly professional and have low fault tolerance.
[0096] As an optional embodiment, based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models, and before the expert model obtains the corresponding output results, the method also includes. The expert model is independently pre-trained with training data matched with the professional field. Based on the expert model after pre-training, the routing module of the improved hybrid expert architecture is trained, wherein the routing module is used to send the query information to the corresponding multiple expert models based on the understanding results and the pre-stored knowledge graph. Based on the trained routing module and the expert model after pre-training, joint optimization is performed through the target optimization function, wherein the target optimization function includes: answer result loss function, routing quality loss function, expert diversity loss function, and knowledge consistency loss function.
[0097] This approach of pre-training the expert model and routing module, followed by joint objective optimization training, allows for rapid and accurate model training, avoiding issues with model training location. Furthermore, multi-objective optimization ensures balanced model development across multiple dimensions, including task performance, routing quality, expert diversity, and knowledge consistency.
[0098] Computational efficiency can also be optimized to increase model response rate, reduce latency, and reduce content occupancy, etc., which will be specifically explained in the implementation method.
[0099] It should be noted that this embodiment also provides an optional implementation method, specifically an example of a practical application of the data processing method of the large language model based on the improved hybrid expert architecture. This implementation method is described in detail below.
[0100] This embodiment provides a large language model based on an improved hybrid expert architecture, such as Figure 2 As shown, this large language model mainly includes the following modules.
[0101] User query / input module: receives queries or questions input by users.
[0102] Query understanding module: Analyzes the semantic intent of user queries and extracts key concepts and demand information.
[0103] Knowledge graph module: stores concepts, relationships, and knowledge systems in professional fields to provide support for routing decisions.
[0104] Expert router: Distributes queries to appropriate expert models based on query understanding results and knowledge graph information.
[0105] Domain Layer: A collection of expert models divided according to professional fields, such as customs experts, tax experts, logistics experts, financial experts, etc.
[0106] Functional layer: A collection of expert models divided according to functional requirements, such as knowledge interpretation, rule reasoning, risk assessment, decision support, etc.
[0107] Expert fusion and output module: Integrates the outputs of multiple expert models to generate the final answer.
[0108] This implementation utilizes a two-tiered expert organizational structure, dividing the expert model into a professional dimension layer and a functional dimension layer. This approach simultaneously meets the requirements for both deep expertise and functional diversity. This hierarchical hybrid expert architecture better adapts to the layered knowledge system within specialized domains, improving the model's knowledge coverage and accuracy within those domains. Furthermore, each layer of the expert model can be based on a refined division of expertise within a specific domain, ensuring clear knowledge boundaries and specialized characteristics, reducing knowledge overlap and gaps.
[0109] Through a semantically-aware expert routing mechanism based on query understanding results, we can accurately understand the intent of specialized queries and distribute them to the most appropriate expert team. Finally, we build an efficient expert collaboration framework that supports inter-expert complementarity, error correction, and collaborative reasoning, enhancing complex problem-solving capabilities. This approach is more applicable to specialized domains and significantly improves the effectiveness of large models in these areas.
[0110] like Figure 3 As shown, the overall data processing flow of the large language model based on the improved hybrid expert architecture is as follows.
[0111] User query input: Receive query information entered by the user.
[0112] Query understanding: The system analyzes the semantic intent of query information and extracts key concepts and requirement information.
[0113] Expert routing: Select an appropriate expert model based on query understanding results and knowledge graph.
[0114] Professional dimension processing: Experts in selected professional fields process queries and provide professional knowledge output.
[0115] Functional dimension processing: Selected functional experts process queries and perform specific functional tasks.
[0116] Expert self-assessment and mutual review: Experts assess the quality of their own outputs, review and correct each other's errors, and reach consensus.
[0117] Expert fusion and final output: Fusion of multiple expert outputs to generate the final answer.
[0118] Throughout the entire process, knowledge graphs are pre-created for professional fields, providing necessary knowledge support for query understanding, expert routing, and expertise processing.
[0119] Specifically, this improved hybrid expert architecture, called the Hierarchical Mixture-of-Experts (HMoE), innovatively divides the complex trade knowledge system into two hierarchical levels: professional and functional. This creates a more sophisticated expert collaboration mechanism, providing enhanced professional performance for trade queries.
[0120] 1. Two-tier division: Professional and functional dimensions: The first tier divides expert models according to professional dimensions such as customs, taxation, logistics, and finance; the second tier divides expert models according to functional dimensions such as knowledge interpretation, rule reasoning, and risk assessment. This two-tier design more accurately addresses the complexity and interdisciplinary nature of the trade knowledge system.
[0121] 2. Knowledge Graph-Based Expert Routing: This innovative approach leverages domain-specific knowledge graphs to assist expert routing decisions. Compared to traditional statistically-based routing mechanisms, this mechanism more accurately identifies query intent and selects the appropriate expert combination. Experiments have shown that this mechanism can improve expert selection accuracy by 15-20%.
[0122] 3. Adaptive Expert Fusion Algorithm: An adaptive fusion algorithm is designed that takes into account the credibility, knowledge coverage, and reasoning ability of experts. It can dynamically adjust the weighted contributions of different experts based on the complexity of the query and the fields involved, solving the limitations of fixed-weight fusion in traditional MoE.
[0123] 4. Expert self-assessment and error correction mechanism: An innovative self-assessment mechanism for expert capabilities is introduced, allowing experts to rate the credibility of their own answers and achieve mutual error correction and supplementation through multi-expert collaboration, significantly improving the accuracy and reliability of the overall system.
[0124] Current mainstream MoE architectures, such as Google's GShard and Switch Transformer, focus primarily on improving computational efficiency and general language capabilities, but pay insufficient attention to the refined expression of specialized domain knowledge. The HMoE architecture proposed in this implementation is a significant innovation in applying existing MoE technology to specialized domains, and is particularly well-suited for complex professional scenarios like trade, which have vast knowledge systems and numerous sub-domains.
[0125] The above two-level division of professional dimension and functional dimension mainly includes professional dimension layer (Domain Layer) and functional dimension layer (Function Layer).
[0126] The professional dimension divides expert models into different sub-fields of trade, including customs experts, tax experts, logistics experts, and financial experts. The functional dimension divides expert models into functional requirements such as knowledge interpretation, rule reasoning, and risk assessment. This dual-level division enables the model to possess both professional depth and functional flexibility.
[0127] Each expert model uses a customized Transformer architecture, such as Figure 4 As shown, Figure 4 FIG. 1 is a schematic diagram of the expert model architecture of an embodiment of the present invention. The expert model includes the following core components.
[0128] Domain-specific terminology embedding layer: This layer uses word embedding representations designed for trade-specific terminology to enhance understanding of the terminology.
[0129] Adaptive multi-head attention layer: A multi-head attention layer that can adaptively adjust attention distribution according to the characteristics of professional fields.
[0130] Expertise enhancement layer: A special layer structure that integrates pre-trained knowledge and expertise.
[0131] Interpretable output layer: An output layer that supports reasoning process tracking and decision-making basis annotation.
[0132] The key technical point is that the professional knowledge enhancement layer adopts the "knowledge-parameter mapping engine" technology to explicitly map structured professional knowledge (such as trade term definitions, rule clauses, etc.) into the model parameter space, so that the model can maintain the flexibility of the neural network while having the ability to accurately express professional knowledge.
[0133] The above-mentioned expert router innovatively combines knowledge graph technology with expert routing mechanism, and designs a dual-wheel drive routing algorithm of "semantic perception + graph reasoning". Figure 5 Schematic diagram of the routing module architecture of the embodiment of the present invention, such as Figure 5 As shown, the routing mechanism consists of four core components.
[0134] Semantic understanding module: Analyzes the semantic intent of user queries and extracts key concepts and requirement types.
[0135] Knowledge graph query module: maps user queries to the trade knowledge graph and identifies relevant knowledge nodes.
[0136] Expert mapping module: Determines the set of expert models that need to be activated based on knowledge nodes and query types.
[0137] Decision integration module: Determines the final expert combination strategy based on query complexity and professional coverage requirements.
[0138] The core routing algorithm consists of three key steps.
[0139] Semantic space mapping: The query content is mapped to a high-dimensional semantic space through a semantic encoder.
[0140] Expert fitness calculation: Calculate the fitness score between the query and each expert model.
[0141] S_i = α sim(Q_emb, E_i_profile) + β KG_relevance(Q, E_i) + γ E_i_performance. S_i is the fit between expert model i and the query. sim(Q_emb, E_i_profile) is the similarity between the semantic features of the query and the expert's semantic profile. KG_relevance(Q, E_i) is the domain relevance between the query and the expert model based on the knowledge graph. E_i_performance is the expert's historical rating on similar queries. α, β, and γ are dynamically adjusted weighting coefficients.
[0142] Optimal expert combination selection: The final expert combination is determined based on Top-K screening and collaborative scoring mechanism.
[0143] E_selected = argmax_{E_subset} [Coverage(E_subset) - Redundancy(E_subset)]. Coverage(E_subset) evaluates the knowledge coverage of the expert subset for the query. Redundancy(E_subset) evaluates the knowledge redundancy between expert subsets.
[0144] Compared with traditional routing methods based on statistical learning, this knowledge graph-based expert routing mechanism can more accurately capture the relevance and hierarchy of professional knowledge and better meet the needs of professional fields.
[0145] The above-mentioned expert fusion and output module adopts an adaptive expert fusion algorithm with the characteristics of "dynamic weight, credibility perception, and conflict coordination". Figure 6 This is a schematic diagram of the fusion architecture of the output results of the embodiment created by the present invention, such as Figure 6 As shown: The adaptive expert fusion algorithm contains four key modules.
[0146] Expert credibility assessment: Assess the credibility of experts based on their professional coverage in the field, historical accuracy, and the certainty of their current output.
[0147] Dynamic weight calculation: The weight of each expert is dynamically calculated based on the credibility score, query relevance and expert collaboration characteristics.
[0148] Conflict detection and coordination: Automatically detect knowledge conflicts between different experts' outputs and coordinate them.
[0149] Fusion strategy optimization: Continuously optimize the fusion strategy based on feedback from fusion results.
[0150] The core mathematical model of expert fusion is shown below.
[0151] Expert credibility calculation: Calculated using the following formula.
[0152] C_i=ω1·Domain_Coverage_i+ω2·Historical_Accuracy_i+ω3·Certainty_ Score_i, in, C_i is the credibility of expert model i, Domain_Coverage_i For professional coverage, Historical_ Accuracy_i is the historical accuracy, Certainty_Score_i is the certainty of the current output, ω1, ω2, ω3 are their respective weights.
[0153] Dynamic weight allocation: W_i= softmax(λ1·C_i+λ2·Relevance_i+λ3· Complementarity_i), Among them, λ1, λ2, λ3 are learnable weights, softmax() is the probability distribution function, W_i is the dynamic weight of expert model i, Relevance_i To query relevance, Complementarity_i is the expert collaboration characteristic, that is, the complementarity parameter.
[0154] Conflict reconciliation function: When knowledge conflicts are detected between experts, the reconciliation function is used for coordination:
[0155] O_harmonized = Harmonize({O_i}, {W_i}, Conflict_Matrix). O_harmonized is the harmonized output, Conflict_Matrix captures the conflicting relationships between expert outputs, O_i is the output of expert model i, and Harmonize() is the harmonization function.
[0156] Final output calculation: O_final = ∑(W_i·O_i) + Adjustment(O_harmonized). O_final is the fused answer, and Adjustment() is the adjustment function.
[0157] This adaptive fusion algorithm can effectively handle knowledge conflicts and complementarity issues in the process of multi-expert collaboration, significantly improving the accuracy and reliability of the system's processing of complex queries.
[0158] The above data processing flow also includes the steps of expert self-assessment and mutual review, which is based on the expert capability self-assessment and error correction mechanism provided by this embodiment for the large language model. Figure 7 Schematic diagram of the self-assessment and mutual assessment architecture of the embodiment of the present invention, such as Figure 7 As shown, the self-assessment mechanism is as follows.
[0159] Each expert model has self-assessment capabilities and evaluates the quality of its own output through the following indicators:
[0160] Knowledge coverage: evaluates the extent to which the expert covers the knowledge required for the current query.
[0161] Coverage_score = Overlap(Query_knowledge_req, Expert_knowledge). Coverage_score is the knowledge coverage, Query_knowledge_req is the knowledge requirement of the query information, Expert_knowledge is the knowledge base of the expert model, and Overlap() is the overlap function.
[0162] Reasoning soundness: Evaluate the logical rigor of the reasoning process.
[0163] Reliability_score = Consistency(Reasoning_steps) * Completeness(Evidence_chain). Reliability_score is the reliability of reasoning, Consistency(Reasoning_steps) is the consistency of reasoning steps, and Completeness(Evidence_chain) is the completeness of the evidence chain.
[0164] Uncertainty quantification: Quantify the uncertainty of the output.
[0165] Uncertainty = 1 - confidence(Output) + entropy(Distribution). Uncertainty is the uncertainty score, confidence(Output) is the output confidence, and entropy(Distribution) is the distribution entropy.
[0166] The process of expert mutual review and error correction is as follows.
[0167] Output mutual review: The output of expert model A is reviewed by expert model B.
[0168] Review(Expert_B, Output_A) → {Correct, Incorrect, Uncertain}. Review() is a peer review feedback function, Expert_B (expert B's judgment), and Output_A (output result A). Correct indicates that expert model B's judgment is completely consistent with expert model A's output result A. Incorrect indicates that there is a clear conflict between expert model B's judgment and output result A. Uncertain indicates that expert model B's judgment is partially consistent with output result A, but there is ambiguity or insufficient evidence.
[0169] Correction Suggestion Generation: When an error is discovered, a specific correction suggestion is generated. if Review == Incorrect: Correction_suggestion = Generate_correction(Expert_B, Output_A). Review == Incorrect means that the correction is triggered when the mutual review feedback result is Incorrect. This means that correction is triggered when there is a clear conflict between the judgment of expert model B and output A. Generate_correction() is the correction suggestion generation function. Correction_suggestion is the correction suggestion.
[0170] Consensus formation: Expert consensus is formed through multiple rounds of mutual review and discussion.
[0171] Consensus = Deliberation({Experts}, {Outputs}, {Reviews}, max_rounds=3). Consensus is the comprehensive consistency score, {Experts} is the expert set, {Outputs} is the expert output set, {Reviews} is the peer review feedback set, and max_rounds=3 means the maximum number of peer review rounds is 3.
[0172] Feedback learning: Feedback learning is conducted based on error correction results to continuously improve expert capabilities.
[0173] Update_expert(Expert_A, Correction_feedback). Expert_A is the expert model A, Correction_feedback is the correction feedback, and Update_expert() is the knowledge update of the expert model.
[0174] This self-assessment and error correction mechanism greatly enhances the model's self-verification capabilities and improves the reliability and credibility of the system. It is particularly suitable for application scenarios such as trade, which are highly professional and have low fault tolerance.
[0175] The large language model of the HMoE architecture also needs to be trained in stages. In view of the characteristics of the HMoE architecture, this implementation method designs a three-stage training strategy, such as Figure 8 As shown, Figure 8 It is a schematic diagram of the training architecture of an embodiment of the present invention.
[0176] Specifically including: Expert pre-training stage: each expert model is independently pre-trained based on domain-specific data.
[0177] Router training phase: fix expert parameters and train expert routers.
[0178] Overall fine-tuning stage: Jointly optimize the expert model and router, where α is a small learning rate coefficient to prevent overfitting.
[0179] Multi-objective joint optimization: A multi-objective joint optimization method is adopted during the training process, and the objective function is designed as follows.
[0180] L=ζ1·L_task+ζ2·L_routing+ζ3·L_diversity+ζ4·L_consistency .in: L_ task Is the main task loss function; L_routing is the routing quality loss function; L_diversity is the expert diversity loss function; L_consistency is the knowledge consistency loss function; ζ1, ζ2, ζ3, ζ4 is the weight coefficient for balancing each loss item.
[0181] This multi-objective optimization ensures the balanced development of the model in terms of task performance, routing quality, expert diversity, and knowledge consistency.
[0182] Computational efficiency can also be optimized. To improve the computational efficiency of HMoE, this project uses the following technologies.
[0183] Sparse activation: Only some experts (usually 2-4) are activated for each inference, significantly reducing computational overhead.
[0184] Expert parallel computing: Leverage GPU parallel capabilities to simultaneously compute the outputs of multiple experts.
[0185] Dynamic batch processing: Adaptively adjust the batch size based on query complexity.
[0186] Model quantization: Perform Int8 quantization on non-critical experts to reduce memory usage.
[0187] Through these optimization technologies, the HMoE architecture maintains high accuracy while controlling inference latency within the range of 100-500ms, meeting the needs of real-time applications.
[0188] The aforementioned knowledge graph is an important supporting technology for the HMoE architecture. This implementation also constructs a specialized trade knowledge graph. The construction of the trade knowledge graph includes the following steps.
[0189] Knowledge extraction: Extract entities and relationships from authoritative materials such as trade regulations, agreement texts, and professional books.
[0190] Knowledge fusion: Integrate knowledge from different sources to eliminate redundancy and contradiction.
[0191] Knowledge reasoning: Expanding the graph through rule reasoning.
[0192] Knowledge Verification: Verify knowledge correctness through expert review and automatic consistency checking.
[0193] Figure 9 This is a schematic diagram of the knowledge graph architecture of the embodiment of the present invention. The core structure of the trade knowledge graph is as follows: Figure 9 As shown, it includes the following main entity types: Trade Terms (Term), Trade Rules (Rule), Rule Conditions (Condition), Organization (Organization), Country / Region (Country / Region), Product Classification (Product), and Trade Process (Process), as well as a variety of relationship types, such as Definitions (defines), Contains (contains), Dependencies (depends_on), Applies to (applies_to), Exemptions (exempts), Amendments (amends), and Replacements (replaces).
[0194] Finally, in this embodiment, when implementing and deploying the large language model of the HMoE architecture, Figure 10 The architecture shown in the figure is constructed. Specifically, it includes the following core services: Expert Service Cluster: A service cluster where each expert model is independently deployed. Routing Service: Responsible for request distribution and expert selection. Fusion Service: Handles the integration of multiple expert outputs. Knowledge Graph Service: Provides knowledge query and reasoning support. Quality Assessment Service: Implements self-assessment and error correction capabilities. API Gateway: Unified service entry and interface management.
[0195] Specifically, this embodiment also provides the key parameters of the HMoE system that has been implemented in the trade field as follows.
[0196] The system hardware environment uses a distributed computing cluster consisting of 16 GPU nodes (NVIDIA A100 80GB), each equipped with 128GB of memory and 1TB of SSD storage. The software environment uses Python 3.8 for model training and inference, based on the PyTorch framework.
[0197] Expert model division: Professional dimension layer: Customs regulations expert: Focus on trade policies and tariff calculations of various countries. Trade agreement expert: Deal with trade rules and trade agreements. Logistics and transportation expert: Optimize international transportation routes and costs.
[0198] Functional dimension layer: Risk assessment experts: Analyze trade risks. Contract review experts: Parse trade contract terms. Decision support experts: Generate trade strategy recommendations.
[0199] Knowledge Graph Construction: Data sources include trade databases, public trade documents, and historical corporate transaction records. Construction tools include the Neo4j graph database, which contains over 100,000 nodes (concepts) and over 500,000 edges (relationships). This facilitates the rapid and efficient construction of a high-coverage knowledge graph.
[0200] Expert Routing Mechanism: Routing algorithm: A two-layer routing algorithm based on an attention mechanism (Top-K professional layer experts + Top-M functional layer experts). For example, when a user enters "Tariff and risk analysis for exporting electronic products from country A to country B," the route is to a customs regulations expert (professional layer) and a risk assessment expert (functional layer).
[0201] Training and Optimization: Training data: 1 million trade-related question-answer pairs and documents. Loss function: Cross-entropy loss + expert load balancing loss (λ = 0.1). Training cycle: 50 epochs, batch size 256.
[0202] The hierarchical hybrid expert architecture large model system and method proposed in this embodiment bring the following significant beneficial effects through innovative technical solutions.
[0203] The coverage and accuracy of professional knowledge are significantly improved: through a two-level expert organization structure and refined expert division, the present invention can cover professional field knowledge more comprehensively and accurately.
[0204] The reasoning ability for complex problems is greatly enhanced: Through the knowledge graph-based expert routing mechanism and adaptive expert fusion algorithm, this implementation can more accurately handle complex problems, especially those involving the intersection of multi-domain knowledge.
[0205] Reliability and stability are significantly improved: Through the expert ability self-assessment and error correction mechanism, this implementation method significantly reduces the system error rate.
[0206] Optimized computing resource utilization efficiency: Through technologies such as sparse activation and parallel computing, this implementation significantly reduces computing overhead while maintaining model performance.
[0207] Enhanced explainability: Through the explainable output layer and evidence chain tracking mechanism, this implementation improves the transparency of system decisions.
[0208] Improved adaptability and scalability: The modular design and domain knowledge enhancement mechanism of this implementation enable the system to adapt more easily to new professional fields and tasks.
[0209] Based on the above-mentioned data processing method of a large language model based on an improved hybrid expert architecture provided by an embodiment of the present invention, an embodiment of the present invention also provides a data processing device of a large language model based on an improved hybrid expert architecture, which includes.
[0210] The input module is used to process the input query information through the semantic understanding module to obtain the understanding result.
[0211] The routing module is used to send the query information to the corresponding multiple expert models based on the understanding results and the pre-stored knowledge graph, and the expert models obtain the corresponding output results. Among them, the multiple expert models are arranged in multiple layers based on the improved hybrid expert architecture. The expert models of different layers are obtained according to different expert dimensions. At least one expert model in each layer of expert models processes the query information and obtains the output results. The professional field of the knowledge graph matches the professional field of the query information.
[0212] The fusion module is used to fuse the output results of multiple expert models to obtain the answer results corresponding to the query information.
[0213] The data processing device for a large language model based on an improved hybrid expert architecture in this embodiment utilizes a multi-layered expert model, divided according to different expert dimensions, to generate output results based on query information. Finally, these output results are integrated to generate the corresponding answer, resulting in higher accuracy and professionalism. Furthermore, based on the semantic understanding of the query information and the knowledge graph matching the professional domain, the query information is routed and sent to the expert model in the improved hybrid expert architecture, achieving even higher accuracy.
[0214] An embodiment of the present invention further provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform the method of the embodiment of the present invention.
[0215] The embodiments of the present invention further provide a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention.
[0216] An embodiment of the present invention further provides an electronic device comprising: at least one processor; and a memory communicatively coupled to the at least one processor. The memory stores a computer program executable by the at least one processor, wherein the computer program, when executed by the at least one processor, causes the electronic device to perform the method of an embodiment of the present invention.
[0217] refer to Figure 11 , a structural block diagram of an electronic device that can be used as a server or client of an embodiment of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0218] like Figure 11 As shown, the electronic device includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. RAM 1103 may also store various programs and data required for the operation of the electronic device. Computing unit 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.
[0219] Multiple components within the electronic device are connected to the I / O interface 1105, including an input unit 1106, an output unit 1107, a storage unit 1108, and a communication unit 1109. The input unit 1106 can be any type of device capable of inputting information into the electronic device. The input unit 1106 can receive input numeric or character information and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 1107 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1108 may include, but is not limited to, a magnetic disk or an optical disk. The communication unit 1109 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0220] Computing unit 1101 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1101 include, but are not limited to, a CPU, a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1101 performs the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention may be implemented as a computer program tangibly embodied in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device via ROM 1102 and / or communication unit 1109. In some embodiments, computing unit 1101 may be configured to perform the above-described methods by any other suitable means (e.g., via firmware).
[0221] The computer programs for implementing the methods of the embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0222] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0223] The above-described embodiments merely represent several implementation methods of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of protection. It should be noted that a person of ordinary skill in the art would be able to make various modifications and improvements without departing from the scope of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A data processing method for a large language model based on an improved hybrid expert architecture, characterized in that: include: The input query information is processed by the semantic understanding module to obtain the understanding result; Based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models, and the expert models obtain corresponding output results, wherein the multiple expert models are arranged in multiple layers based on an improved hybrid expert architecture, and the expert models in different layers are obtained according to different expert dimensions. At least one expert model in each layer of the expert model processes the query information and obtains an output result, and the professional field of the knowledge graph matches the professional field of the query information; The output results of the multiple expert models are fused to obtain the answer result corresponding to the query information.
2. The method according to claim 1, characterized in that Based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models, and the expert models obtain corresponding output results, including: According to the understanding result, the query information is mapped to the knowledge graph to determine relevant knowledge nodes of the query information; Calculating the degree of fit between the query information and each of the expert models based on the relevant knowledge nodes; The expert models that meet the requirements are included in the expert model set and screened to obtain the expert combination; The query information is sent to the multiple expert models of the expert combination, and the multiple expert models of the expert combination output corresponding output results.
3. The method according to claim 2, characterized in that Calculating the degree of fit between the query information and each of the expert models based on the relevant knowledge nodes includes: Based on the relevant knowledge nodes, calculating the domain relevance between the query information and the expert model; Determining the similarity between the semantic features in the understanding result of the query information and the semantic profile of the expert model; Obtaining historical scores of the expert model for similar queries to which the query information belongs; The degree of adaptation is determined based on the domain relevance, the similarity, and the historical score in combination with corresponding dynamic weights.
4. The method according to claim 2, characterized in that The expert models that meet the requirements are included in the expert model set and screened to obtain the expert combination, including: Determining a plurality of candidate subsets from the expert model set by a screening algorithm; Calculating the knowledge coverage of the expert model in each candidate subset for the query information, as well as the knowledge redundancy between different expert models; According to the knowledge coverage and the knowledge redundancy, an optimal candidate subset is calculated using an optimal solution algorithm as the expert combination.
5. The method according to claim 2, characterized in that The query information is sent to the multiple expert models of the expert combination, and the multiple expert models of the expert combination output corresponding output results, including: Inputting the query information into the expert model at the first layer in the expert combination and outputting a first result, wherein the first layer is the first layer of the improved hybrid expert architecture; Sending the first result to the expert model of the next layer, outputting a second result, and inputting the second result layer by layer to the expert model of the next layer in the expert combination to obtain output results of multiple expert models of the expert combination in each layer; Among them, the improved hybrid expert architecture includes at least: an upper professional dimension layer, and a lower functional dimension layer. The expert model of the professional dimension layer includes at least one of the following: customs experts, tax experts, logistics experts, and financial experts. The expert model of the functional dimension layer includes at least one of the following: knowledge interpretation experts, rule reasoning experts, risk assessment experts, and decision support experts.
6. The method according to claim 5, characterized in that Inputting the query information into the expert model at the first level of the expert combination and outputting a first result includes: Embed the query information into professional terms through a domain professional term embedding layer to obtain a professional term embedding result; Processing the term embeddings using an adaptive multi-head attention layer to obtain a context-aware representation. Mapping the structured expertise in the context-aware representation to a model parameter space using a knowledge-parameter mapping engine to obtain a knowledge-enhanced representation; Obtaining a corresponding first result based on the knowledge-enhanced representation using an interpretable output layer; Among them, the expert model of the next layer receives the results of the upper layer, and uses the domain professional term embedding layer, the adaptive multi-head attention layer, the knowledge-parameter mapping engine and the interpretable output layer to obtain the corresponding output results.
7. The method according to claim 1, characterized in that The output results of the multiple expert models are fused to obtain the answer result corresponding to the query information, including: Based on the professional coverage of the knowledge graph of the expert model for the knowledge nodes of the query information, the accuracy of the historical output of the expert model, and the certainty of the output results of the query information, the credibility of the expert model and the corresponding output results is evaluated to obtain the credibility; Calculating a weight of each of the output results based on the credibility, the query relevance between the expertise vector of the expert model and the query vector of the query information, and a complementarity parameter between different expert models; reconciling conflicting output results according to the credibility to obtain an adjusted output result; The answer result is determined based on the adjusted output result and the corresponding weight.
8. The method according to claim 1, characterized in that Based on the understanding results and the pre-stored knowledge graph, the query information is sent to the corresponding multiple expert models. After the expert models obtain the corresponding output results, the following steps are included: Calculating a self-assessment score based on the expert model's knowledge coverage of the query information, and the inference reliability and quantitative uncertainty of the output result; Through expert peer review, we determine whether there are errors in the output results of the expert model. If errors are found, the corresponding peer review experts will provide correction suggestions for error correction and feedback learning; Through multiple rounds of mutual review, the output results of multiple expert models that form a consensus are obtained, which are then used for subsequent fusion to obtain the answer result.
9. The method according to claim 1, characterized in that Based on the understanding result and the pre-stored knowledge graph, the query information is sent to corresponding multiple expert models, and before the expert models obtain corresponding output results, the method further includes: Independently pre-train the expert model using training data matching the professional field; Based on the pre-trained expert model, the routing module of the improved hybrid expert architecture is trained, wherein the routing module is used to send the query information to the corresponding multiple expert models based on the understanding results and the pre-stored knowledge graph; Based on the trained routing module and the pre-trained expert model, joint optimization is performed through the target optimization function, wherein the target optimization function includes: answer result loss function, routing quality loss function, expert diversity loss function, and knowledge consistency loss function.
10. A computer program product comprising a computer program, wherein The computer program is used to cause the computer to execute the method according to any one of claims 1 to 9 when executed by a processor of the computer.
11. An electronic device comprising: A processor and a memory storing a program, wherein the program comprises instructions which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Task processing method and system based on multiple expert layers, terminal and medium
CN119416822A
Fusion question and answer method, device and equipment of mixed expert large language model and medium
CN119692477A
Retrieval method, retrieval device and retrieval equipment based on knowledge graph and hybrid expert model
CN120030120A
Cross-border trade risk solution generation method and device, equipment and medium
CN120410232A
Cited By
Information consultation method and system based on artificial intelligence
CN121031796A
Artificial intelligence-based information consulting method and system
CN121031796B