Interview method and device based on AI main and standby expert models

By employing a dynamic adaptive routing strategy that combines gated confidence and Sinkhorn double random matrices in the interview process, the problem of inflexible allocation of computing resources in hybrid expert models is solved, achieving efficient allocation of computing resources and load balancing, and improving the accuracy and efficiency of interview results.

CN121301032AActive Publication Date: 2026-01-09BEISEN CLOUD COMPUTING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511852071.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-01-09
Estimated Expiration
2045-12-10

AI Technical Summary

Technical Problem

Existing hybrid expert models suffer from inflexible allocation of computational resources during the interview process, resulting in low computational efficiency and uneven load distribution, making it difficult to achieve optimal allocation of computational resources and model performance.

Method used

An interview method based on an AI master-slave expert model is adopted. A dynamic adaptive routing strategy combining gated confidence and Sinkhorn double random matrix is ​​used to dynamically select the expert set, ensuring adaptive allocation and load balancing of computing resources and avoiding unnecessary computational overhead.

Benefits of technology

It achieves improved computational efficiency while ensuring the accuracy of interview results, avoids unnecessary computational overhead, and ensures balanced load distribution, preventing experts from being idle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301032A_ABST
    Figure CN121301032A_ABST
Patent Text Reader

Abstract

The invention provides an interview method and device based on an AI main and standby expert model, and relates to the technical field of artificial intelligence. When an expert set corresponding to each piece of interview feature information of a candidate object is determined, a main expert is determined based on the sequence of affinity scores and a preset cumulative confidence threshold; and when the number of the main experts is smaller than the preset minimum number of the experts, determining the standby experts through double random matrix conversion of the Sinkhorn algorithm. In this way, dynamic expert activation is driven by accumulating the confidence threshold value, self-adaptive computing resource allocation is achieved, and under the condition that the accuracy of an interview result output by the model is guaranteed, the computing efficiency is improved, and unnecessary computing overhead is avoided; and when the number of the activated experts is insufficient, the expert with the best global balance is selected as the standby expert through double random matrix conversion of the Sinkhorn algorithm, so that the load distribution balance is ensured, and the experts are prevented from being idle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an interview method and apparatus based on an AI master-slave expert model. Background Technology

[0002] In recent years, with the rapid development of Artificial Intelligence (AI) in various fields, such as the application of AI interviewers in the recruitment process, and especially the breakthroughs in basic large-scale models, the demands on model capabilities are increasing daily. However, traditional intensive deep learning models (referred to as dense models) face severe challenges in pursuing higher performance. While increasing the number of model parameters can improve model capacity (i.e., the model's ability to learn and represent complex information), it also leads to an exponential increase in computational resource consumption. For example... Figure 1 (a) and Figure 2 As shown, Dense models require activating all their parameters for computation when processing each input, making it increasingly difficult to train and deploy larger-scale models, and even unsustainable in some scenarios. Furthermore, modern datasets are becoming increasingly diverse and complex, often containing multimodal data and intricate structural relationships, making it difficult for a single Dense model to effectively integrate heterogeneous or even conflicting knowledge.

[0003] To address these challenges, the Mixture of Experts (MoE) model architecture has emerged and regained widespread attention. The core idea of ​​the MoE model is to employ a "divide and conquer" strategy, allowing the model to significantly increase the total number of parameters while keeping the computational cost of each forward propagation (inference) within a controllable range. Its design motivation lies in breaking the strong coupling between traditional model capacity and computational cost, achieving a better balance between model performance and efficiency. Figure 1 (b) and Figure 3 As shown, the MoE model (a sparse model) processes each input by dynamically selecting and activating the most relevant subset of sub-models (i.e., "experts"), avoiding the enormous computational overhead of activating the entire network and reducing latency. Research shows that the MoE architecture can significantly improve model performance and efficiency with fewer computational resources, especially excelling in handling large-scale, multimodal, heterogeneous, and complex data.

[0004] However, existing MoE models typically employ a fixed Top-K expert routing strategy. For each input token (the basic unit of text data, which can be a word, sub-word, or a single character), the routing mechanism (i.e., a gating network) calculates its affinity score with all N experts, selects the K experts with the highest scores for activation, and combines them with shared experts for computation. K is pre-set. However, this routing mechanism lacks flexibility and struggles to achieve optimal allocation of computational resources. Summary of the Invention

[0005] The purpose of this invention is to provide an interview method and apparatus based on an AI master-slave expert model, so as to achieve adaptive allocation of computing resources, improve computing efficiency while ensuring the accuracy of the interview results output by the model, avoid unnecessary computing overhead, and ensure balanced load distribution to avoid idle experts.

[0006] In a first aspect, the present invention provides an interview method based on an AI master-slave expert model, comprising: Obtain at least one interview characteristic information of the candidate; The affinity score between each interview feature and multiple pre-set candidate experts is calculated. Based on the affinity scores of each candidate expert for each interview feature, the expert set corresponding to each interview feature is determined; wherein, the expert set includes the main experts determined based on the ranking of affinity scores and a preset cumulative confidence threshold; when the number of main experts is less than the preset minimum number of experts, the expert set also includes backup experts determined by the double random matrix transformation through the Sinkhorn algorithm. By activating the expert set corresponding to each interview feature information, the interview results of the candidate are obtained by calculating at least one interview feature information.

[0007] In an optional implementation, an affinity score is calculated between each interview feature and multiple preset candidate experts, including: For each interview feature, the original routing score of the interview feature for each candidate expert is calculated based on the interview feature and the expert gating vector of each candidate expert. Based on the raw routing score of each candidate expert using interview feature information, the affinity score of the interview feature information for each candidate expert is obtained by normalization using the Sigmoid function.

[0008] In an optional implementation, the formula for calculating the original routing score is: ; in, Indicates the first t The interview feature information for the first iThe raw routing scores of each candidate expert. Indicates the first t Interview characteristics information, Indicates the first i The expert gating vector of each candidate expert. Indicates the first i Trainable biases for each candidate expert; The formula for calculating affinity score is: ; in, Indicates the first t The interview feature information for the first i The affinity score of each candidate expert. This indicates that the Sigmoid function is applied to z.

[0009] In an optional implementation, based on the affinity score of each candidate expert for each interview feature, the expert set corresponding to each interview feature is determined, including: For each interview feature, the affinity scores of each candidate expert are sorted in descending order to obtain the sorting result. The affinity scores are accumulated according to the sorting result, and the candidate experts whose accumulated affinity scores reach the cumulative confidence threshold, or whose number of candidate experts participating in the accumulation reaches the preset maximum allowed number of experts, are determined as the main experts corresponding to the interview feature. Interview feature information where the number of lead experts is less than the minimum number of experts is identified as simple feature information. With lead experts masked, the normalized score matrix corresponding to the simple feature information is obtained through the double random matrix transformation of the Sinkhorn algorithm. The normalized score matrix includes the normalized score of the simple feature information for each candidate expert. The candidate experts with the largest normalized scores in the normalized score matrix are identified as backup experts corresponding to the simple feature information.

[0010] In an optional implementation, affinity scores are accumulated based on the ranking results, and the candidate experts whose accumulated affinity scores reach the cumulative confidence threshold, or whose number of participating candidate experts reaches the preset maximum allowed number of experts, are determined as the principal experts, including: Extract the affinity scores in order of sorting results; Determine whether the sum of the currently retrieved affinity scores reaches the cumulative confidence threshold; If the cumulative confidence threshold is not reached, determine whether the number of candidate experts corresponding to the currently extracted affinity score has reached the maximum allowed number of experts; If the maximum allowed number of experts has not been reached, repeat the step of retrieving affinity scores in order of sorting results; If the cumulative confidence threshold or the maximum allowed number of experts is reached, the candidate expert corresponding to the currently extracted affinity score will be determined as the principal expert.

[0011] In an optional implementation, with the master expert masked, the normalized score matrix corresponding to the simple feature information is obtained through the double random matrix transformation of the Sinkhorn algorithm, including: Obtain the original routing scores for each candidate expert based on simple feature information, and reset the original routing score corresponding to the master expert to negative infinity; Through exponential transformation, the original routing scores of each candidate expert based on simple feature information are transformed into an exponential transformation score matrix. The exponentially transformed fractional matrix is ​​transformed into a normalized fractional matrix using the Sinkhorn algorithm's double random matrix transformation.

[0012] In an optional implementation, by activating the expert set corresponding to each interview feature information, calculating at least one interview feature information, the interview results of the candidate are obtained, including: For each interview feature, the gating weight of the interview feature on each activated expert in the corresponding expert set is obtained by normalization based on the affinity score of the interview feature on each activated expert. The interview feature information is input into the corresponding expert set for processing to obtain the expert results output by each activated expert; The expert results output by each activation expert are weighted and aggregated based on the gating weights of the interview feature information on each activation expert to obtain the target output result of the interview feature information. Based on the target output of each interview feature information, the interview results of the candidate are generated.

[0013] Secondly, the present invention provides an interview device based on an AI master-slave expert model, comprising: The feature acquisition module is used to acquire at least one interview feature information of the candidate; The score calculation module is used to calculate the affinity score between each interview feature and multiple preset candidate experts; The expert set determination module is used to determine the expert set corresponding to each interview feature based on the affinity score of each candidate expert. The expert set includes the main experts determined based on the ranking of affinity scores and a preset cumulative confidence threshold. When the number of main experts is less than the preset minimum number of experts, the expert set also includes backup experts determined by the double random matrix transformation of the Sinkhorn algorithm. The result determination module is used to calculate at least one interview feature by activating the expert set corresponding to each interview feature information to obtain the interview result of the candidate.

[0014] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the interview method based on an AI master-slave expert model as described in any of the foregoing embodiments.

[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the interview method based on an AI master-slave expert model as described in any of the foregoing embodiments.

[0016] This invention provides an interview method and apparatus based on an AI master-slave expert model. The method includes: acquiring at least one interview feature information of a candidate; calculating the affinity score between each interview feature information and multiple preset candidate experts; determining the expert set corresponding to each interview feature information based on the affinity score of each candidate expert; wherein the expert set includes master experts determined based on the affinity score ranking and a preset cumulative confidence threshold; when the number of master experts is less than a preset minimum number of experts, the expert set also includes backup experts determined by a double random matrix transformation using the Sinkhorn algorithm; by activating the expert set corresponding to each interview feature information, calculating at least one interview feature information, and obtaining the interview result of the candidate. This dynamic expert activation driven by the cumulative confidence threshold achieves adaptive allocation of computational resources, improving computational efficiency and avoiding unnecessary computational overhead while ensuring the accuracy of the interview results output by the model; and when the number of activated experts is insufficient, the double random matrix transformation using the Sinkhorn algorithm selects the most globally balanced expert as a backup expert, ensuring balanced load distribution and avoiding expert idleness. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 Here are schematic diagrams of the network structures of the Dense model and the MoE model; Figure 2 This is a schematic diagram of the architecture of the Dense model; Figure 3 This is a schematic diagram of the architecture of a sparse model; Figure 4This is a schematic diagram of the MoE model architecture; Figure 5 A flowchart illustrating an interview method based on an AI master-slave expert model provided in an embodiment of the present invention; Figure 6 A flowchart illustrating another interview method based on an AI master-slave expert model provided in an embodiment of the present invention; Figure 7 A schematic diagram of the structure of an interview device based on an AI master-slave expert model provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Hybrid Expert (MoE) modeling strategies decompose complex tasks or data spaces, assigning tasks to multiple specialized "expert" networks. Unlike traditional dense models that activate all parameters when processing each input, MoE models dynamically select and activate a small subset of experts most relevant to the current input through a gating network or router. Figure 4 As shown, the MoE model typically includes the following core components: Expert networks: Each expert is typically an independent neural network, and there are many of them. In the Transformer architecture, they replace the feed-forward network (FFN) layer. Expert networks focus on learning unique, specialized knowledge and are responsible for processing domain-specific inputs.

[0021] Shared expert networks: Independent of expert networks, with less data, these shared experts are always active for all inputs and are responsible for processing all inputs. Shared experts focus on cross-domain common knowledge, effectively reducing the overlap of learning content among routing experts and alleviating the problem of knowledge redundancy.

[0022] Gating Network: It receives input and decides which experts to route the input to and how to combine the outputs of these experts, typically generating weights for each expert.

[0023] Routing Mechanism: Selecting experts based on the output of the gating network. The most common method is Top-K routing, which selects the K experts with the highest scores.

[0024] Although current mainstream MoE hybrid expert models are more efficient than traditional intensive architectures, and most of them adopt a fixed Top-K expert routing strategy (such as...), Figure 4 As shown in the diagram, while this approach is simple and effective, it ignores the fact that different tokens may have varying degrees of complexity or diversity. For simple tokens, activating K experts may waste computational resources; while for tokens requiring complex reasoning or knowledge, a fixed number of K experts may not provide sufficient computational power or model capacity, thus limiting the model's performance ceiling. This "one-size-fits-all" approach clearly lacks flexibility and makes it difficult to achieve optimal allocation of computational resources.

[0025] To overcome the limitations of traditional Top-K expert routing strategies, this invention provides an interview method and apparatus based on an AI master-slave expert model. It employs a dynamic adaptive routing strategy combining gated confidence and a Sinkhorn double random matrix, which can improve the large-model inference capability of the AI ​​interviewer. The double random matrix is ​​a square matrix where the sum of the elements in each row and column is equal to 1.

[0026] The core idea of ​​this invention is to drive dynamic expert activation through confidence. Specifically, a cumulative confidence threshold (e.g., 0.8) is set, and the threshold is incremented sequentially based on expert scores until it is met. When the number of activated experts is insufficient, the Sinkhorn algorithm is used to transform the scoring matrix of the remaining experts into an approximately doubly random distribution. This allows for the selection of the most globally balanced expert as a candidate, ensuring a fair load distribution and preventing expert idleness. This ensures that enough experts are activated and provides a more uniform gradient signal for backpropagation.

[0027] Unlike Top-K, which selects a fixed number of experts, the dynamic adaptive routing strategy based on the combination of gated confidence and double random matrices is more refined and intelligent in its adaptive allocation of computational resources compared to the static expert selection strategy. This improves computational efficiency, avoids unnecessary computational overhead, and balances sparsity and balance, thereby enhancing model performance.

[0028] To facilitate understanding of this embodiment, a detailed description of an interview method based on an AI master-slave expert model disclosed in this embodiment of the invention will be provided first.

[0029] This invention provides an interview method based on an AI-powered primary and backup expert model, which can be executed by an electronic device with data processing capabilities. See also... Figure 5 The diagram shows a flowchart of an interview method based on an AI master-slave expert model. The method mainly includes the following steps S510 to S540: Step S510: Obtain at least one interview feature information of the candidate.

[0030] The aforementioned candidates can be applicants for a specific position.

[0031] In some possible embodiments, step S510 above may include: acquiring an interview video of the candidate, the interview video including audio of question answers and interview image data; performing feature extraction on the interview video to obtain at least one interview feature information, the interview feature information may include one or more of text features, audio features, facial expression features, and behavioral features. Feature extraction may be implemented, but is not limited to, through a neural network (such as a Transformer).

[0032] In practice, the spoken answers to questions can be converted into text, and natural language processing techniques can be used to analyze the textual features. By analyzing the vocabulary, grammatical structure, and semantics used by candidates when answering questions, their communication skills, professional knowledge, and logical thinking can be assessed.

[0033] It can perform voice feature recognition on the voice of the answer to the question, such as pitch, rhythm, fluency, etc., to obtain voice features, in order to assess one or more of the candidate's emotional state, level of confidence, and level of nervousness.

[0034] Facial expression recognition and behavioral analysis can be performed on interview image data to obtain expression and behavioral characteristics. Computer vision technology can be used to identify micro-expressions on a person's face and infer the candidate's emotional response, such as excitement, nervousness, and honesty. Nonverbal behaviors of candidates, such as gestures and eye contact, can be analyzed to obtain information about the candidate's personality traits.

[0035] Step S520: Calculate the affinity score between each interview feature and multiple preset candidate experts.

[0036] Candidate experts are pre-trained expert networks (such as neural networks). Different candidate experts process the input differently, and therefore their suitable interview feature information may differ. Thus, it is necessary to match a suitable expert to each interview feature information. An interview feature information can be processed by one or more experts. To determine the experts suitable for the interview feature information, it is necessary to first calculate the affinity score between each interview feature information and multiple pre-set candidate experts. The affinity score is used to characterize the degree of matching between the interview feature information and the corresponding candidate expert, thereby selecting the expert set corresponding to each interview feature information based on the affinity score.

[0037] In some possible embodiments, each expert has a gating vector, which is used to calculate an affinity score with the input token. The expert gating vector is a parameter weight matrix trained and learned to represent the characteristics of the corresponding expert. Based on this, step S520 above may include: for each interview feature information, calculating the original routing score of the interview feature information for each candidate expert based on the interview feature information and the expert gating vector of each candidate expert; and obtaining the affinity score of the interview feature information for each candidate expert by normalizing it using the Sigmoid function based on the original routing score of the interview feature information for each candidate expert.

[0038] A linear computation (e.g., dot product plus bias) can be performed on the input token and the expert gating vector, and then the affinity score of each token to each candidate expert can be obtained through the sigmoid function. Based on this, in one possible implementation, the formula for calculating the original routing score can be: ; in, Indicates the first t The interview feature information for the first i The raw routing scores of each candidate expert. Indicates the first t Interview characteristics information, Indicates the first i The expert gating vector of each candidate expert. Indicates the first i Trainable bias terms for each candidate expert. Right now and The dot product of represents the first . t The first interview characteristic information and the first i The degree of similarity or matching between candidate experts. A non-assisted loss strategy is adopted (it does not directly rely on an externally defined loss function to adjust the bias, but instead optimizes indirectly based on the results of route selection). It only affects the route decision and is used to adjust the activation probability of experts to avoid route collapse (always select several experts with higher weights), thereby dynamically achieving load balancing (if the utilization rate of an expert is low, the bias is increased, and vice versa).

[0039] The original routing score can be normalized to a score in the interval [0,1] using the Sigmoid function. In one possible implementation, the affinity score can be calculated as follows: ; in, Indicates the first t The interview feature information for the first iThe affinity score of each candidate expert. This indicates that the Sigmoid function is applied to z.

[0040] In another possible implementation, a temperature hyperparameter can be introduced when calculating the affinity score. Lower temperatures result in a sharper output, making the selected expert more certain; conversely, higher temperatures result in a smoother output, making the selected expert more random. That is, to control the sharpness or smoothness of the output, a temperature hyperparameter T can be introduced into the sigmoid function. Adjusting this hyperparameter can affect the certainty of the expert's selection. Based on this, the formula for calculating the affinity score can be: ; When T is low: the output of the sigmoid function is sharper, meaning that some experts have much higher affinity scores than others. This makes the model more inclined to select the few experts with the highest affinity scores, thus increasing the certainty of the selection; When T is high: the output of the sigmoid function is smoother, and the differences in affinity scores among all experts decrease. In this case, the probability of choosing any one expert increases, making the expert selection more random.

[0041] Step S530: Determine the expert set corresponding to each interview feature based on the affinity score of each candidate expert for each interview feature; wherein, the expert set includes the main experts determined based on the ranking of affinity scores and a preset cumulative confidence threshold; when the number of main experts is less than the preset minimum number of experts, the expert set also includes backup experts determined by the double random matrix transformation of the Sinkhorn algorithm.

[0042] When determining the expert set corresponding to each interview feature, a primary expert can be selected using Confidence-Based Routing, and backup experts can be selected using Sinkhorn Routing. When selecting the primary expert, the affinity scores of each candidate expert can be sorted in descending order, and then the scores are accumulated starting from the highest score until the cumulative sum is greater than or equal to a preset threshold (cumulative confidence threshold). The candidate experts selected during the accumulation process are marked as the initially selected primary experts, and their original affinity scores and weights are recorded. The weights can be obtained by normalizing the affinity scores of each primary expert. When selecting backup experts, it can be checked whether the number of initially selected primary experts is less than the preset minimum number of experts (min). kIf the number of experts is not less than the minimum number, then there is no need to select backup experts; if the number of experts is less than the minimum number, then all candidate experts that were not initially selected are used, and the selected primary expert is masked (e.g., the original routing score logits of the primary expert is set to negative infinity). Then, Sinkhorn normalization is applied to the logits of these unselected candidate experts to obtain backup experts.

[0043] In some possible embodiments, a maximum allowed number of experts is preset. This maximum allowed number of experts is the maximum number of experts that can be activated, used to prevent the selection and activation of too many experts. The sum of the number of backup experts and the number of primary experts should not be less than the minimum number of experts and not greater than the maximum allowed number of experts. Based on this, the above step S530 may include: for each interview feature information, sorting the affinity scores of each candidate expert for the interview feature information in descending order to obtain a sorting result; accumulating the affinity scores according to the sorting result, and determining the candidate experts whose accumulated affinity scores reach the accumulated confidence threshold, or whose number of candidate experts participating in the accumulation reaches the preset maximum allowed number of experts, as the primary experts corresponding to the interview feature information; determining the interview feature information with a number of primary experts less than the minimum number of experts as simple feature information; in the case of masking the primary experts, obtaining the normalized score matrix corresponding to the simple feature information through the double random matrix transformation of the Sinkhorn algorithm, the normalized score matrix including the normalized score of each candidate expert for the simple feature information; determining the several candidate experts with the largest normalized scores in the normalized score matrix as backup experts corresponding to the simple feature information.

[0044] In one possible implementation, the principal expert can be determined through the following process: First, affinity scores are extracted sequentially according to the ranking results. Second, it is determined whether the sum of the currently extracted affinity scores reaches the cumulative confidence threshold. Third, if the cumulative confidence threshold is not reached, it is determined whether the number of candidate experts corresponding to the currently extracted affinity scores reaches the maximum allowed number of experts. Fourth, if the maximum allowed number of experts is not reached, the step of extracting affinity scores sequentially according to the ranking results is repeated. Fifth, if the cumulative confidence threshold or the maximum allowed number of experts is reached, the candidate expert corresponding to the currently extracted affinity score is determined as the principal expert.

[0045] In one possible implementation, the normalized score matrix corresponding to the simple feature information can be obtained through the following process: obtain the original routing scores of the simple feature information for each candidate expert, and reset the original routing scores corresponding to the master expert to negative infinity to mask the master expert; transform the original routing scores of the simple feature information for each candidate expert into an exponential transformation score matrix through exponential transformation to ensure that all elements in the exponential transformation score matrix (i.e., the exponential transformation scores corresponding to each candidate expert) are non-negative; transform the exponential transformation score matrix into a normalized score matrix through the double random matrix transformation of the Sinkhorn algorithm.

[0046] The above exponential transformation can be expressed as: ; If the first i The original routing score z corresponding to each candidate expert i If the value is -∞ (negative infinity), then its exponential transformation fraction Q i =0.

[0047] The Sinkhorn algorithm is a method that iteratively transforms any non-negative matrix into a double-stochastic matrix (a matrix where the sum of each row and column is 1). By alternately normalizing the rows and columns of the matrix, it transforms the original unbalanced assignment matrix into a balanced, differentiable, and approximately optimal assignment structure. The Sinkhorn algorithm effectively achieves load balancing and globally optimal assignment. The data in the normalized score matrix obtained by the Sinkhorn algorithm represents the probability or weight of being routed to different experts. This matrix allows the input data (i.e., various interview feature information) to be assigned to the corresponding expert networks for processing, effectively selecting backup experts for the interview feature information.

[0048] Step S540: By activating the expert set corresponding to each interview feature information, at least one interview feature information is calculated to obtain the interview result of the candidate.

[0049] For each interview feature, the final weight of each expert in the expert set can be obtained first through normalization. Then, the result is output based on the final weight of each expert in the expert set, and backpropagation is used for optimization. Specifically, the weights of each expert in the expert set corresponding to each Token (i.e., interview feature) (such as the affinity score of the main expert and the affinity score / normalized score of the backup expert) can be normalized to make their sum equal to 1, thus obtaining the gating weight of each expert in the expert set corresponding to each Token. For each Token, the gating weight of the main expert obtained by confidence routing and the gating weight of the backup expert obtained by Sinkhorn routing are combined with the expert output to calculate the final output. Based on the final output of each Token, the interview results of the candidate are generated. According to the interview results of the candidate and the preset real labels, the loss signal is calculated through the selected loss function. According to the loss signal, the network parameters and routing gating parameters of the corresponding experts are updated through backpropagation. The routing gating parameters are used to determine the expert gating vector.

[0050] In some possible embodiments, step S540 above may include: for each interview feature information, obtaining the gating weight of the interview feature information on each activated expert in the corresponding expert set through normalization processing based on the affinity score of the interview feature information on each activated expert; inputting the interview feature information into the corresponding expert set for processing to obtain the expert results output by each activated expert; performing weighted aggregation on the expert results output by each activated expert according to the gating weight of the interview feature information on each activated expert to obtain the target output result of the interview feature information; and generating the interview results of the candidate object based on the target output result of each interview feature information.

[0051] The interview method based on an AI master-slave expert model provided in this invention achieves adaptive allocation of computing resources by driving dynamic expert activation through an accumulated confidence threshold. This improves computational efficiency and avoids unnecessary computational overhead while ensuring the accuracy of the interview results output by the model. Furthermore, when the number of activated experts is insufficient, the Sinkhorn algorithm's double random matrix transformation selects the most globally balanced expert as a backup expert, ensuring balanced load distribution and preventing expert idleness.

[0052] The interview method based on an AI master-slave expert model provided in this embodiment of the invention mainly includes the following steps: 1. Calculate the affinity score between the input token and the expert gating network.

[0053] 2. Sort the affinity scores.

[0054] 3. Determine the initial activation master expert based on the cumulative affinity score and cumulative confidence threshold.

[0055] 4. Check if the number of activated experts meets the minimum requirement (min). k .

[0056] 5. If less than min k For inactive experts, calculate the Sinkhorn normalized probability and select backup experts until the minimum condition is met. k .

[0057] 6. Combine the confidence level and the experts selected by Sinkhorn for weight normalization.

[0058] 7. Return the final route weights and the set of experts that are activated.

[0059] 8. The expert output results are combined with the weights to calculate the final result, and the network parameters are updated through backpropagation.

[0060] For ease of understanding, please refer to the following: Figure 6 This paper introduces the process of determining the expert set in the aforementioned interview method based on an AI-powered master-slave expert model. For example... Figure 6 As shown, this interview method based on an AI master-slave expert model includes the following steps: Step S610: Calculate the affinity score between the token and the expert gating vector.

[0061] Step S620: Sort the expert affinity scores in descending order.

[0062] Step S630: Determine if there is an affinity score that reaches the cumulative confidence threshold. If not, proceed to step S640; if yes, proceed to step S650.

[0063] Step S640: Accumulate affinity scores and select the lead expert. Then repeat step S630.

[0064] Step S650: Determine the lead expert. Then proceed to step S660.

[0065] Step S660: Check if the number of chief experts meets the minimum number of experts. If not, proceed to step S670; if yes, proceed to step S690.

[0066] Step S670: Apply the Sinkhorn routing policy to allocate experts.

[0067] Step S680: Select the expert.

[0068] Step S690: Determine the routing expert set.

[0069] To facilitate understanding, the above-mentioned interview method based on AI master-slave expert models will be introduced as an example below.

[0070] The implementation plan is set as follows: Input token quantity: 3 (i.e., 3 interview feature pieces), denoted as Each vector feature dimension D=4 (actually, there is also the batch size). , B (This refers to the batch size). The token will be calculated with the gating network and the expert network.

[0071] Gated network: N=4, including expert gating vectors , Each expert has a gating vector, which is used to perform a vector dot product with the token to calculate the affinity score.

[0072] Expert network: Experts ultimately selected based on the gating network.

[0073] Cumulative confidence threshold: τ =0.8, the cumulative score threshold for confidence routing. Expert selection stops when this threshold is reached or exceeded. A higher threshold usually means that more experts need to be selected and activated to meet the confidence requirement.

[0074] Minimum number of active experts (i.e., minimum number of experts): mink=2, the minimum number of active experts. If the number of experts selected for confidence routing is less than this, the Sinkhorn backup mechanism will be activated. Configurable.

[0075] Maximum number of activated experts (i.e., maximum allowed number of experts): maxk=3, the maximum number of activated experts. If the number of experts selected for confidence routing exceeds this number, selection will stop to prevent the selection of too many activated experts. Configurable. Entropy can also be introduced as a dynamic loss. Higher entropy indicates that the candidates are more dispersed, with similar scores among experts, resulting in lower expert certainty; conversely, lower entropy indicates that expert scores are more concentrated, meaning that one or a few experts are significantly higher than others, resulting in higher expert certainty. Entropy is used here as a metric to evaluate the distribution of candidate expert scores; the entropy is calculated using the scores of the selected experts and added as an additional term to the total loss function. In this way, the model not only optimizes the main task loss during training but also considers the certainty of expert selection.

[0076] Selecting the Activation Expert Set: The expert set selected when the cumulative probability reaches the cumulative confidence threshold, together with the expert set selected through Fallback (i.e., the backup mechanism), constitutes the final activation expert set of this token.

[0077] Step 1: Calculate the raw routing score Logits for the gated network.

[0078] For each token Calculate the dot product score with all experts: ; in, It is the first t Each token represents a semantic feature of the current input (i.e., interview feature information), such as the output vector from the previous Transformer layer. It is the first i Each expert gating vector is a parameter weight matrix learned during training, representing the features of the expert. This represents the similarity or matching degree (dot product) between the token and the expert. It is the first i Trainable biases for each expert. It is a linear combination score of the token and the expert gating network.

[0079] For ease of demonstration and to simplify the calculation process, let's assume: The expert gating vector (d=4) is: , , , .

[0080] The input token vector is: , , .

[0081] Expert gating bias (initial value, trainable): .

[0082] Taking x1 as an example, calculate the dot product with each expert gate plus the bias: .

[0083] Step 2: Calculate the expert-gated affinity Sigmoid.

[0084] Calculate the input token With each expert gating network The affinity score (i.e., affinity rating) is calculated and normalized to a score within the [0,1] interval using the Sigmoid function. The formula is as follows: ; in, , No. t Feature vectors of input tokens ( d dimension); , No. i One expert-gated vector; , No. i A learnable bias gating system; , The linear combination score of the token and the expert gating network; Applying the Sigmoid function to the original score z maps the score to the (0,1) interval and converts the score into a probability. The expert score output by the Sigmoid function represents the token. t With experts i The affinity is the degree to which the token matches or the probability of matching the expert.

[0085] ; z=0.37 indicates the degree of match between the token and the expert. 1,1 =0.591 is token x 1 was assigned to an expert e Affinity of 1. All. A value close to 1 indicates "greater confidence" in the expert's handling of the matter.

[0086] Other affinities were calculated, and the final affinity calculation results are shown in Table 1 below: Table 1

[0087] Step 3: Select experts based on confidence level (cumulative affinity score).

[0088] by x For example, 1: s i,t The affinity scores, ranked from highest to lowest, are: [0.591, 0.552, 0.527, 0.512].

[0089] Detection s i,t The maximum value in s max Is it greater than or equal to the cumulative confidence threshold τ=0.8? s max If ≥τ, then choose s max The corresponding experts. Conversely, the accumulation continues until it is greater than or equal to τ, or until the preset maximum allowed number of experts is reached (max). k The formula is: ; in, It is a cumulative expert affinity score; This means accumulating the next expert affinity score until the cumulative confidence threshold is met or the maximum allowed number of experts is exceeded.

[0090] The scores are accumulated sequentially until they exceed the threshold τ = 0.8. First expert (0.591) → Cumulative = 0.591; The second expert (0.552) → cumulative = 0.591 + 0.552 = 1.143 ≥ 0.8; The experts that are accumulated are the main experts that are selected and activated, so Top-2 selects experts {e1,e2}.

[0091] Similarly, experts were selected for x2 and x3, and the final expert selection results are shown in Table 2 below: Table 2

[0092] Here, for each token, experts are selected based on their confidence scores, satisfying the condition of ≥2 experts (min). k =2), and does not exceed max k =3, no fallback is needed. This reflects the sparsity of conditional computation (i.e., expanding computation only when needed, thereby reducing the total computational cost).

[0093] Step 4: Normalize expert gating weights.

[0094] The selected expert gating scores are weighted and normalized for use in the weighted expert output. The formula is as follows: ; in, Represents token t In activated experts i The gating values ​​(weights) are used to weight the expert output; This represents the original affinity, used in weight calculation; This represents the total score of all K selected experts activated by the current token, used for normalization; the normalization result satisfies... ,and .

[0095] Token x1 normalized gating weight: ; Token x2 normalized gating weight: ; Token x3 normalized gating weight: .

[0096] The gating weights determine how much each activated expert contributes to the output.

[0097] Step 5: Activate the selection of experts with a number less than mink Then, the fallback mechanism is activated.

[0098] A backup expert will be assigned to each token (based on the balance principle).

[0099] Based on the previous settings, use the following configuration: Input token x4 is [0.9, -0.8, 0.2, 0.5]; The confidence threshold becomes τ=0.7; The minimum number of experts is still min k =2.

[0100] Based on the above formula, the affinity score of token x4 is shown in Table 3 below: Table 3

[0101] Scores sorted from highest to lowest: [0.774, 0.670, 0.625, 0.283].

[0102] The confidence threshold τ = 0.7. The fourth expert, e4, directly exceeds τ, so only one chief expert is selected, but the minimum requirement is not met. k =2, triggering the Sinkhorn fallback mechanism.

[0103] Step 6: Disable Mask to obtain index score.

[0104] Currently, the primary expert e4 is selected based on confidence level. Then, the primary expert e4 is masked and its logits (original dot product score) is set to -∞ to prevent duplicate selection of this expert. Then, one backup expert is selected from the remaining experts e1, e2, and e3.

[0105] Based on the above calculations, before selecting an expert, the inactive logits are [0.71, -0.93, 0.51, 1.23]. After masking, the masked_logits are [0.71, -0.93, 0.51, -inf], which means e4 is set to negative infinity (-∞). Then, a Sinkhorn algorithm is applied to the logits of the remaining experts, and the logits are normalized.

[0106] Exponential transformation ensures all elements are non-negative: if not masked, Q i =exp(z i If z i Q is -∞ i =0.

[0107] The calculation result is: ; .

[0108] Q The score is given to the exponential transformation vector (i.e., the exponential transformation score matrix).

[0109] Step 7: Sinkhorn iterative normalization.

[0110] By iterating through the Sinkhorn regularization multiple times, it is made to approximately satisfy the following formula, so that the sum of its rows and the sum of its columns are 1: , .

[0111] The normalization formula is: .

[0112] The calculation result is: ; .

[0113] There is only one token here, that is If all tokens are equal, then the minimum number of experts is met; if there are other tokens that do not meet the minimum number of experts requirement, such as e5, then... Then, both rows and columns need to be normalized to ensure that the total weight assigned to each token is 1 (per row) and the load receiving ratio of each expert is consistent (per column). This process makes matrix Q approach a double random matrix (both rows and columns are 1) to achieve the goal of a double random matrix and balance the load among multiple fallback tokens. It is the normalized score (i.e., the normalized score).

[0114] Step 8: Select backup experts and normalize their weights.

[0115] The original primary expert was e4, with a score of 0.774. The fallback calculation results are shown in step 7. After sorting, the top-1 expert is selected, i.e., backup expert e1 is chosen, and finally, experts {e4, e1} are selected. Compared to Top-K, which can only select a fixed K, the combination of the above two methods (i.e., confidence-based routing strategy and Sinkhorn routing strategy) provides more flexible and finer-grained control.

[0116] The scores of the two selected experts were: e4 = 0.774, e1 = 0.497.

[0117] Calculate the gating weights: .

[0118] Step 9: Output the routing expert and perform backpropagation gradient update.

[0119] After selecting the above experts, enter the token. x t They will be sent to these expert networks for processing, resulting in their respective outputs. y i,t Experts who are not selected are not included in the calculation. Then, the calculated gating weights and output are used. y i,t Weighted aggregation yields the final output of the MoE layer. y t The final output is then backpropagated through the loss function to update the parameters of the expert gating network and the expert network. This process forms a direct feedback mechanism that influences the weights of the final output. During backpropagation, gradient signals are received, enabling the gating network to learn how to make better routing and weight allocation decisions, thereby optimizing the performance of the entire MoE layer.

[0120] The formula is: .

[0121] In summary, the key technical points of the embodiments of the present invention include: Confidence routing strategy: Abandoning the fixed top-k, it uses a threshold-based confidence score accumulation method to dynamically select experts. That is, it prioritizes experts with high affinity to the input token, accumulates their affinity scores until a preset confidence threshold is reached, thereby identifying a group of "high confidence" master experts and improving decision adaptability.

[0122] Sinkhorn routing (Fallback) strategy: It ensures that sparse experts can still be distributed in a balanced load by using a global normalized distribution. That is, when there are not enough confidence routing experts, Sinkhorn normalization is used to provide a globally balanced probability distribution for the remaining experts, which serves as the basis for selecting supplementary experts and thus obtaining backup experts.

[0123] Expert deduplication and masking mechanism: ensures that fallback experts and confidence experts are complementary and do not conflict; ensures that experts selected by Sinkhorn will not duplicate experts already selected by confidence, avoiding redundant calculations and weight conflicts.

[0124] Weight balancing ensures backpropagation trainability: The final routing weights are normalized to ensure they are valid probability distributions and that gradients can propagate back steadily.

[0125] This approach combines confidence-based routing and Sinkhorn routing strategies, which can further improve routing quality and expert utilization while maintaining the lightweight nature of the DeepSeek-MoE architecture, thereby improving model performance and ultimately enhancing the accuracy and reliability of the final interview results.

[0126] Corresponding to the above-described interview method based on an AI-based primary and secondary expert model, this embodiment of the invention also provides an interview device based on an AI-based primary and secondary expert model. See [link to related document]. Figure 7 The diagram shown illustrates the structure of an interview device based on an AI master-slave expert model. The device includes: The feature acquisition module 701 is used to acquire at least one interview feature information of the candidate; The score calculation module 702 is used to calculate the affinity score between each interview feature and multiple preset candidate experts; The expert set determination module 703 is used to determine the expert set corresponding to each interview feature based on the affinity score of each candidate expert for each interview feature; wherein, the expert set includes the main experts determined based on the ranking of affinity scores and a preset cumulative confidence threshold; when the number of main experts is less than the preset minimum number of experts, the expert set also includes the backup experts determined by the double random matrix transformation of the Sinkhorn algorithm. The result determination module 704 is used to calculate at least one interview feature by activating the expert set corresponding to each interview feature information to obtain the interview result of the candidate.

[0127] The interview device based on the AI ​​master-slave expert model provided in this embodiment of the invention achieves adaptive allocation of computing resources by driving dynamic expert activation through the cumulative confidence threshold. While ensuring the accuracy of the interview results output by the model, it improves computing efficiency and avoids unnecessary computing overhead. Furthermore, when the number of activated experts is insufficient, the Sinkhorn algorithm is used to transform the double random matrix to select the expert with the most global balance as the backup expert, ensuring balanced load distribution and avoiding expert idleness.

[0128] Further, the score calculation module 702 is specifically used for: for each interview feature information, calculating the original routing score of the interview feature information for each candidate expert based on the interview feature information and the expert gating vector of each candidate expert; and obtaining the affinity score of the interview feature information for each candidate expert by normalizing it with the Sigmoid function based on the original routing score of the interview feature information for each candidate expert.

[0129] Furthermore, the formula for calculating the original routing score is as follows: ; in, Indicates the first t The interview feature information for the first i The raw routing scores of each candidate expert. Indicates the first t Interview characteristics information, Indicates the first i The expert gating vector of each candidate expert. Indicates the first i Trainable biases for each candidate expert; The formula for calculating the affinity score is: ; in, Indicates the first t The interview feature information for the first i The affinity score of each candidate expert. This indicates that the Sigmoid function is applied to z.

[0130] Further, the expert set determination module 703 is specifically used for: for each interview feature information, sorting the affinity scores of each candidate expert in descending order to obtain a sorting result; accumulating the affinity scores according to the sorting result, and determining the candidate experts whose accumulated affinity scores reach the accumulated confidence threshold, or whose number of participating candidate experts reaches the preset maximum allowed number of experts, as the main experts corresponding to the interview feature information; determining the interview feature information with a number of main experts less than the minimum number of experts as simple feature information; obtaining the normalized score matrix corresponding to the simple feature information through the double random matrix transformation of the Sinkhorn algorithm when the main experts are masked, the normalized score matrix including the normalized score of each candidate expert for the simple feature information; and determining the candidate experts with the largest normalized scores in the normalized score matrix as backup experts corresponding to the simple feature information.

[0131] Furthermore, the expert set determination module 703 is also used to: sequentially extract affinity scores according to the sorting result; determine whether the sum of the currently extracted affinity scores reaches the cumulative confidence threshold; if the cumulative confidence threshold is not reached, determine whether the number of candidate experts corresponding to the currently extracted affinity scores reaches the maximum allowed number of experts; if the maximum allowed number of experts is not reached, re-execute the step of sequentially extracting affinity scores according to the sorting result; if the cumulative confidence threshold or the maximum allowed number of experts is reached, determine the candidate expert corresponding to the currently extracted affinity score as the master expert.

[0132] Furthermore, the expert set determination module 703 is also used to: obtain the original routing scores of the simple feature information for each of the candidate experts, and reset the original routing scores corresponding to the master expert to negative infinity; transform the original routing scores of the simple feature information for each of the candidate experts into an exponential transformation score matrix through exponential transformation; and transform the exponential transformation score matrix into the normalized score matrix through the double random matrix transformation of the Sinkhorn algorithm.

[0133] Further, the result determination module 704 is specifically used for: for each of the interview feature information, obtaining the gating weight of the interview feature information on each of the activated experts in the corresponding expert set through normalization processing based on the affinity score of the interview feature information on each activated expert; inputting the interview feature information into the corresponding expert set for processing to obtain the expert results output by each activated expert; performing weighted aggregation on the expert results output by each activated expert based on the gating weight of the interview feature information on each activated expert to obtain the target output result of the interview feature information; and generating the interview results of the candidate based on the target output results of each of the interview feature information.

[0134] The interview device based on the AI ​​master-slave expert model provided in this embodiment has the same implementation principle and technical effect as the aforementioned interview method based on the AI ​​master-slave expert model. For the sake of brevity, any parts not mentioned in the embodiment of the interview device based on the AI ​​master-slave expert model can be referred to the corresponding content in the aforementioned interview method embodiment based on the AI ​​master-slave expert model.

[0135] like Figure 8 As shown, an electronic device 800 provided in this embodiment of the invention includes: a processor 801, a memory 802 and a bus. The memory 802 stores a computer program that can run on the processor 801. When the electronic device 800 is running, the processor 801 and the memory 802 communicate through the bus, and the processor 801 executes the computer program to implement the above-mentioned interview method based on an AI master-slave expert model.

[0136] Specifically, the memory 802 and processor 801 mentioned above can be general-purpose memory and processor, without any specific limitations.

[0137] This invention also provides a computer-readable storage medium storing a computer program. When a processor runs this computer program, it executes the interview method based on an AI master-slave expert model described in the preceding method embodiments. The computer-readable storage medium includes various media capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), RAM, magnetic disk, or optical disk.

[0138] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0139] In all examples shown and described herein, any specific values ​​should be interpreted as merely exemplary and not as limitations; therefore, other examples of exemplary embodiments may have different values.

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0142] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0143] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An interview method based on an AI master-slave expert model, characterized in that, include: Obtain at least one interview characteristic information of the candidate; The affinity score between each of the interview feature information and multiple preset candidate experts is calculated; Based on the affinity score of each of the interview feature information for each of the candidate experts, an expert set corresponding to each of the interview feature information is determined; wherein, the expert set includes the main experts determined based on the ranking of affinity scores and a preset cumulative confidence threshold; when the number of the main experts is less than the preset minimum number of experts, the expert set also includes the backup experts determined by the double random matrix transformation of the Sinkhorn algorithm. By activating the expert set corresponding to each of the interview feature information, the interview results of the candidate are obtained by calculating the at least one interview feature information.

2. The interview method based on an AI master-slave expert model according to claim 1, characterized in that, The calculation yields an affinity score between each interview feature and multiple preset candidate experts, including: For each of the interview feature information, the original routing score of the interview feature information for each of the candidate experts is calculated based on the interview feature information and the expert gating vector of each candidate expert; Based on the original routing score of each candidate expert according to the interview feature information, the affinity score of the interview feature information for each candidate expert is obtained by normalization using the Sigmoid function.

3. The interview method based on an AI master-slave expert model according to claim 2, characterized in that, The formula for calculating the original routing score is as follows: ; in, Indicates the first t The interview feature information for the first i The raw routing scores of each candidate expert. Indicates the first t Interview characteristics information, Indicates the first i The expert gating vector of each candidate expert. Indicates the first i Trainable biases for each candidate expert; The formula for calculating the affinity score is: ; in, Indicates the first t The interview feature information for the first i The affinity score of each candidate expert. This indicates that the Sigmoid function is applied to z.

4. The interview method based on an AI master-slave expert model according to claim 1, characterized in that, The step of determining the expert set corresponding to each interview feature based on the affinity score of each candidate expert for each interview feature includes: For each interview feature, the affinity scores of each candidate expert are sorted in descending order to obtain a sorting result; the affinity scores are accumulated according to the sorting result, and the candidate experts whose accumulated affinity scores reach the accumulated confidence threshold, or whose number of candidate experts participating in the accumulation reaches the preset maximum allowed number of experts, are determined as the main experts corresponding to the interview feature. The interview feature information where the number of lead experts is less than the minimum number of experts is identified as simple feature information. With lead experts masked, a normalized score matrix corresponding to the simple feature information is obtained through a double random matrix transformation using the Sinkhorn algorithm. The normalized score matrix includes the normalized score of the simple feature information for each candidate expert. The candidate experts with the largest normalized scores in the normalized score matrix are identified as backup experts corresponding to the simple feature information.

5. The interview method based on an AI master-slave expert model according to claim 4, characterized in that, The step of accumulating affinity scores based on the ranking results, and determining the candidate experts whose cumulative affinity scores reach the cumulative confidence threshold, or whose number of participating candidate experts reaches the preset maximum allowed number of experts, as the main experts, includes: Extract the affinity scores sequentially according to the sorting results; Determine whether the sum of the currently extracted affinity scores reaches the cumulative confidence threshold; If the cumulative confidence threshold is not reached, determine whether the number of candidate experts corresponding to the currently extracted affinity score has reached the maximum allowed number of experts; If the maximum allowed number of experts is not reached, repeat the step of retrieving affinity scores sequentially according to the sorting results; If the cumulative confidence threshold or the maximum number of allowed experts is reached, the candidate expert corresponding to the currently extracted affinity score will be determined as the master expert.

6. The interview method based on an AI master-slave expert model according to claim 4, characterized in that, The step of obtaining the normalized score matrix corresponding to the simple feature information through the double random matrix transformation of the Sinkhorn algorithm, under the condition of masking the master expert, includes: Obtain the original routing scores of each candidate expert based on the simple feature information, and reset the original routing score corresponding to the master expert to negative infinity; Through exponential transformation, the original routing scores of each candidate expert based on the simple feature information are transformed into an exponential transformation score matrix. The exponentially transformed fractional matrix is ​​transformed into the normalized fractional matrix through the double random matrix transformation of the Sinkhorn algorithm.

7. The interview method based on an AI master-slave expert model according to claim 1, characterized in that, The step of activating the expert set corresponding to each of the interview feature information, calculating the interview result of the candidate by activating the expert set corresponding to each of the interview feature information, includes: For each of the interview feature information, the gating weight of the interview feature information on each of the activated experts in the corresponding expert set is obtained by normalization based on the affinity score of the interview feature information on each activated expert. The interview feature information is input into the corresponding expert set for processing to obtain the expert results output by each activated expert; The expert results output by each of the activation experts are weighted and aggregated according to the gating weights of the interview feature information on each of the activation experts to obtain the target output result of the interview feature information; Based on the target output results of each of the interview feature information, the interview results of the candidate are generated.

8. An interview device based on an AI master-slave expert model, characterized in that, include: The feature acquisition module is used to acquire at least one interview feature information of the candidate; The score calculation module is used to calculate the affinity score between each of the interview feature information and multiple preset candidate experts; The expert set determination module is used to determine the expert set corresponding to each of the interview feature information based on the affinity score of each of the candidate experts; wherein, the expert set includes the main experts determined based on the ranking of affinity scores and a preset cumulative confidence threshold; when the number of the main experts is less than the preset minimum number of experts, the expert set also includes the backup experts determined by the double random matrix transformation of the Sinkhorn algorithm. The result determination module is used to calculate the interview result of the candidate by activating the expert set corresponding to each of the interview feature information.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the interview method based on the AI ​​master-slave expert model as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when executed by the processor, performs the interview method based on the AI ​​master-slave expert model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Hybrid expert model routing network optimization method, product, device and medium

    CN118410851A

  • Data analysis method and device, computer equipment, readable storage medium and program product

    CN119670742A

  • Data processing method and system of low-energy-consumption large language model based on momentum mechanism and multiple types of experts

    CN120450054A