A dynamic selection method for cold start portrait acquisition by an educational intelligent agent with double output of single response

CN122548040APending Publication Date: 2026-08-11BEIJING PROSHINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0008](c) 采集环节与服务环节相互独立,未能利用“教育智能体在交付教学服务的同一次响应中即可顺带产出画像观测”这一教育场景特有的机会

Benefits of technology

[0026]第一,由于被选智能体以单次响应同时产出教学服务内容与画像观测,画像采集嵌入于教学服务交付之内完成,不占用独立的、不交付教学服务的交互轮次。相较于提问选择方案需要消耗专门的、不交付教学的提问轮次来逐一采集各维度,本发明在系统达到画像稳态前不交付教学的轮次开销为零。在本说明书实施例所示 K=4、N=5 的场景下(观测精度 ρ_k=0.5、终止阈值 τ_k=0.20),本发明约经 13 次均同时交付教学服务的交互即使所有关键维度方差降至阈值以下,其间不交付教学的专用采集轮次为 0;而采用每轮提出一个不含教学服务的诊断问题的提问式采集方案,将同样四个维度降至阈值约需 36 个不交付教学的专用提问轮次。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention discloses a dynamic selection method for cold-start profile acquisition using educational agents that achieve dual output in a single response, relating to the field of artificial intelligence education technology. The method constructs a K-dimensional profile uncertainty vector containing mean and variance for new users, using a cluster prior of historical user clustering as the initial value. Each agent in the agent library can simultaneously produce teaching service content and profile observations in a single response, and a K-dimensional information gain vector, independent of the current posterior and bound to the agent, is pre-calibrated and stacked in rows to form an N×K information gain matrix. Before each interaction, an agent is selected based on a comprehensive score of "expected reward + information value − resource cost," with the information value driven by the pre-calibrated information gain matrix. During the exploration phase, a calibration probe targeting the dimension with the maximum current uncertainty is injected into the selected agent. The selected agent simultaneously delivers teaching services and produces a profile observation for that dimension in a single response, which is then precisely updated using Bayesian methods to affect the variance of that dimension. The cold start is considered complete when the variances of all key dimensions fall below a threshold. This invention enables profile acquisition to avoid occupying independent interaction rounds that do not deliver teaching services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence education technology, specifically to a dynamic selection method for cold-start profile acquisition using an educational agent that produces both teaching service content and profile observations in a single response. In the initial stage, when user historical interaction data is scarce or sparse, the method selects an educational agent capable of simultaneously producing teaching service content and profile observations in a single response. Information value is calculated based on a pre-calibrated agent-by-agent, dimension-by-dimensional information gain matrix, and calibration probes are embedded in the selected agent's response. This ensures that profile acquisition does not occupy independent, non-teaching service interaction rounds, thereby reducing the non-teaching acquisition overhead required for the system to reach a steady-state profile. Background Technology

[0002] Multi-agent educational systems based on large language models are gradually replacing single, general-purpose intelligent teaching assistants. These systems typically pre-build a library of intelligent agents with various functions, including Q&A, Socratic guidance, question generation, grading, emotional support, and knowledge graph explanation. The system then selects agents to provide services based on student profiles. During the cold start phase (typically the first 5-30 interactions) when new users enter the system, the quality of agent selection methods based on contextual bandit (such as LinUCB) or semantic similarity degrades to near-random levels due to the sparseness of profile data.

[0003] To alleviate the cold start problem, several solutions have been developed in existing technologies, but all of them have shortcomings: First, scalar uncertainty exploration schemes explore with a single scalar uncertainty—using trade-offs—which cannot characterize the objective fact that "students have different uncertainties across different profile dimensions," resulting in a waste of exploration resources on relatively certain dimensions.

[0004] Secondly, cold start questioning solutions based on preference structures (such as the Pep method published in 2026) model the cold start problem as a "question selection" problem: offline learning of the correlation structure between preference dimensions from group data, online maintenance of dimension-wise Bayesian posteriors for new users, and successively selecting "which preference question to ask the user" based on the criterion of maximizing information gain, inferring the user's complete preference profile with fewer questions.

[0005] The selection scheme for this type of question has the following fundamental limitations, which are precisely what this invention aims to overcome:

[0006] (a) The selected object is an independent question; the questioning action only produces information and does not produce teaching services. Continuously asking diagnostic questions to new users to collect profiles will occupy dedicated interaction rounds that do not deliver teaching services. In educational scenarios, the more rounds required for data collection, the greater the overhead of the rounds that do not deliver teaching services before the system reaches a steady state of profile collection.

[0007] (b) Information gain is calculated online at each step based on the current posterior, and its calculation object is the potential user embedding and propagates with the relevance structure; there is no "dimensional information coverage characterization that is independent of the current posterior and the relevance structure and is bound to a specific heterogeneous service agent". When the optional actions are educational agents with different functions, there is a lack of structured representation that directly links "selecting a certain agent" with "which dimensions in the profile are revealed faster".

[0008] (c) The data collection and service processes are independent of each other, failing to take advantage of the unique opportunity in the educational scenario where “the educational agent can generate a profile observation in the same response as delivering teaching services.”

[0009] Therefore, there is an urgent need for a dynamic selection method for cold-start educational agents, which enables the selected agents to simultaneously produce teaching services and profile observations in a single response. The selection is driven by a pre-calibrated information gain matrix bound to heterogeneous agents, and profile collection is completed in an embedded manner in the service response. Summary of the Invention (I) Purpose of the Invention

[0010] The purpose of this invention is to overcome the shortcomings of the prior art and provide a dynamic selection method for cold start profile collection through a single-response dual-output educational agent. This method simultaneously and embeddedly reduces profile uncertainty while delivering teaching services to new users, so that profile collection does not occupy independent interaction rounds that do not deliver teaching services. (II) Definition of Terms

[0011] Single-response dual-output: This refers to a selected educational agent producing two types of content in a single response—(i) user-oriented teaching service content, and (ii) profile observations obtained after signal acquisition and structured analysis, used for posterior profile updates; both originate from the same response and do not require a separate interaction without teaching services for data acquisition. Educational agents possessing this property are also referred to as dual-output agents in this specification. (III) Core Concept and Distinguishing Technical Features of the Invention

[0012] The core difference between this invention and existing cold-start questioning and selection schemes (such as Pep) lies in the following three complementary technical features, which together constitute the inventive point of this invention:

[0013] Distinguishing Feature 1 (Single Response, Dual Output Mechanism): During the execution phase, the selected educational agent's single response simultaneously produces teaching service content and profile observations. This allows the expected reward representing the value of the teaching service and the information acquisition representing the value of profile collection to be achieved simultaneously in the same response, thus ensuring that profile collection does not occupy a separate interaction round that does not deliver teaching services. This mechanism differs from the execution structure in the question selection scheme, where "the question action only produces information and not teaching services."

[0014] Distinguishing Feature Two (Information Gain Matrix Bound to Heterogeneous Agents and Decoupled from the Posterior): Each agent in the agent library is pre-labeled with an inherent K-dimensional information gain vector m_i. These N vectors are stacked row-wise to form an N×K information gain matrix M. The value of M is bound to a specific heterogeneous educational agent, characterizing "the variance reduction capability of selecting a certain agent for each dimension of the profile." Furthermore, its value is independent of the current user's posterior and preference correlation structure and is determined before selection. This representation differs from the approach in question selection where information gain is based on the current posterior, calculated online regarding potential users, and propagated along with the correlation structure. It is a structured object that only holds true in scenarios with multiple service agents possessing different functions.

[0015] The third distinguishing feature (embedded calibration probe and its driven variance update closed loop): While the selected agent generates the main response, a calibration probe targeting the current dimension with the greatest uncertainty, k*, is injected into it. The response of this probe, after analysis, directly drives a precise update of the Bayesian variance of dimension k*, forming a closed loop of "dimensional selection → probe embedding → observation acquisition → precise update of the variance of this dimension". The inventiveness of this invention lies in this closed loop, not in the form of probe injection itself. (iv) Technical files

[0016] To achieve the above objectives, the present invention provides the following technical solution:

[0017] A dynamic selection method for cold-start profile collection using an educational agent with dual output in a single response includes the following steps:

[0018] S1. Construction of Profile Uncertainty Vector: Construct a K-dimensional profile representation (K ≥ 3) for each target user, including a mean vector μ and a variance vector σ², where σ²_k represents the current estimation uncertainty of the corresponding profile dimension k; the profile dimension includes at least the subject ability estimate, learning motivation level, preference difficulty, and emotion or anxiety tendency;

[0019] S2. Construction of the Agent Library and Information Gain Matrix: Construct an agent library A={a_1, …, a_N} containing N educational agents. Each agent can simultaneously produce teaching service content and profile observations in a single response. For each agent a_i, a K-dimensional information gain vector m_i=(m_{i,1}, …, m_{i,K}) is pre-labeled, where m_{i,k}∈ [0,1] represents the expected variance reduction ratio of profile dimension k after a single interaction is completed by selecting a_i. The N vectors are stacked row-wise to form an N×K information gain matrix M. The value of M is independent of the current user's posterior and is determined before agent selection.

[0020] S3. Cluster Prior Injection: Based on C cluster priors obtained from historical user clustering (C ranges from 5 to 20), new users are assigned to the nearest neighbor cluster and the profile parameters of that cluster are used as the initial profile values;

[0021] S4. Comprehensive Score Calculation: Before the t-th interaction, calculate a comprehensive score for each agent: Score_t(a_i) = R̂_t(a_i) + λ(t) · IV_t(a_i) − δ · Cost(a_i) Where R̂_t(a_i) is the prediction of the value of teaching services, and IV_t(a_i) is the prediction of the value of profile collection, both determined by the pre-calibrated information gain matrix M and the current variance vector. IV_t(a_i) = Σ_{k=1}^{K} w_k · m_{i,k} · σ²_k(t) w_k is the importance weight of profile dimension k and Σ w_k=1; the dynamic tradeoff coefficient λ(t)=λ_0 · (Σ_k σ²_k(t) / Σ_k σ²_k(0)) decays as the total uncertainty of the profile decreases, and λ_0 is the initial exploration intensity;

[0022] S5. Agent Selection, Embedded Probe Injection, and Dual Output of Single Response: Select the agent a* with the highest Score_t; if λ(t) > λ_thresh, identify the current maximum uncertainty dimension k* = argmax_k σ²_k(t), select a diagnostic probe p* for dimension k* from the probe template library T_{a*} corresponding to the agent, and inject it into the response generation context of a* using system prompts; a* generates a single response accordingly, which simultaneously produces: user-oriented teaching service content, and a profile observation for dimension k* derived from the injected probe; the probe length is significantly shorter than the teaching service content and consistent with the main functional semantics of the agent; "diagnostic" means that the probe response can reduce the posterior variance of dimension k* with non-zero mutual information;

[0023] S6. Multi-channel signal acquisition: Acquire multi-channel signals from user feedback triggered by a single response of the selected agent, including: main response feedback signal; probe response signal, which is structured and parsed to obtain the observation value y_{k*} and observation precision ρ_{k*} for dimension k*; and implicit behavior signals; the above acquisition and teaching service delivery originate from the same response, and no independent acquisition interaction without teaching services is initiated;

[0024] S7. Profile Post-hoc Update and Termination Determination: For dimensions k* directly covered by the probe, perform Bayesian exact update based on the collected real observations (y_{k*}, ρ_{k*}); for dimensions k not directly covered by the probe, perform an upper bounded approximate update based on the accompanying signals generated during the user's use of the selected agent service, according to the information gain m_{a*,k} of the agent; if all critical dimensions k ∈ K_critical satisfy σ²_k(t+1) < τ_k, then the cold start is determined to end, λ(t) is set to 0 and probe injection is stopped, otherwise return to S4. (v) Beneficial effects

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] First, since the selected agent simultaneously produces teaching service content and profile observations in a single response, profile acquisition is embedded within the delivery of teaching services, without occupying separate interaction rounds that do not deliver teaching services. Compared to the question selection scheme, which requires dedicated question rounds that do not deliver teaching services to collect each dimension one by one, the overhead of non-teaching rounds in this invention is zero before the system reaches a steady state of profile acquisition. In the scenario of K=4 and N=5 shown in the embodiments of this specification (observation accuracy ρ_k=0.5, termination threshold τ_k=0.20), this invention delivers teaching services simultaneously in approximately 13 interactions, even if the variance of all key dimensions drops below the threshold, with zero dedicated acquisition rounds that do not deliver teaching services during this period; while using a question-based acquisition scheme that proposes a diagnostic question without teaching services in each round, it would require approximately 36 dedicated question rounds that do not deliver teaching services to reduce the same four dimensions to the threshold.

[0027] Second, since the information value is driven by the pre-calibrated matrix M, which is bound to heterogeneous agents and decoupled from the current posterior, the agent selection can structurally associate "selecting a certain agent" with "which dimensions in the profile are revealed faster", so that the selection takes into account both the quality of real-time teaching services and the efficiency of profile revelation. Since M is independent of the posterior and is determined before selection, there is no need to solve for the information gain online for the current posterior at each step.

[0028] Third, through the closed loop of "selecting dimensions → embedding probes → collecting observations → accurately updating the variance of the dimension" formed by the embedded calibration probes, the dimensions covered by the probes obtain accurate posterior updates based on real observations, so that the collected information directly and directionally acts on the most uncertain dimension at present.

[0029] Fourth, replacing single scalar uncertainty with a dimensional variance vector enables the exploration to have dimensional differentiation capabilities, avoiding wasting probe rounds on already relatively certain dimensions.

[0030] Fifth, this invention can be directly embedded as a cold start submodule into existing intelligent agent scheduling systems based on student profiles, and has good portability and scalability.

[0031] It should be noted that the update of the dimension not directly covered by the probe in the above embodiments is a prior compact estimation with an upper limit based on the accompanying signal of the agent service process; the accurate update based on real observations is only applied to the dimension covered by the probe. Therefore, the reduction of the profile variance in this invention is mainly based on the real observations collected by the probe. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the overall process of the method described in this invention;

[0033] Figure 2This is a schematic diagram illustrating the correspondence between the image uncertainty vector and the N×K information gain matrix described in this invention.

[0034] Figure 3 This is a schematic diagram of the comprehensive scoring calculation and agent selection process described in this invention;

[0035] Figure 4 This is a schematic diagram of the dual-output mechanism of embedded calibration probe generation, injection, and single-response described in this invention;

[0036] Figure 5 This is a schematic diagram illustrating the posterior update and cold start termination determination of the image described in this invention;

[0037] Figure 6 This is a schematic diagram of the decay curves of the variance of each dimension of the image with the number of interactions in an embodiment of the present invention. Detailed Implementation

[0038] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited to the following embodiments. Example 1: Cold Start Scenario for K12 Mathematics 1. System Configuration

[0039] The profile dimensions K=4: p_1 (mathematical ability), p_2 (learning motivation), p_3 (preference for difficulty), p_4 (anxiety tendency); the dimension importance weights w=(0.30, 0.25, 0.20, 0.25). The agent library contains 5 agents, each capable of simultaneously producing teaching services and profile observations in a single response: a_1 answering questions, a_2 Socratic guidance, a_3 generating questions, a_4 providing emotional support, and a_5 explaining knowledge graphs; the pre-calibrated information gain vectors for each agent are shown in Table 1. Table 1. Pre-calibrated information gain vectors m_i for each agent

[0040] Initial parameters: λ_0=1.0, λ_thresh=0.10, δ=0, termination threshold for all dimensions τ_k=0.20, K_critical={1,2,3,4}. New user Zhang is assigned to the "Medium to weak mathematical foundation" cluster, with initial profile μ(0)=(0.45, 0.50, 0.40, 0.55), σ²(0)=(1.0, 1.0, 1.0, 1.0). For simplicity, the following calculations assume R̂_t(a_i)≈0.50 is the same for all agents, Cost is ignored, and probe response accuracy ρ_k=0.5. 2. Calibration method for information gain matrix M

[0041] Method A (Expert Annotation): Educational experts score each agent's expected variance reduction ratio in each dimension using the Delphi method and take the average. Method B (Historical Data Calibration): Using historical interaction logs as input, calculate the sample mean for the relative reduction of the posterior variance of the profile dimension k before and after each use of a_i: m_{i,k}=E_history[1 − σ²_k_after / σ²_k_before | agent=a_i]. A combination of A and B is actually used: m_{i,k}_new=(1−η)·m_{i,k}_old + η·m_{i,k}_stat,η ∈ [0.05,0.20]. Once M is calibrated, it is fixed before selection; during selection, it is directly read from a table, and its value does not change with the current user's posterior variance—this is a structural feature that distinguishes this invention from online information gain calculation schemes. 3. Calculation method of probe response accuracy ρ_k

[0042] For rating-based probes (such as 1-5 component scales), ρ_k is the reciprocal of the noise variance of human raters, typically ρ_k=0.5; for open-ended probes, the confidence score conf ∈ [0,1] is output from the structured extraction small model and mapped according to ρ_k=ρ_max·conf (ρ_max ∈ [0.5,1.0]); invalid or empty responses are set to ρ_k=0 and are not updated. 4. Mechanism illustration: Step-by-step calculation of the trajectory (t=0 to t=13)

[0043] The table below illustrates the collaborative computation process of each step in this invention, rather than serving as a performance benchmark against baseline methods. Iterations S4-S7: The probe-covered dimension k* is precisely updated based on actual observations: σ²_{k*}←1 / (1 / σ²_{k*}+ρ. The remaining dimensions are approximately updated based on the selected agent m rows: σ²_k←σ²_k·(1−m_{a*,k}). "Selected" refers to the selected agent, and k* represents the dimension covered by the probe. Table 2. Step-by-step computation trajectory (mechanism illustration) for Example 1

[0044] During the t=0–11 phase (λ>λ_thresh=0.10), probes were continuously injected, with the system selecting a_2 / a_3 / a_4 to work alternately, prioritizing the reduction of the current largest uncertainty dimension. From t=12 onwards, when λ dropped below 0.10, probe injection stopped. At t=13, when the variance of all four key dimensions dropped below τ_k=0.20, the system determined that the cold start had ended. In this trajectory, each interaction simultaneously delivered teaching services to the user; no independent data collection interactions without teaching services were initiated. 5. Comparison of overhead with question-based data collection methods

[0045] In contrast, a question-based data collection scheme that presents a diagnostic question without any teaching services to the user in each round and performs a precise update on the questioned dimensions requires approximately 36 dedicated question rounds to reduce all four dimensions to below τ_k=0.20 under the same ρ_k and threshold values, and none of these rounds deliver teaching services. In comparison, this invention completes the same profile collection in approximately 13 interactions, with 0 dedicated collection rounds that do not deliver teaching services—the difference in collection overhead stems from the single-response dual-output mechanism of this invention, which merges data collection and teaching service delivery into a single response. Example 2: Cold Start Scenario for University Programming Courses

[0046] The profile dimension has been adjusted to K=5, with the addition of p_5 (programming language proficiency); the agent library has been adjusted to include 5 agents capable of single-response dual output: code explanation, code review, debugging guidance, concept explanation, and project planning; the probe template library has been adjusted accordingly, for example, the probe template for the code review agent includes "Have you been exposed to unit testing before?" (indirectly probing using the p_5 dimension). The remaining processes are the same as in Implementation Example 1. Example 3: System Implementation

[0047] This invention can also be implemented as a system, including: a profile uncertainty management module, an information gain matrix management module, a cluster prior injection module, a comprehensive scoring calculation module, an agent selection and probe generation module, a multi-channel signal acquisition module, and a profile posterior update module; the modules communicate with each other through shared storage and message queues. This invention can also be stored as a computer program product in a non-volatile computer-readable storage medium and loaded and executed by a processor.

Claims

1. A dynamic selection method for cold start profiling by a single-response dual-output educational agent, characterized in that This includes the following steps: S1. Construct a K-dimensional profile uncertainty vector for the target user, containing a mean vector μ and a variance vector σ², where K ≥ 3, and σ²_k represents the current estimated uncertainty of the corresponding profile dimension k; S2. Construct an agent library A={a_1, …, a_N} containing N educational agents. Each agent can simultaneously produce user-oriented teaching service content and profile observations for profile updates in a single response. For each agent a_i, a K-dimensional information gain vector m_i is pre-labeled, m_{i,k} ∈ [0,1] representing the expected variance reduction ratio of profile dimension k after selecting a_i to complete an interaction. The N vectors are stacked in rows to form an N×K information gain matrix M. The value of M is independent of the current user's posterior and is determined before agent selection. S3. Based on multiple cluster priors obtained from historical user clustering, new users are assigned to the nearest neighbor cluster, and the profile parameters of that cluster are used as the initial profile values; S4. For each agent, calculate the comprehensive score Score_t(a_i)=R̂_t(a_i)+λ(t)·IV_t(a_i)−δ·Cost(a_i), where the information value IV_t(a_i)=Σ_k w_k·m_{i,k}·σ²_k(t), and λ(t) is a dynamic trade-off coefficient that decays as the total uncertainty of the profile decreases; S5. Select the agent a* with the highest Score_t; if λ(t) is higher than the probe injection threshold λ_thresh, then identify the current maximum uncertainty dimension k*=argmax_k σ²_k(t), select a diagnostic probe p* for dimension k* from the probe template library T_{a*} corresponding to the agent, and inject it into the response generation context of a* in the form of system prompt words; a* generates a single response accordingly, which simultaneously produces user-oriented teaching service content and a profile observation for dimension k* induced by the injected probe. The probe length is significantly shorter than the teaching service content and is consistent with the main functional semantics of the agent; S6. Collect the main response feedback signal, probe response signal, and implicit behavior signal from the user feedback triggered by the single response of the selected agent. Perform structured analysis on the probe response to obtain the observed value y_{k*} and the observation accuracy ρ_{k*}. The collection and teaching service delivery originate from the same response without initiating independent collection interactions without teaching services. S7. For dimensions k* directly covered by the probe, perform Bayesian exact updates based on the actual observations (y_{k*}, ρ_{k*}). For dimensions not directly covered by the probe, perform bounded approximate updates based on the accompanying signals of the selected agent's service process according to its information gain m_{a*,k}. If all critical dimensions k ∈ K_critical satisfy σ²_k(t+1) < τ_k, then the cold start is considered complete, λ(t) is set to 0 and probe injection is stopped; otherwise, return to S4.

2. The method of claim 1, wherein In steps S5 and S6, the same response of the selected agent constitutes both the delivery of teaching services to the user and the output of profile observations for the posterior update of the profile through structured parsing, so that profile collection does not occupy independent interaction rounds that do not deliver teaching services.

3. The method of claim 1, wherein The information gain matrix M is bound to the specific educational agent and is independent of the current user's posterior and preference correlation structure. It is determined before the agent is selected, and the selection is performed by directly looking up m_{i,k} in the table without having to calculate the information gain online based on the current posterior in each interaction. The value of m_i is obtained by one or a combination of the following methods: (a) manually labeled by domain experts based on the agent's functional labels; (b) based on historical interaction data, using the average relative reduction of the posterior variance of each dimension of the profile after using a_i as a statistical estimate; (c) obtained by mixing (a) and (b) through a weighted moving average.

4. The method of claim 1, wherein The information gain vector m_i is differentiated in dimension, such that the difference between the maximum and minimum components of each agent is max_k m_{i,k} − min_k m_{i,k} ≥ 0.3, that is, each agent has its own advantageous dimension in terms of information coverage.

5. The method according to claim 1, characterized in that... The dynamic tradeoff coefficient λ(t) = λ_0·( Σ_kσ²_k(t) / Σ_k σ²_k(0) ), where λ_0 is the initial exploration intensity; the injection of the embedded probe is stopped when λ(t) does not exceed the probe injection threshold λ_thresh.

6. The method according to claim 1, characterized in that... Steps S5 to S7 form a closed loop for the current dimension of maximum uncertainty: identify the dimension of maximum uncertainty k*, inject probes for k* into the selected agent, collect observations for k* induced by the probes, and accurately update the posterior variance of dimension k* based on the observations; the probe template library T_{a_i} is a set of diagnostic templates pre-configured for each agent and indexed by the profile dimension, and each probe template is consistent with the main functional semantics of the agent and is significantly shorter than the teaching service content.

7. The method according to claim 1, characterized in that... The calculation method for the observation accuracy ρ_{k*} in step S6 includes: for rating-based probe responses, the reciprocal of the human rater noise variance is used as ρ_{k*}; for open probe responses, the confidence score conf is output by the structured extraction sub-model and mapped according to ρ_{k*}=ρ_max·conf; for invalid or empty responses, ρ_{k*}=0 is set.

8. The method according to claim 1, characterized in that... In step S7, the update rule for dimension k* directly covered by the probe is: σ²_{k*}(t+1) = 1 / ( 1 / σ²_{k*}(t) + ρ_{k*} ) μ_{k*}(t+1) = ( μ_{k*}(t) / σ²_{k*}(t) + y_{k*}·ρ_{k*} ) · σ²_{k*}(t+1) The approximate update rule for dimension k not directly covered by the probe is σ²_k(t+1)=σ²_k(t)·(1−m_{a*,k}) and μ_k remains unchanged. This approximate update has a lower variance limit to prevent unfounded continuous contraction in the absence of real observations.

9. The method according to claim 1, characterized in that... The profile dimension set includes at least the subject ability estimation dimension, learning motivation level dimension, preference difficulty dimension, and emotion or anxiety tendency dimension; the key dimension set K_critical is a non-empty subset of the profile dimension set and includes at least the subject ability dimension.

10. A dynamic selection system for cold-start profile acquisition using an educational intelligent agent with dual output in a single response, characterized in that... The method comprises: a profile uncertainty management module, an information gain matrix management module, a cluster prior injection module, a comprehensive score calculation module, an agent selection and probe generation module, a multi-channel signal acquisition module, and a profile posterior update module, for performing the method described in any one of claims 1 to 9.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that... When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 9.

12. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that... When the processor executes the computer program, it implements the method according to any one of claims 1 to 9.