Context example selection method and system based on factorial hidden variable and common cause generation
By constructing a factorial latent state space and training with a multi-head variational encoder, task features are decoupled as independent factors. The sample size is calculated by combining the coverage number theory, which solves the bias problem in example selection in large language models, achieves efficient and accurate example selection, and improves the model's performance in complex tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-03-25
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies in large language models cannot effectively distinguish multi-dimensional task features due to the example selection method, resulting in fluctuating model performance. Furthermore, the reliance on manually preset causal directions limits the model's generalization ability in complex tasks.
By constructing a factorial latent state space, decoupling task features into independent latent factor vectors, using a multi-head variational encoder for training and a fully correlated decoupling penalty term, and combining coverage number theory to calculate the sample size, accurate example selection is achieved.
It improves the performance of large language models in few-shot reasoning in complex tasks, ensures the relevance and comprehensiveness of example selection, reduces retrieval costs, and improves the accuracy and interpretability of feature representation.
Smart Images

Figure CN121920353A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to a method and system for selecting contextual examples based on factorial latent variables and common factors. Background Technology
[0002] With the rapid development of Large Language Models (LLMs), contextual learning has become a mainstream paradigm that can adapt to downstream tasks without requiring parameter updates. In this process, users guide the model to complete specific prediction tasks by providing a small number of input-output examples in prompts. However, real-world applications show that the inference performance of LLMs is extremely sensitive to the selection and arrangement of examples. Due to the lack of systematic selection criteria, random or simple similarity-based examples often lead to drastic fluctuations in model performance, especially in complex tasks involving multi-step inference or specific format requirements. Automatically selecting the optimal combination of examples from massive amounts of data has become a major bottleneck restricting the deployment of large-scale models.
[0003] The effectiveness of contextual learning stems from the model's ability to implicitly infer latent task concepts from observed examples. However, existing theoretical modeling often oversimplifies this complex task concept into a single, static latent variable. The single-variable assumption ignores the inherent structured nature of natural language tasks; a complete task is actually determined by multiple independent latent factors, such as format paradigms, domain semantics, and reasoning logic. When examples vary in quality across these dimensions, traditional example selection methods based on single-variable estimation fail to distinguish these subtle differences, thus making it difficult to capture the full picture of the task.
[0004] Existing technologies suffer from two main technical shortcomings. First, in feature modeling, existing methods employ point estimation strategies to compress all task information into a single vector, leading to confusion between features of different dimensions and an inability to identify and combine functionally complementary examples. Second, regarding causality, existing methods rely on human experience to presuppose a binary causal direction where the input determines the label or the label determines the input. This rigid assumption violates the essence of generative models, ignoring the fact that in the Transformer architecture, the input text and output label are essentially observations co-generated by the underlying task context.
[0005] The aforementioned shortcomings prevent existing example selection methods from providing accurate task guidance for large language models, making it difficult for them to infer the true underlying task concept from the context. According to statistical learning theory, the predictive performance of a large language model can only approach the theoretical upper limit when its inference of the underlying task concept is completely accurate; however, the inference bias introduced by existing techniques directly limits the generalization ability of large language models when facing complex tasks. Summary of the Invention
[0006] The main objective of this invention is to provide a method and system for selecting contextual examples based on factorial latent variables and common factors, aiming to solve the technical problems of existing technologies that compress task features such as format and semantics into a single static variable, leading to ambiguity in model understanding, and that relying on manually preset input-output causal directions, resulting in difficulties in transfer.
[0007] To achieve the above objectives, this invention provides a context example selection method based on factorial latent variables and common factors. This method is applied to natural language text processing and includes the following steps: Task features are extracted from the original context learning task data of the large language model, and a set of factorial latent variables is defined in the continuous embedding space of the large language model to form a factorial latent state space. The original context learning task data includes the input text and the corresponding output label for context learning. The task features are mapped and decoupled into mutually independent latent factor vectors in the set of latent variables in the latent factor state space, so as to reconstruct the input text and output labels into observations generated by the combined effect of the latent factors. The latent factor vectors are used to capture feature attributes of different dimensions in the context example. A multi-head variational encoder is constructed, and the multi-head variational encoder is trained using an objective function with a fully correlated decoupling penalty term. The multi-head variational encoder is configured to receive the concatenated embedding of input text and output labels and predict the distribution parameters of different latent factor vectors through multiple output heads. A sample size calculation model is constructed based on the coverage number theory, and the total retrieval budget parameters are configured through the sample size calculation model. Based on the total retrieval budget parameters and the trained multi-head variational encoder, a retrieval module is configured. The retrieval module is used to map the test input text to the factorial latent state space and parse the posterior distribution parameter set of the test input text in each latent factor dimension. In response to receiving test input text, the test input text is input to the retrieval module to obtain a posterior distribution parameter set. Based on the posterior distribution parameter set, the target point is determined, and a complementary example set covering all target points is retrieved from the candidate example pool through a centroid approximation strategy.
[0008] Optionally, the step of mapping and decoupling the task features into mutually independent latent factor vectors in the set of latent factorial variables within the latent factorial state space, so as to reconstruct the input text and output labels into observations generated by the combined effect of the latent factors, includes: The task features are mapped to the factorial latent state space, and the task features in the factorial latent state space are decoupled by dimension to obtain a subset of task features. Each task feature subset is mapped to the corresponding vector of the factorial latent variable set to obtain the initial latent factor vector; Define independent standard normal marginal distributions for the initial latent factor vectors, and take the product of the marginal distributions as the joint prior distribution of the factorial latent variable set, as shown in the following formula: in, Describes the set of latent variables for factorial. The joint prior distribution, This represents the number of latent factor vectors. Indicates the first A vector of potential factors, Indicates the first Marginal prior distributions of potential factor vectors This represents a vector with zero mean and a covariance matrix that is the identity matrix. The normal distribution; The initial latent factor vector is optimized based on the joint prior distribution to obtain mutually independent latent factor vectors in the factorial latent state space. A generative topology is established with the set of factorial latent variables as global conditions, and the input text and output labels are concatenated into a joint observation sequence; The input text and output labels are reconstructed into observations generated by the combined effect of the latent factors through the generated topology.
[0009] Optionally, the construction of the multi-head variational encoder, which involves training the multi-head variational encoder using an objective function with a fully correlated decoupling penalty term, includes: An approximate posterior distribution controlled by preset parameters is introduced, and this approximate posterior distribution is decomposed into independent sub-distribution structures corresponding to each latent factor vector, as shown in the following formula: in, Indicates that it is based on preset parameters The approximate posterior distribution of the control. This indicates the input text. Indicates preset parameters. Indicates the output label. Represents the set of latent variables for factorial. Indicates the first The mean vector of the potential factor vectors Indicates the first The variance vector of the latent factor vectors Represents the diagonal covariance matrix; Based on the aforementioned independent sub-distribution structure, a multi-output head architecture for a multi-head variational encoder is designed, and a multi-head variational encoder is constructed. Based on the principles of Bayesian learning, an evidence lower bound optimization objective is constructed, referring to the following formula: in, This indicates the objective term for optimizing the lower bound of the evidence. This represents the expectation operation over an approximate posterior distribution. Representing a given latent variable Time observation data The conditional log-likelihood, The parameters representing the large language model. KL divergence is used to measure the difference between two probability distributions. Indicates the first Approximate posterior distribution of the potential factor vectors, Indicates the first Sampled values of a latent factor vector, Indicates the first The standard deviation of the latent factor vector Let represent the auxiliary noise variable, which follows a standard normal distribution; A fully correlated decoupling penalty term is added to the lower bound of evidence optimization objective to form the objective function, as shown in the following formula: in, This represents the total loss term of the objective function. This represents a hyperparameter that controls the strength of decoupling. This represents the fully correlated decoupling penalty term. Denotes the aggregate posterior distribution of the set of latent variables for factorial. Indicates the first Marginal posterior distribution of each latent factor vector; The multi-head variational encoder is trained using the objective function so that it can predict the distribution parameters of different latent factor vectors through multiple output heads.
[0010] Optionally, the step of constructing a sample size calculation model based on coverage number theory and configuring the total retrieval budget parameters through the sample size calculation model includes: A sample complexity function is constructed based on the coverage number theory. This function is used to calculate the minimum number of examples required to control the inference error within a preset threshold under a pre-set confidence level. The sample complexity function is defined by the following formula: in, Indicates the reliability of the preset settings To control the distribution error to not exceed a preset threshold Minimum required sample size Indicates the distribution error threshold. The complement of the confidence level, Indicates a constant factor. In factorial latent state space , with norm For measurement, radius is The coverage number reflects the factorial latent state space. The geometric complexity, Describing the latent state space of factorial The measure of entropy; The parameter space features of the factorial latent state space are imported into the sample complexity function to construct a sample size calculation model; The sample size calculation model is used to calculate the sample requirement under the single-unit characterization framework, which serves as a benchmark reference value. The sample size calculation model is used to calculate the lower bound of the example under the factorial representation framework; A sample budget allocation inequality is established based on the benchmark reference value and the sample lower limit quantity, and the total retrieval budget parameters are configured based on the sample budget allocation inequality, referring to the following formula: in, This indicates the retrieval of the total budget parameter. Indicates the first The minimum number of representative examples required for each feature dimension This represents the sample requirement under the factorial representation framework. This indicates the sample requirement under the single-unit representation framework.
[0011] Optionally, the retrieval module configuration based on the total retrieval budget parameter and the trained multi-head variational encoder includes: Import the total search budget parameter into the search module to configure the sample search quantity constraint of the search module; Based on the trained multi-head variational encoder, a parameter extraction submodule is built in the retrieval module. This submodule is used to extract the posterior distribution parameter set of the test input text in each latent factor dimension, as shown in the following formula: in, This represents the posterior distribution of the test input text on the m-th latent factor vector. This represents the test input text. This represents the central demand of the test input text on the m-th latent factor vector. This represents the variance vector of the test input text on the m-th latent factor vector. This indicates that the multi-head variational encoder is used for the test input text. The output.
[0012] Optionally, the step of determining the target point based on the posterior distribution parameter set and retrieving a complementary example set covering all target points from the candidate example pool using a centroid approximation strategy includes: The mean and variance parameters of each latent factor dimension are extracted from the posterior distribution parameter set and integrated to form a feature parameter set; The set of feature parameters is used as the feature identifier of the test input text in the factorial latent state space and is determined as the target point. A weighted second-order Wasserstein distance metric is defined to calculate the distributional difference between the test input text and candidate examples in the candidate example pool, as shown in the following formula: in, Indicates the test input text With candidate examples A comprehensive measure of difference in the factorial latent state space. Indicates the first The weight coefficients of each potential factor dimension. Denotes the squared second-order Wasserstein distance between two Gaussian distributions. This indicates that the test input text is at the [number]th [position]. The posterior distribution over the latent factor vectors Indicates the candidate example in the th Posterior distribution over _ potential factor vectors; An optimization function is constructed based on the first-order moment centroid approximation strategy and the second-order Wasserstein distance metric, as shown in the following formula: in, This represents the final set of complementary examples. Represents the candidate example pool, Indicates the first One candidate example, Indicates the size of the candidate example pool. This indicates that the test input text is at the [number]th [position]. The mean of the potential factor vectors Indicates the first The candidate example in the first The mean of the potential factor vectors; Based on the optimization function, candidate examples that can correct set bias are selected from the candidate example pool, and the selected candidate examples are combined to obtain a complementary example set covering all target points.
[0013] Optionally, the step of extracting task features from the original context learning task data of the large language model and defining a set of factorial latent variables in the continuous embedding space of the large language model to form a factorial latent state space includes: We acquire input text and output label pairing data for context learning scenarios in large language models, and perform data cleaning and standardization to form raw context learning task data. The continuous embedding space of the large language model is determined based on the parameter matrix of the large language model, and multi-dimensional task features are extracted from the original context learning task data within the continuous embedding space. Define a set of factorial latent variables consisting of multiple vectors in the continuous embedding space, and introduce orthogonality constraints for the set of factorial latent variables. By combining the set of factorial latent variables with orthogonality constraints and the task characteristics, a factorial latent state space is formed.
[0014] Furthermore, to achieve the above objectives, this invention also proposes a context example selection system based on the generation of latent factors and common factors, which applies the context example selection method based on the generation of latent factors and common factors as described above. The context example selection system based on the generation of latent factors and common factors includes: The latent space construction module is used to extract task features from the original context learning task data of the large language model and define a set of factorial latent variables in the continuous embedding space of the large language model to form a factorial latent state space. The original context learning task data includes the input text and the corresponding output labels for context learning. The task reconstruction module is used to map and decouple the task features into mutually independent latent factor vectors in the set of latent variables in the latent factorial state space, so as to reconstruct the input text and output label into observations generated by the combined effect of the latent factors. The latent factor vectors are used to capture feature attributes of different dimensions in the context example. The encoding configuration module is used to construct a multi-head variational encoder. The multi-head variational encoder is trained using an objective function with a fully correlated decoupling penalty term. The multi-head variational encoder is configured to receive the concatenation and embedding of input text and output labels and predict the distribution parameters of different latent factor vectors through multiple output heads. The budget configuration module is used to construct a sample size calculation model based on the coverage number theory, and to configure the total budget parameters for retrieval through the sample size calculation model. The retrieval configuration module is used to configure the retrieval module based on the total retrieval budget parameter and the trained multi-head variational encoder. The retrieval module is used to map the test input text to the factorial latent state space and parse the posterior distribution parameter set of the test input text in each latent factor dimension. The example retrieval module is used to respond to receiving test input text, input the test input text into the retrieval module, obtain a posterior distribution parameter set, determine the target point based on the posterior distribution parameter set, and retrieve a complementary example set that combines and covers all target points from the candidate example pool through a centroid approximation strategy.
[0015] Optionally, the task reconstruction module is further configured to map the task features to the factorial latent state space, perform dimensionality decoupling processing on the task features in the factorial latent state space to obtain a subset of task features; map each subset of task features to the corresponding vector of the factorial latent variable set to obtain an initial latent factor vector; define an independent standard normal marginal distribution for the initial latent factor vector, and use the product of each marginal distribution as the joint prior distribution of the factorial latent variable set, referring to the following formula: in, Describes the set of latent variables for factorial. The joint prior distribution, This represents the number of latent factor vectors. Indicates the first A vector of potential factors, Indicates the first Marginal prior distributions of potential factor vectors This represents a vector with zero mean and a covariance matrix that is the identity matrix. The normal distribution; The initial latent factor vector is optimized based on the joint prior distribution to obtain mutually independent latent factor vectors in the factorial latent state space; a generative topology is established with the set of factorial latent variables as global conditions, and the input text and output labels are concatenated into a joint observation sequence; the input text and output labels are reconstructed into observations generated by the combined action of the latent factors through the generative topology.
[0016] Optionally, the encoding configuration module is further configured to introduce an approximate posterior distribution controlled by preset parameters, and decompose the approximate posterior distribution into independent sub-distribution structures corresponding to each latent factor vector, as shown in the following formula: in, Indicates that it is based on preset parameters The approximate posterior distribution of the control. This indicates the input text. Indicates preset parameters. Indicates the output label. Represents the set of latent variables for factorial. Indicates the first The mean vector of the potential factor vectors Indicates the first The variance vector of the latent factor vectors Represents the diagonal covariance matrix; Based on the aforementioned independent sub-distribution structure, a multi-output head architecture for a multi-head variational encoder is designed, and a multi-head variational encoder is constructed. An evidence lower bound optimization objective is constructed based on the Bayesian learning principle, referring to the following formula: in, This indicates the objective term for optimizing the lower bound of the evidence. This represents the expectation operation over an approximate posterior distribution. Representing a given latent variable Time observation data The conditional log-likelihood, The parameters representing the large language model. KL divergence is used to measure the difference between two probability distributions. Indicates the first Approximate posterior distribution of the potential factor vectors, Indicates the first Sampled values of a latent factor vector, Indicates the first The standard deviation of the latent factor vector Let represent the auxiliary noise variable, which follows a standard normal distribution; A fully correlated decoupling penalty term is added to the lower bound of evidence optimization objective to form the objective function, as shown in the following formula: in, This represents the total loss term of the objective function. This represents a hyperparameter that controls the strength of decoupling. This represents the fully correlated decoupling penalty term. Denotes the aggregate posterior distribution of the set of latent variables for factorial. Indicates the first Marginal posterior distribution of each latent factor vector; The multi-head variational encoder is trained using the objective function so that it can predict the distribution parameters of different latent factor vectors through multiple output heads.
[0017] This invention constructs a factorial latent state space to refine the modeling of task features, decoupling abstract task concepts into independent latent factors. Through common-cause causal modeling, it abandons the binary causal assumption, treating input text and output labels as observations jointly generated by latent factors. During the training phase, a fully correlated decoupling penalty term is introduced to optimize the variational lower bound, training a multi-head encoder to identify and separate each independent factor. During the inference phase, by calculating the demand distribution of the test input text in each factor space, a complementary example set covering all target factors is retrieved and combined from the candidate pool. This invention effectively mines the multi-dimensional features of contextual examples, accurately capturing the examples through feature extraction, decoupling, and modeling. The core feature attributes enhance the accuracy and interpretability of feature representation; improve the efficiency and accuracy of contextual example selection by rationally allocating retrieval resources through a sample size calculation model and combining a centroid approximation strategy to achieve accurate example selection, avoiding redundancy and omissions, and reducing retrieval costs; strengthen the relevance and comprehensiveness of example selection by constructing complementary example sets to ensure that examples can fully cover all feature targets of the test input, providing high-quality example support for contextual learning of large language models; thus effectively solving the problem of biased example selection in complex contextual learning tasks, significantly improving the few-shot inference performance of large language models, achieving accurate example selection, and improving the matching degree between examples and test input text. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the hardware operating environment of the embodiment of the present invention, which is a context example selection device based on factorial latent variables and common factors. Figure 2 This is a flowchart illustrating the first embodiment of the context example selection method based on factorial latent variables and common factors generation of the present invention. Figure 3 This is a flowchart illustrating the second embodiment of the context example selection method based on factorial latent variables and common factors generation of the present invention. Figure 4 This is a flowchart illustrating the third embodiment of the context example selection method based on factorial latent variables and common factors generation of the present invention. Figure 5 This is a structural block diagram of the first embodiment of the context example selection system based on factorial latent variables and common factors generated by the present invention.
[0020] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0022] Reference Figure 1 , Figure 1 This is a schematic diagram of the device structure selected for the context example of the hardware operating environment based on factorial latent variables and common factors generated in the embodiments of the present invention.
[0023] like Figure 1 As shown, the context example selection device based on factorial latent variables and common factors can include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to implement communication between these components. The user interface 1003 can include a display screen, an input unit such as a keyboard, and can also include standard wired or wireless interfaces. The network interface 1004 can optionally include standard wired or wireless interfaces (such as Wireless-Fidelity (Wi-Fi) interfaces). The memory 1005 can be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 can also optionally be a storage system independent of the aforementioned processor 1001.
[0024] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the selection device for the context example generated based on factorial latent variables and common factors, and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0025] like Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a context example selection program based on factorial latent variables and common factors.
[0026] exist Figure 1 In the context example selection device based on factorial latent variables and common factors, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user. The processor 1001 and memory 1005 in the context example selection device based on factorial latent variables and common factors can be set in the context example selection device. The context example selection device based on factorial latent variables and common factors calls the context example selection program based on factorial latent variables and common factors stored in the memory 1005 through the processor 1001, and executes the context example selection method based on factorial latent variables and common factors provided in the embodiment of the present invention.
[0027] This invention provides a method for selecting contextual examples based on factorial latent variables and common factors generation, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the context example selection method based on factorial latent variables and common factors generated according to the present invention.
[0028] In this embodiment, the context example selection method based on factorial latent variables and common factors includes the following steps: Step S10: Extract task features from the original context learning task data of the large language model, and define a set of factorial latent variables in the continuous embedding space of the large language model to form a factorial latent state space.
[0029] It should be understood that the executing entity of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a terminal electronic device capable of performing the above functions. The following uses the example of selecting a device based on a contextual example generated from factorial latent variables and common factors (selection device) to illustrate this embodiment and the following embodiments.
[0030] It should be noted that the original context learning task data refers to the basic data that the large language model relies on when performing context learning, which may include: the input text for context learning (i.e., the original input information to be processed by the model, such as question description, text fragments, etc.) and the corresponding output labels (i.e., the standard results corresponding to the input text, such as question answers, text classification labels, etc.).
[0031] The factorial latent variable set can be a collection of multiple independent latent variables. Each latent variable corresponds to a feature of a specific dimension in the context example. The variables are unrelated and can be represented independently. The set as a whole is used to comprehensively cover the multi-dimensional feature attributes of the context example.
[0032] The factorial latent state space refers to the feature representation space constructed in the continuous embedding space of a large language model with the factorial latent variable set as the core. It is used to standardize and structure the representation of task features, and provide a unified spatial carrier for subsequent feature decoupling and retrieval.
[0033] Task features can be feature information extracted from the original context learning task data that can characterize the core attributes of the task, covering the semantic features and syntactic features of the input text, as well as the category features and association features of the output label.
[0034] It is understandable that this embodiment addresses the real-world need for large language models to accurately extract multidimensional task features from a limited number of examples in context learning scenarios. It maps abstract task intents into structured vectors in the model embedding space and reconstructs the probabilistic dependencies of context sequences based on autoregressive generation mechanisms.
[0035] When dealing with complex tasks involving multiple constraints such as specific formats, domain-specific knowledge, and reasoning logic, the task representation within a large language model often exhibits a distribution in a high-dimensional space. To address the problem of task features being confused and difficult to distinguish in a single vector representation, a factorial latent variable set is defined. The set consists of It consists of several independent latent factor vectors, which are used to capture and control feature attributes of different dimensions in the context example.
[0036] Furthermore, to improve data reliability and avoid feature omissions and limitations of a single dimension, step S10 above may include: Step S101: Obtain the input text and output label pairing data in the context learning scenario of the large language model, and perform data cleaning and standardization to form the original context learning task data; Step S102: Determine the continuous embedding space of the large language model based on the parameter matrix of the large language model, and extract multi-dimensional task features from the original context learning task data within the continuous embedding space; Step S103: Define a set of factorial latent variables consisting of multiple vectors in the continuous embedding space, and introduce orthogonality constraints for the set of factorial latent variables; Step S104: Combine the set of factorial latent variables with orthogonality constraints and the task characteristics to form a factorial latent state space.
[0037] It is understandable that this embodiment ensures the diversity and scenario adaptability of the original data through the data acquisition process, which fits the actual needs of the scenario; data cleaning removes invalid and abnormal data, improves data quality, and avoids interference from invalid data with subsequent feature extraction and model training; data standardization achieves the unification of data format and dimensions, eliminates the impact of data differences, and provides a standardized and high-quality data source for subsequent task feature extraction and potential state space construction.
[0038] It should be noted that the parameter matrix can be the core matrix used to represent the semantics of text in a large language model, including word embedding matrix, attention matrix, etc. Its dimension and value determine the feature space range of text embedding and are used to determine the continuous embedding space.
[0039] Multidimensional task features refer to feature information extracted from the original context learning task data that covers multiple dimensions, such as the semantic, grammatical, sentiment, and tag association features of e-commerce review texts, which can comprehensively represent the core attributes of the review.
[0040] In some embodiments, a pre-trained large language model is selected, and the word embedding parameter matrix of the model is extracted. Based on the range and dimension of the parameter matrix, the continuous embedding space of the large language model is determined to ensure that the space can accommodate all features of the comment text and sentiment label. The standardized original context learning task data (e.g., comment text + sentiment label) is input into the large language model, and multi-dimensional task features are extracted using the model's built-in embedding layer and feature extraction module.
[0041] It should be noted that orthogonality constraints are mathematical constraints used to constrain the relationship between latent variable vectors. They require that any two latent variable vectors in the set are mutually orthogonal (the dot product is 0). The purpose is to strengthen the independence of latent variable vectors and avoid interference between features of different dimensions.
[0042] Step S20: Map and decouple the task features into mutually independent latent factor vectors in the set of latent factor variables in the latent factor state space, so as to reconstruct the input text and output labels into observations generated by the combined effect of the latent factors.
[0043] It should be noted that the latent factor vector can be the feature vector corresponding to each independent latent variable in the set of latent variables in the factorial latent state space. It is specifically used to capture the feature attributes of a certain dimension (such as semantic dimension, category dimension, etc.) in the context example, and the value of the vector corresponds to the specific representation of the feature of that dimension.
[0044] The observations refer to the equivalent representations of the input text and output labels in the original context learning task data, which are generated by the combined action of multiple latent factor vectors after feature mapping and decoupling. They have a one-to-one correspondence with the original input and output data and retain the core feature information of the original data.
[0045] For any given context task, its latent representation is defined as an ordered set of vectors in this space: In this definition, each Corresponding to a set of continuous vectors to be learned (i.e., latent factor vectors), its dimension It is much smaller than the hidden layer dimension of a large language model.
[0046] It is understood that this embodiment achieves effective decoupling of task features, separating complex task features into multiple independent single-dimensional latent factor vectors, eliminating interference between different feature dimensions, and improving the accuracy of feature representation. By reconstructing the observations, it ensures that the decoupled latent factor vectors can completely retain the core information of the original task data, providing a reliable feature foundation for subsequent encoder training and example retrieval, while improving the interpretability of features and facilitating targeted processing of features of different dimensions.
[0047] In some embodiments, the selected device may employ linear mapping or neural network mapping methods to map the extracted task features into the factorial latent state space; through a decoupling algorithm, the mapped feature vectors are decomposed into mutually independent latent factor vectors in the factorial latent variable set, ensuring that each latent factor vector corresponds to only one feature dimension and does not contain interference information from other dimensions; based on all the latent factor vectors obtained from the decomposition, they are combined through a common factor generation mechanism to reconstruct the observations corresponding to the original input text and output labels, thus verifying the effectiveness of decoupling and feature preservation.
[0048] Step S30: Construct a multi-head variational encoder and train the multi-head variational encoder using an objective function with a fully correlated decoupling penalty term. The multi-head variational encoder is configured to receive the concatenation embedding of input text and output labels and predict the distribution parameters of different latent factor vectors through multiple output heads.
[0049] It should be noted that a multi-head variational encoder can be a neural network model based on a variational autoencoder. It contains multiple output heads, each of which corresponds to a latent factor vector. It is used to receive the concatenation and embedding of the input text and the output label. By predicting the distribution parameters of different latent factor vectors through different output heads, it can achieve accurate modeling of multi-dimensional latent factors.
[0050] The objective function of the fully correlated decoupling penalty term can be a loss function used to constrain the training direction of the model during the training of a multi-head variational encoder. Its core function is to force each latent factor vector to remain independent by penalizing the correlation between latent factor vectors, thereby further enhancing the feature decoupling effect and reducing the training error of the model.
[0051] Distribution parameters are parameters (such as mean, variance, etc.) used to describe the distribution characteristics of latent factor vectors. They are predicted by the output head of a multi-head variational encoder and can accurately characterize the value patterns and distribution characteristics of each latent factor vector, providing a basis for subsequent posterior distribution analysis.
[0052] Understandably, this embodiment achieves simultaneous prediction of the distribution parameters of multiple independent latent factor vectors through a multi-head structure design, improving the efficiency and accuracy of latent factor modeling. The introduction of a fully correlated decoupling penalty term further strengthens the independence of latent factor vectors, avoids mutual interference between features of different dimensions, and improves the effect of feature decoupling. The trained multi-head variational encoder can accurately capture the distribution pattern of latent factor vectors, providing reliable model support for the mapping of subsequent test inputs and posterior distribution analysis, and reducing errors in subsequent retrieval stages.
[0053] In some embodiments, a device is selected to construct a multi-head variational encoder comprising an input layer, a hidden layer, and multiple output heads. The input layer receives the concatenated embedding of input text and output labels. The hidden layer performs feature extraction and transformation on the input through operations such as convolution and fully connected layers. Each output head corresponds to a latent factor vector. An objective function including a fully correlated decoupling penalty term is defined, integrating decoupling constraints, reconstruction error, KL divergence, and other indicators to measure the model training effect. The gradient descent algorithm is used to iteratively train the multi-head variational encoder using the reconstructed observations and their corresponding latent factor vectors as training data until the model converges, ensuring that the distribution parameters predicted by each output head can accurately match the true distribution of the latent factor vector.
[0054] Step S40: Construct a sample size calculation model based on the coverage number theory, and configure the total retrieval budget parameters through the sample size calculation model.
[0055] It should be noted that the covering number theory is a mathematical theory used to measure the complexity of a set. Its core is to quantify the distribution characteristics and complexity of a set by calculating the minimum number of covering units that can cover the target set. In this embodiment, it is used to quantify the distribution complexity of latent factors in the factorial latent state space.
[0056] The sample size calculation model is a mathematical model based on the coverage number theory. It is used to calculate the minimum sample size required to achieve effective retrieval based on the complexity of the factorial latent state space, the distribution characteristics of latent factors, and the retrieval accuracy requirements, thus providing a basis for configuring the total retrieval budget.
[0057] The total retrieval budget parameter is a core parameter used to constrain the retrieval process. Essentially, it is the maximum number of candidate examples that can be called during the retrieval process. It is calculated by the sample size calculation model and is used to balance retrieval accuracy and retrieval efficiency, avoiding waste or insufficiency of retrieval resources.
[0058] In some embodiments, the selected device is based on the coverage number theory, and a sample size calculation model is constructed by combining the dimension of the factorial latent state space, the distribution range and density of the latent factor vectors. Parameters such as the distribution complexity of latent factors, the accuracy of the retrieval target, and the retrieval efficiency requirements are input into the model, and the minimum sample size required to effectively cover all latent factor dimensions is obtained through model calculation. Based on this minimum sample size, and combined with the resource constraints of the actual retrieval scenario, the total retrieval budget parameter is reasonably configured to ensure that the budget parameter can meet the retrieval accuracy requirements without causing redundant consumption of retrieval resources.
[0059] Furthermore, in order to achieve accurate planning of the number of examples, step S40 above may include: Step S401: Construct a sample complexity function based on the coverage number theory. The sample complexity function is used to calculate the minimum number of examples required to control the inference error within a preset threshold under a preset confidence level. Step S402: Import the parameter space features of the factorial latent state space into the sample complexity function to construct a sample size calculation model; Step S403: Calculate the sample requirement under the single-unit characterization framework using the sample size calculation model, and use it as a benchmark reference value; Step S404: Calculate the lower bound of examples under the factorial representation framework using the sample size calculation model; Step S405: Establish a sample budget allocation inequality based on the benchmark reference value and the sample lower limit quantity, and configure the total retrieval budget parameters based on the sample budget allocation inequality.
[0060] Understandably, to achieve accurate planning of the number of examples, a sample size calculation model based on coverage theory is first established. In context learning, the inference performance of large language models highly depends on the encoder's understanding of latent task concepts. The estimation accuracy. To achieve the encoder's inference estimate... With the concept of real tasks Distribution error between Controlled within the allowable threshold To achieve the target within a certain range, a sufficient number of observation samples must be provided. Define the sample complexity function. Used to calculate confidence level The minimum number of examples required. This function is proportional to the metric entropy of the latent parameter space, and its formal definition is as follows: in, Indicates the reliability of the preset settings To control the distribution error to not exceed a preset threshold Minimum required sample size Indicates the distribution error threshold. The complement of the confidence level, Indicates a constant factor. In factorial latent state space , with norm For measurement, radius is The coverage number reflects the factorial latent state space. The geometric complexity, Describing the latent state space of factorial The measure of entropy.
[0061] In this computational model Entropy, representing the metric required to cover the latent space, is used to evaluate the mapping between the complexity of the current task representation and the required number of examples. If the geometric complexity of the latent space is too high, leading to computationally inefficient representations... If the context window limit of a large language model is exceeded, dimensionality reduction or decoupling strategies must be adopted to reduce sample requirements.
[0062] Using the above computational model, the sample requirement under the traditional single-unit representation assumption is calculated as a benchmark. Under the single-unit assumption, all feature dimensions of the task are considered as an indivisible whole, compressed to a total dimension of [value missing]. In the high-dimensional vector. At this point, the encoder is forced to operate in the entire joint Cartesian product space. The search will proceed within this scope. Sample size required. The following exponential dependency relationship is presented: To address the sample size overflow issue caused by the singleton assumption, the budget configuration module applies the factorial representation framework established in step one to perform the final sample budget configuration. Because the parameter space is decoupled within the factorial framework... For each independent submanifold, according to the additivity principle of statistical learning theory, the metric entropy of the joint space is equal to the sum of the metric entropies of each subspace. Therefore, the total sample size required for recalculation is... This is transformed into a linear superposition of the sample sizes required for each sub-dimension. The logical derivation of this calculation is as follows: Based on this calculation result, the budget allocation module establishes the following sample budget allocation inequality as the final determination of the retrieval quantity. Basis: in, This indicates the retrieval of the total budget parameter. Indicates the first The minimum number of representative examples required for each feature dimension This represents the sample requirement under the factorial representation framework. This represents the sample requirement under the single-unit representation framework. This inequality indicates that the budget allocation module only needs to consider each individual feature dimension. Retrieve a small number of representative examples. This allows for the comprehensive coverage of the entire task space through combination. Based on this, the total retrieval budget... The configuration is the sum of the minimum requirements for each dimension, rather than their product. This ensures that within a very limited context window, effective cue words that satisfy the convergence condition of generalization error are constructed by combining examples from different dimensions, achieving the engineering goal of approximating the optimal prediction performance of Bayes with minimal sample cost.
[0063] Step S50: Configure the retrieval module based on the total retrieval budget parameters and the trained multi-head variational encoder.
[0064] It should be noted that the retrieval module can be built based on the total retrieval budget parameters and the trained multi-head variational encoder. It is used to receive test input text, map it to the factorial latent state space, and parse the posterior distribution parameter set of the test input in each latent factor dimension, providing the core basis for subsequent example retrieval.
[0065] It should be noted that the test input text refers to the input text (without output labels) that the large language model needs to match when it actually performs context learning. It is the target object in the retrieval process and needs to find a matching context example through the retrieval module.
[0066] The posterior distribution parameter set can be a set of distribution parameters of the test input in each latent factor dimension in the factorial latent state space. Each distribution parameter corresponds to a latent factor dimension, which can accurately characterize the feature distribution of the test input in that dimension and is the basis for determining the target point.
[0067] In some embodiments, the selected device sets parameters such as the retrieval range and retrieval accuracy threshold of the retrieval module based on the total retrieval budget parameter; the trained multi-head variational encoder is integrated into the retrieval module as a core component for test input mapping and distribution parameter prediction; the test input is segmented and encoded into an embedding vector consistent with the training data format, and input into the multi-head variational encoder in the retrieval module; through the multiple output heads of the encoder, the posterior distribution parameters of the test input in each latent factor dimension are predicted respectively, and the posterior distribution parameters of all dimensions are integrated to form a posterior distribution parameter set, thus completing the feature parsing of the test input.
[0068] Furthermore, in order to achieve accurate example matching, step S50 above may include: Step S501: Import the total retrieval budget parameter into the retrieval module to configure the sample retrieval quantity constraint of the retrieval module; Step S502: Based on the trained multi-head variational encoder, a parameter extraction function submodule is built in the retrieval module. The parameter extraction function submodule is used to extract the posterior distribution parameter set of the test input text in each latent factor dimension.
[0069] Understandably, to achieve accurate example matching, the retrieval module first needs to analyze the specific requirements of the current test input in each latent feature dimension. Using the multi-head variational encoder trained in step two, the retrieval module, in inference mode, will process the test input... Mapped to the factorial latent state space.
[0070] For a given test input The multi-head variational encoder outputs its... The set of posterior distribution parameters over *n* independent latent factors. This set quantifies the specific features and uncertainties of the current task in terms of format, semantics, and logic, and its mathematical expression is: in, This indicates that the test input text is at the [number]th [position]. The posterior distribution over the latent factor vectors This represents the test input text. This indicates that the test input text is at the [number]th [position]. Central demand on each potential factor vector This indicates that the test input text is at the [number]th [position]. The variance vector over the latent factor vectors This indicates that the multi-head variational encoder is used for the test input text. The output.
[0071] In the above formula, Characterizes the test input at the 1st The central need in each characteristic dimension (e.g., "the need for medical knowledge"), and This characterizes the uncertainty or tolerance of the demand. This set of distribution parameters constitutes the target point for subsequent retrieval operations.
[0072] Step S60: In response to receiving test input text, input the test input text into the retrieval module to obtain a posterior distribution parameter set, determine the target point based on the posterior distribution parameter set, and retrieve a complementary example set that combines all target points from the candidate example pool through a centroid approximation strategy.
[0073] It should be noted that the target point can be a feature target point determined based on the set of posterior distribution parameters of the test input, which needs to be covered by contextual examples. Each target point corresponds to a core feature of a latent factor dimension, which is the core basis for retrieving examples and ensures that the retrieved examples can match the feature requirements of the test input.
[0074] The centroid approximation strategy is an optimization strategy for example retrieval. Its core is to calculate the feature centroid of each example in the factorial latent state space in the candidate example pool. By approximating the centroid corresponding to the posterior distribution parameter set of the test input, the example that best matches the features of the test input is selected, while ensuring the complementarity between examples.
[0075] It should be noted that the candidate example pool stores a set of all context examples that can be used for retrieval. Each example has been processed through the above steps and transformed into a latent factor vector in the factorial latent state space, which facilitates the retrieval module to perform fast matching and filtering.
[0076] A complementary example set refers to a set of examples retrieved from the candidate example pool that can be combined to cover all target points. The examples in the set are complementary, and each example corresponds to one or more target points. When combined, they can completely cover all feature dimensions of the test input, providing comprehensive example support for contextual learning of large language models.
[0077] It is understandable that this embodiment achieves accurate screening of candidate examples through a centroid approximation strategy, thereby improving the matching degree between examples and test inputs; it ensures that the retrieved complementary example set can completely cover all target points of the test input, avoiding poor context learning results caused by incomplete example coverage; the complementarity between examples can reduce redundant examples, improve the efficiency of context learning, and at the same time provide comprehensive and accurate example support for large language models, helping the model to better understand the features of the test input and improve the performance of context learning.
[0078] In some embodiments, the device is selected based on the posterior distribution parameter set, extracts the core features of each latent factor dimension, determines the target points to be covered, and clarifies the feature requirements corresponding to each target point; a centroid approximation strategy is adopted to calculate the feature centroid of each example in the factorial latent state space in the candidate example pool, and at the same time calculates the matching degree between each example and each target point; combined with the total retrieval budget parameter, examples with high matching degree and mutual complementarity are selected from the candidate example pool to ensure that the selected example set can be combined to cover all target points, and finally form a complementary example set, which is output to the large language model for context learning.
[0079] Furthermore, in order to effectively measure the difference between candidate examples and target requirements and improve the accuracy of example selection, step S60 above may include: Step S601: Extract the mean and variance parameters of each latent factor dimension from the posterior distribution parameter set and integrate them to form a feature parameter set; Step S602: Use the set of feature parameters as the feature identifier of the test input text in the factorial latent state space, and determine it as the target point; Step S603: Define a weighted second-order Wasserstein distance metric function, which is used to calculate the distribution difference between the test input text and the candidate examples in the candidate example pool; Step S604: Construct an optimization function based on the first-order moment centroid approximation strategy and the second-order Wasserstein distance metric function; Step S605: Based on the optimization function, select candidate examples from the candidate example pool that can correct the set bias, and combine the selected candidate examples to obtain a complementary example set covering all target points.
[0080] Understandably, after establishing the target, the retrieval module needs to define a metric that can effectively measure the difference between candidate examples and the target requirement. Since latent factors are modeled as probability distributions rather than deterministic point vectors, traditional Euclidean distance cannot capture the differences brought about by the shape of the distribution (i.e., variance information). Therefore, the retrieval module introduces the Wasserstein-2 distance from optimal transport theory as the core metric function to calculate the geometric distance between two Gaussian distributions on a Riemannian manifold.
[0081] Considering that different tasks have varying sensitivities to different feature dimensions, the system defines a weighted distribution distance metric function. For the test input... With any candidate example Their combined difference in the latent space is defined as the weighted sum of the Wasserstein distances of each subspace: in, Indicates the test input text With candidate examples A comprehensive measure of difference in the factorial latent state space. Indicates the first The weight coefficients of each potential factor dimension. Denotes the squared second-order Wasserstein distance between two Gaussian distributions. This indicates that the test input text is at the [number]th [position]. The posterior distribution over the latent factor vectors Indicates the candidate example in the th The posterior distribution over the potential factor vectors; for the diagonal covariance matrix, its analytical solution is... Weighting coefficients The retrieval module automatically adjusts the weighting based on the posterior variance of each dimension (the smaller the variance, the higher the confidence level, and the greater the weight). This metric function serves as a theoretical guide, ensuring that the system accurately quantifies the degree of matching between examples and task requirements at the distribution level.
[0082] Based on the aforementioned metrics, the retrieval module performs the final example selection operation. Unlike existing techniques that search for the "Top-k" examples most similar to the test input, the goal of this system is to construct a complementary set of examples. The core logic is that individual examples often exhibit a "specialization" phenomenon, while an excellent combination of examples should complement each other so that the mixed feature distribution of the combination can cover the feature requirements of the test input.
[0083] To achieve this goal, a centroid approximation strategy based on first-order moments is adopted as an engineering simplification of the Wasserstein distance optimization. The retrieval module no longer minimizes the distance between a single example and the target, but rather minimizes the deviation between the aggregate mean center of the example set and the target requirement. Mathematically, this is formalized as a combinatorial optimization problem with cardinality constraints: in, This represents the final set of complementary examples. Represents the candidate example pool, Indicates the first One candidate example, Indicates the size of the candidate example pool. This indicates that the test input text is at the [number]th [position]. The mean of the potential factor vectors Indicates the first The candidate example in the first The mean of the potential factor vectors; The physical meaning of this optimization objective is extremely profound: it requires the selection of... The "centroid" formed by each example in the vector space must fall as close as possible to the target feature point of the test input. Although the formula is formally simplified to mean matching, the aforementioned weights... The variance adjustment mechanism essentially retains the constraint effect of the second moment on optimization. During execution, an iterative greedy algorithm is used, selecting an example that can maximally correct the current set's bias at each step. For example, if the current set is well-matched in terms of "format factor" but biased in terms of "logic factor," the algorithm will tend to select an example that is extremely strong in "logic factor," even if it is weak in "format factor," to balance the overall distribution. This mechanism achieves true "complementary selection," ensuring that the final generated contextual prompt provides accurate guidance for the large language model in all key dimensions, including format, semantics, and logic, thereby maximizing inference performance.
[0084] This embodiment constructs a factorial latent state space to perform refined modeling of task features, decoupling the abstract task concept into mutually independent latent factors. Through common-cause causal modeling, it abandons the binary causal assumption, treating the input text and output labels as observations jointly generated by latent factors. During the training phase, a fully correlated decoupling penalty term is introduced to optimize the variational lower bound, training a multi-head encoder to identify and separate each independent factor. During the inference phase, by calculating the demand distribution of the test input text in each factor space, a complementary example set covering all target factors is retrieved and combined from the candidate pool. This embodiment effectively mines the multi-dimensional features of contextual examples, accurately capturing examples through feature extraction, decoupling, and modeling. The core feature attributes are improved to enhance the accuracy and interpretability of feature representation; the efficiency and accuracy of context example selection are improved by rationally allocating retrieval resources through sample size calculation models and combining centroid approximation strategies to achieve accurate example selection, avoiding redundancy and omissions, and reducing retrieval costs; the relevance and comprehensiveness of example selection are strengthened by constructing complementary example sets to ensure that examples can fully cover all feature targets of the test input, providing high-quality example support for context learning of large language models; thus, the problem of biased example selection in complex context learning tasks is effectively solved, significantly improving the few-shot inference performance of large language models, achieving accurate example selection, and improving the matching degree between examples and test input text.
[0085] In some embodiments, this invention can be applied to a contextual learning scenario for sentiment classification of Chinese e-commerce reviews using a large language model. The core requirement of this scenario is that the large language model needs to quickly learn to classify the sentiment of product reviews on Chinese e-commerce platforms (distinguishing between positive, negative, and neutral sentiments) using a small number of contextual examples, without requiring large-scale retraining of the model. This is suitable for real-time sentiment analysis of e-commerce reviews, rapid categorization of user feedback, and monitoring of product reputation, helping e-commerce platforms efficiently process massive amounts of review data and improve user experience and operational efficiency.
[0086] The learning task in this scenario is Chinese e-commerce review sentiment classification, meaning that the large language model learns from contextual examples to master the ability to "input Chinese e-commerce review text and output corresponding sentiment labels (positive / negative / neutral)". The original contextual learning task data can include: input text is real product reviews from e-commerce platforms (e.g., "These headphones have clear sound quality and sufficient battery life, I'm very satisfied"), and the corresponding output label is the sentiment category of the review (e.g., "positive"). All such combinations of "review text + sentiment label" constitute the original contextual learning task data used for feature extraction and model training.
[0087] In this scenario, the task features extracted from the original context learning task data, and the decoupled latent factor vectors (single-dimensional features), specifically include four categories, each corresponding one-to-one with the factorial latent variable set: (1) Semantic sentiment features: The words and semantic tendencies that directly express emotions in the comments, such as “satisfied”, “not easy to use”, and “so-so”, correspond to a latent factor vector, which is used to capture the core sentiment tendency of the comments; (2) Characteristics of the review object: The product function / attribute targeted by the review, such as "sound quality", "battery life", "logistics" and "price", corresponds to a latent factor vector, which is used to distinguish the focus of the review; (3) Tone intensity features: The tone and intensity of the comment, such as "very good", "extremely bad" and "okay", correspond to a latent factor vector to capture the intensity of emotional expression; (4) Sentence structure characteristics: The sentence type of the comment (declarative sentence, exclamatory sentence, rhetorical question), such as "It's such a good deal!" (exclamatory sentence) and "Is this quality acceptable?" (rhetorical question), corresponds to a potential factor vector, which is used to help judge the sentiment tendency.
[0088] The above four types of features are the task features extracted from the original context learning task data. After decoupling, they form four independent latent factor vectors, which together constitute the core content of the factorial latent state space.
[0089] In this scenario, the example is a specific combination of "e-commerce review text + sentiment tags", as shown in the following example: (1) Some examples in the candidate example pool (covering different potential factor dimensions): Example 1: Input text "This face cream is very moisturizing, and my skin is very smooth after using it. I recommend buying it!", output label "positive" (core features: positive semantic sentiment, the object of the comment is the texture of the face cream, strong tone, and exclamatory sentence). Example 2: Input text "The logistics were too slow and the goods were damaged. I am very dissatisfied". Output label "negative" (core features: negative semantic sentiment, the object of the comment is logistics + the condition of the goods, strong tone, and the sentence structure is a declarative sentence). Example 3: Input text "The sound quality of the headphones is okay, the battery life is average, neither amazing nor bad", output label "neutral" (core features: semantic sentiment neutral, the comment object is sound quality + battery life, weak tone intensity, sentence structure is declarative sentence); Example 4: Input text "This price is too good to pass up for this quality, right?", output label "positive" (core features: positive semantic sentiment, comment object is price + quality, strong tone, sentence structure is rhetorical question).
[0090] (2) Complementary example set: Assuming the test input is "This phone's photo quality is average, but the battery life is okay" (labeled "neutral"), its target points are "semantic sentiment neutral, the comment object is photo + battery, weak tone intensity, and the sentence structure is declarative sentence". The complementary example set obtained by the centroid approximation strategy can be composed of Example 3 (covering neutral sentiment, sound quality + battery life is similar to photo + battery, weak tone, and declarative sentence) and Example 1 (assisting in covering the sentence structure norms of product attribute comments), ensuring complete coverage of all target points of the test input and providing accurate contextual reference for the model.
[0091] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the context example selection method based on factorial latent variables and common factors generated by the present invention.
[0092] Based on the first embodiment described above, in this embodiment, step S20 further includes: Step S201: Map the task features to the factorial latent state space, and perform dimensional decoupling processing on the task features in the factorial latent state space to obtain a subset of task features.
[0093] It should be noted that a task feature subset refers to a set of multiple single-dimensional features obtained after decoupling the original task features by dimension. Each subset corresponds to a specific feature dimension (such as the semantic sentiment dimension or the comment object dimension of e-commerce reviews). The subset only contains feature information of that dimension and does not contain interference from other dimensions.
[0094] It is understandable that this embodiment realizes the transformation of task features from "multi-dimensional fusion" to "single-dimensional splitting". By decoupling dimensions, interference between different feature dimensions is eliminated, and the purity of features is improved. The resulting subset of task features avoids potential factor modeling errors caused by the mixing of multi-dimensional features.
[0095] In some embodiments, taking the Chinese e-commerce review sentiment classification scenario as an example, the extracted task features (including four types of fused features: semantic sentiment, review object, tone intensity, and sentence structure) are mapped to the constructed factorial latent state space through a linear mapping algorithm (such as linear regression mapping). This ensures that the mapped feature vector matches the dimension of the latent state space and retains the core information of the original task features. The task features mapped to the latent state space are then decoupled according to dimensions. Based on the four preset feature dimensions (corresponding to four latent variables), the fused features are split into four independent task feature subsets: semantic sentiment feature subset (containing only sentiment words and semantic tendencies in the review), review object feature subset (containing only the product attributes targeted by the review), tone intensity feature subset (containing only the tone intensity information of the review), and sentence structure feature subset (containing only the sentence type information of the review).
[0096] Step S202: Map each task feature subset to the corresponding vector of the factorial latent variable set to obtain the initial latent factor vector.
[0097] It should be noted that the initial latent factor vector can be the initial feature vector obtained by mapping each task feature subset to the corresponding latent variables in the factorial latent variable set. It has not undergone distribution optimization and only initially corresponds to the feature representation of each latent variable. It needs to be further optimized before it can be used as the final latent factor vector.
[0098] In some embodiments, multiple task feature subsets are mapped one-to-one with latent variables in the factorial latent variable set. For example, the semantic sentiment feature subset corresponds to the semantic sentiment latent variable, and the comment object feature subset corresponds to the comment object latent variable. A neural network mapping method is used to map each task feature subset to the vector space of the corresponding latent variable, transforming it into an initial latent factor vector of fixed dimension. For example, the semantic sentiment feature subset (containing lexical features such as "satisfied" and "not easy to use") is mapped to a semantic sentiment latent factor vector, and the vector value corresponds to the strength of sentiment tendency.
[0099] Step S203: Define independent standard normal marginal distributions for the initial latent factor vectors, and use the product of each marginal distribution as the joint prior distribution of the factorial latent variable set.
[0100] It should be noted that the standard normal marginal distribution refers to a normal distribution with a mean of 0 and a variance of 1. In this embodiment, each initial latent factor vector is defined separately to constrain the distribution pattern of each latent factor vector and ensure the independence of each latent factor vector.
[0101] The joint prior distribution refers to the distribution obtained by multiplying all the marginal distributions based on the standard normal marginal distributions of each initial latent factor vector. It is used to describe the overall distribution law of the factorial latent variable set and to provide a constraint basis for the optimization of the initial latent factor vectors.
[0102] It is understood that this embodiment achieves precise optimization of the initial latent factor vector by constraining the standard normal marginal distribution and the joint prior distribution, ensuring that the optimized latent factor vectors are independent of each other and completely eliminating the interference of features of different dimensions; the optimized latent factor vectors have a regular distribution pattern, which improves the stability and accuracy of feature representation.
[0103] Step S204: Optimize the initial latent factor vector based on the joint prior distribution to obtain mutually independent latent factor vectors in the factorial latent state space.
[0104] In some embodiments, the device selection may employ the KL divergence minimization algorithm, which uses the joint prior distribution as a constraint to iteratively optimize the initial latent factor vector, adjust the distribution of the vector values so that the distribution of each latent factor vector is infinitely close to its standard normal marginal distribution, and there is no correlation between the vectors.
[0105] Step S205: Establish a generative topology structure with the set of factorial latent variables as global conditions, and concatenate the input text and output labels into a joint observation sequence.
[0106] It should be noted that the generative topology refers to the feature generation network structure built with the factorial latent variable set as the global condition. It is used to receive latent factor vectors and, according to preset generation rules, transform the features of the latent factor vectors into a representation form corresponding to the original input and output data.
[0107] The joint observation sequence can be a unified sequence formed by concatenating the input text (e.g., e-commerce review text) and the corresponding output label (e.g., sentiment label) in the original context learning task data. It contains complete information about the input and output and serves as an input reference for generating the topology.
[0108] It is understood that the generated topology structure constructed in this embodiment provides a reliable network carrier for the transformation of latent factor vectors into observations, ensuring that the features of latent factors can be effectively transformed into equivalent representations of the original input and output; the splicing of joint observation sequences integrates the complete information of input text and output labels, providing an accurate reference standard for observation reconstruction and avoiding the disconnect between the reconstructed observations and the original data.
[0109] Step S206: Reconstruct the input text and output labels into observations generated by the combined effect of the latent factors through the generated topology structure.
[0110] In the specific implementation, in order to ensure that different latent factor vectors encode different attributes of the task and prevent feature redundancy, orthogonality constraints are introduced during the construction of the feature space so that the directions of latent factor vectors of different dimensions in the feature space are kept as orthogonal as possible.
[0111] The joint prior distribution is defined as the distribution of each independent latent factor. Product of marginal distributions: in, Describes the set of latent variables for factorial. The joint prior distribution, This represents the number of latent factor vectors. Indicates the first A vector of potential factors, Indicates the first Marginal prior distributions of potential factor vectors This represents a vector with zero mean and a covariance matrix that is the identity matrix. The normal distribution; By forcing prior distributions to be independent, large language models are compelled to decompose and map the complex correlations in the observed data to mutually independent latent factors during the learning process, thereby achieving fine-grained control over task features.
[0112] To address the characteristic of large language models processing long sequences through attention mechanisms, a method based on factorial latent variables is established. This is the topology for generating global conditions. In this structure, the input text sequence within the context... and target label sequence No longer viewed as an independent entity with a unidirectional causal relationship, but rather as a entity with underlying task intent. The observed symbol sequence is generated under the combined effect of these factors. Therefore, the generation process of the observed data is formalized as a joint probability distribution, obtained by multiplying the prior probability of the factorial latent variable by the conditional generation probability given the factorial latent variable: in, The representative is the parameter. The likelihood function is defined by the pre-trained large language model. To adapt to the word-by-word autoregressive generation mechanism of the large language model, the joint observation sequence is denoted as... Furthermore, the generation probabilities of the above conditions are expanded into a product form based on time steps. The generation probability of each word not only depends on historical words but is also explicitly affected by the set of latent factors. Conditions and constraints: In this formula, Represents the first in the sequence Each word element, Represents the historical context prior to the current moment. Latent factor vector. Through feature fusion function Integrated into the input layer of a large language model, the large language model uses an attention mechanism to adjust the output probability distribution of each word: Contextual learning is essentially a reverse inference process, which uses the observed sequence of input and output examples to calculate the factorial latent factor that maximizes the generation probability. This, in turn, guides the reasoning for new tasks.
[0113] This embodiment eliminates interference between different feature dimensions through dimensional decoupling, improves feature purity, and achieves accurate association between task feature subsets and factorial latent variables. It transforms single-dimensional feature subsets into initial latent factor vectors that meet the requirements of the latent state space. Through constraints of standard normal marginal distribution and joint prior distribution, it achieves accurate optimization of the initial latent factor vectors, ensuring that the optimized latent factor vectors are independent of each other. By generating topological structures and common factor generation mechanisms, it achieves accurate reconstruction of latent factor vectors into observations, ensuring that the observations are equivalent to the original input and output data, completely preserving the original task features, and avoiding feature loss due to feature decoupling or optimization. This provides high-quality training data for the subsequent training of multi-head variational encoders.
[0114] refer to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the context example selection method based on factorial latent variables and common factors generated by the present invention.
[0115] Based on the above embodiments, in this embodiment, step S30 further includes: Step S301: Introduce an approximate posterior distribution controlled by preset parameters, and decompose the approximate posterior distribution into independent sub-distribution structures corresponding to each latent factor vector.
[0116] It should be noted that the approximate posterior distribution refers to the probability distribution used to approximate the true posterior distribution of the latent factor vector. It is controlled by preset parameters (such as initial mean and variance). Its function is to reduce the inference complexity of the posterior distribution, facilitate the prediction of subsequent distribution parameters and model training, and is a core component of variational inference.
[0117] Preset parameters refer to the parameters used to control the initial state of the approximate posterior distribution, including the initial mean and initial variance of the distribution. They can be preset according to the distribution law of features in the actual learning scenario (such as the sentiment classification scenario of e-commerce reviews), and can be further optimized through model training.
[0118] The independent sub-distribution structure can be achieved by splitting the approximate posterior distribution into multiple independent sub-distributions according to the dimension of the latent factor vector. Each sub-distribution corresponds to a latent factor vector, ensuring that there is no correlation between the sub-distributions, which matches the independence of the factorial latent variable set.
[0119] Understandably, given the extremely large parameter space of a large language model, directly calculating a given input sequence... With target label Lower factorial latent variables The true posterior distribution involves integral operations over a high-dimensional latent space, which is too complex. Therefore, this embodiment designs a variational inference framework. It approximates the true posterior distribution by constructing a parameterized multi-head variational encoder, and further constructs an optimization mechanism that enforces orthogonal decoupling of features by combining the full correlation constraint from information theory.
[0120] To transform uncomputable posterior inference into an optimizable function approximation problem, a parameter is introduced. Approximate posterior distribution of control A multi-head variational encoder was constructed as the recognition model. This encoder was designed to support flexible input modes, capable of processing complete observation data. Alternatively, a masking mechanism can be used to process some of the observation data. During the training phase, the encoder receives... and splicing and embedding and through Each output head predicts the distribution parameters of different latent factors to prevent information mixing between different feature channels and to divert the observed information to a predefined orthogonal subspace. Formally, the approximate posterior distribution is decomposed into... The product of 10 independent sub-distributions is given by the following formula: in, Indicates that it is based on preset parameters The approximate posterior distribution of the control. This indicates the input text. Indicates preset parameters. Indicates the output label. Represents the set of latent variables for factorial. Indicates the first The mean vector of the potential factor vectors Indicates the first The variance vector of the latent factor vectors Represents the diagonal covariance matrix; and The encoder's first The calculation is derived from the individual output heads, each representing the... The central trend and range of uncertainty of each potential factor.
[0121] Step S302: Based on the independent sub-distribution structure, design a multi-output head architecture for the multi-head variational encoder and construct the multi-head variational encoder.
[0122] It should be noted that the multi-output head architecture is the core architecture of the multi-head variational encoder. It consists of multiple independent output heads, each of which corresponds to an independent sub-distribution (i.e., a latent factor vector). Its core function is to predict the distribution parameters of the corresponding latent factor vectors, thereby achieving synchronous modeling of multi-dimensional latent factors.
[0123] In some embodiments, the selected device constructs a complete multi-head variational encoder, comprising a three-layer core structure: Input layer: receives the concatenated embedding of e-commerce review text and sentiment tags (the dimension is consistent with the extracted task feature dimension); Hidden layer: consists of two fully connected neural networks, using the ReLU activation function, used to extract and transform features from the concatenated embedding, and outputting an intermediate feature vector of fixed dimension; Output layer: namely, four independent output heads, each using a linear activation function, used to output the distribution parameters (mean and variance) of the corresponding latent factor vector.
[0124] Step S303: Construct an evidence lower bound optimization objective based on the Bayesian learning principle.
[0125] It should be noted that the Bayesian learning principle infers the posterior distribution from the prior distribution and observed data. In this embodiment, it is used to construct the optimization objective of the multi-head variational encoder to ensure that the model training can approximate the true distribution of the latent factor vector.
[0126] The Evidence Lower Bound Optimization Objective (ELBO) is an optimization objective built on the principles of Bayesian learning. It is used to approximate the logarithmic evidence of the posterior distribution. Its core function is to reduce the training error of the model and ensure that the model can accurately capture the distribution pattern of the latent factor vector.
[0127] Understandably, to simultaneously optimize generation quality and inference accuracy, based on Bayesian learning principles, the ultimate goal of multi-head variational encoder optimization is to maximize the log-marginal likelihood of the observed data. However, directly optimizing this likelihood function is extremely difficult. Using Jensen's inequality, the objective can be transformed into maximizing its lower bound of evidence, and a balance mechanism between reconstruction error and regularization term can be designed. Specifically, the mechanism for maximizing its lower bound of evidence consists of two parts: the first part is the reconstruction expectation, which forces the inferred latent factors to accurately reconstruct the observed context examples; the second part is the KL divergence, which constrains the approximate posterior distribution to not deviate from the prior distribution, preventing the model from overfitting to individual samples. Its mathematical expression is as follows: In actual optimization, since the expectation term involves factorial latent variables... The sampling operation can lead to gradient breakage, making it impossible to update the encoder parameters using the backpropagation algorithm. To address this, a reparameterization technique is introduced, transferring randomness from the sampling process to an independent auxiliary noise variable. Above, immediately ordered ,in This allows the sampling operation to be differentiable with respect to the mean and variance of the encoder output, enabling end-to-end joint training from the generation loss function of the large language model back to the gradient flow of the multi-head variational encoder.
[0128] in, This indicates the objective term for optimizing the lower bound of the evidence. This represents the expectation operation over an approximate posterior distribution. Representing a given latent variable Time observation data The conditional log-likelihood, The parameters representing the large language model. KL divergence is used to measure the difference between two probability distributions. Indicates the first Approximate posterior distribution of the potential factor vectors, Indicates the first Sampled values of a latent factor vector, Indicates the first The standard deviation of the latent factor vector Let represent the auxiliary noise variable, which follows a standard normal distribution.
[0129] Step S304: Add a fully correlated decoupling penalty term to the lower bound of evidence optimization objective to form the objective function.
[0130] Understandably, while maximizing the lower bound of evidence based on the mean-field hypothesis can force factor independence when dealing with individual samples, correlations between factors may still occur when aggregating the posterior distribution across the entire dataset, resulting in features not achieving true decoupling at the global level. For example, the model might repeatedly encode the same text length information in both the "format factor" and "semantic factor," and this information redundancy weakens the model's generalization ability when faced with new task combinations. To address this, a higher-order penalty term based on information theory—full correlation—is introduced into the loss function. This term measures a strict metric of the dependency between multiple random variables, and its value is zero if and only if these variables are statistically completely independent. By minimizing the full correlation between latent factors, the multi-head variational encoder is forced to strip away common information from the observed data, retaining only the attributes specific to each dimension. The modified objective function is formalized as follows: in, This represents the total loss term of the objective function. This represents a hyperparameter that controls the strength of decoupling. This represents the fully correlated decoupling penalty term. Denotes the aggregate posterior distribution of the set of latent variables for factorial. Indicates the first The marginal posterior distribution of the nth latent factor vector; if the nth... The first factor has already encoded the format feature of "question-answer pairs", which will prevent the second factor from being used. Factors ( It re-encodes any format-related information. This fundamentally guarantees the purity and orthogonality of the latent representation, enabling the multi-head variational encoder to learn truly independent task primitives and perform zero-shot reasoning on unknown task combinations.
[0131] Step S305: Train the multi-head variational encoder using the objective function so that the multi-head variational encoder can predict the distribution parameters of different latent factor vectors through multiple output heads.
[0132] Understandably, this embodiment optimizes the network parameters of the multi-head variational encoder through the constraint of the objective function and iterative training, thereby improving the model's fitting accuracy and ensuring that each output head can accurately predict the distribution parameters of the corresponding latent factor vectors. The verification process during training ensures the training effect of the model and avoids problems such as overfitting, underfitting, or poor decoupling. The trained multi-head variational encoder can accurately capture the distribution patterns of latent factors in various dimensions in the context learning scenario, providing reliable model support for the mapping of subsequent test input text and posterior distribution analysis, thus improving the accuracy and efficiency of subsequent retrieval processes.
[0133] In some embodiments, the selection device can select training data for the Chinese e-commerce review sentiment classification scenario, namely the reconstructed observations (vector form) and their corresponding optimized latent factor vectors, which are divided into a training set and a validation set in an 8:2 ratio. The training set is used for model training, and the validation set is used to verify the training effect. The gradient descent algorithm is used as the optimizer, with a learning rate preset to 0.001, a number of iterations preset to 1000, and a convergence threshold preset to 0.0001. Training stops when the loss value of the objective function is lower than the convergence threshold for 10 consecutive iterations. The training set data is input into the constructed multi-head variational encoder. The loss value of the objective function is calculated through the model. The gradient descent algorithm is used for backpropagation to adjust the network parameters of each layer of the model. After each iteration, the training effect of the model is verified using the validation set. The correlation between the loss value of the validation set and the latent factor vector is calculated to ensure that the model training direction is correct.
[0134] This embodiment reduces the complexity of posterior inference by introducing and decomposing an approximate posterior distribution, while simultaneously matching the independence requirements of factorial latent variables, providing a scientific basis for encoder architecture design. Through the construction of a multi-head variational encoder with a multi-output head architecture, it achieves synchronous prediction of the distribution parameters of multi-dimensional latent factor vectors, improving the efficiency and relevance of latent factor modeling and meeting the modeling needs of multi-dimensional features in e-commerce reviews. By training with an objective function that balances fitting accuracy and decoupling effect, it ensures that the trained multi-head variational encoder can accurately capture the distribution patterns of latent factors, strengthening feature decoupling and providing high-quality model support for subsequent test input parsing and example retrieval. This further improves the logical closed loop of the entire technical solution and enhances its feasibility and practicality.
[0135] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a context example selection program generated based on factorial latent variables and common factors. When the context example selection program generated based on factorial latent variables and common factors is executed by a processor, it implements the steps of the context example selection method generated based on factorial latent variables and common factors as described above.
[0136] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0137] The aforementioned computer-readable storage medium may be included in a context example selection device based on factorial latent variables and common causes; or it may exist independently and not assembled into a context example selection device based on factorial latent variables and common causes.
[0138] Furthermore, this invention also proposes a computer program product, including a context example selection program generated based on factorial latent variables and common factors. When the context example selection program generated based on factorial latent variables and common factors is executed by a processor, it implements the steps of the context example selection method generated based on factorial latent variables and common factors as described above.
[0139] The specific implementation of the computer program product of the present invention is basically the same as the various embodiments of the context example selection method based on factorial latent variables and common factors generated above, and will not be repeated here.
[0140] Reference Figure 5 , Figure 5 This is a structural block diagram of the first embodiment of the context example selection system based on factorial latent variables and common factors generated by the present invention.
[0141] like Figure 5 As shown, the context example selection system based on factorial latent variables and common factor generation proposed in this embodiment of the invention includes: The latent space construction module 10 is used to extract task features from the original context learning task data of the large language model and define a set of factorial latent variables in the continuous embedding space of the large language model to form a factorial latent state space. The original context learning task data includes the input text and the corresponding output label for context learning. The task reconstruction module 20 is used to map and decouple the task features into mutually independent latent factor vectors in the set of latent variables in the latent factor state space, so as to reconstruct the input text and output label into observations generated by the combined action of the latent factors. The latent factor vectors are used to capture feature attributes of different dimensions in the context example. The encoding configuration module 30 is used to construct a multi-head variational encoder. The multi-head variational encoder is trained using an objective function with a fully correlated decoupling penalty term. The multi-head variational encoder is configured to receive the concatenation and embedding of input text and output labels and predict the distribution parameters of different latent factor vectors through multiple output heads. Budget configuration module 40 is used to construct a sample size calculation model based on the coverage number theory, and to configure the total budget parameters for retrieval through the sample size calculation model. The retrieval configuration module 50 is used to configure the retrieval module based on the total retrieval budget parameter and the trained multi-head variational encoder. The retrieval module is used to map the test input text to the factorial latent state space and parse the test input text to obtain the posterior distribution parameter set of each latent factor dimension. Example retrieval module 60 is used to respond to receiving test input text, input the test input text into the retrieval module, obtain a posterior distribution parameter set, determine the target point based on the posterior distribution parameter set, and retrieve a complementary example set that combines and covers all target points from the candidate example pool through a centroid approximation strategy.
[0142] This embodiment constructs a factorial latent state space to perform refined modeling of task features, decoupling the abstract task concept into mutually independent latent factors. Through common-cause causal modeling, it abandons the binary causal assumption, treating the input text and output labels as observations jointly generated by latent factors. During the training phase, a fully correlated decoupling penalty term is introduced to optimize the variational lower bound, training a multi-head encoder to identify and separate each independent factor. During the inference phase, by calculating the demand distribution of the test input text in each factor space, a complementary example set covering all target factors is retrieved and combined from the candidate pool. This embodiment effectively mines the multi-dimensional features of contextual examples, accurately capturing examples through feature extraction, decoupling, and modeling. The core feature attributes are improved to enhance the accuracy and interpretability of feature representation; the efficiency and accuracy of context example selection are improved by rationally allocating retrieval resources through sample size calculation models and combining centroid approximation strategies to achieve accurate example selection, avoiding redundancy and omissions, and reducing retrieval costs; the relevance and comprehensiveness of example selection are strengthened by constructing complementary example sets to ensure that examples can fully cover all feature targets of the test input, providing high-quality example support for context learning of large language models; thus, the problem of biased example selection in complex context learning tasks is effectively solved, significantly improving the few-shot inference performance of large language models, achieving accurate example selection, and improving the matching degree between examples and test input text.
[0143] The context example selection system based on factorial latent variables and common factors provided in this application, employing the context example selection method based on factorial latent variables and common factors in the above embodiments, can solve the technical problem of context example selection based on factorial latent variables and common factors. Compared with the prior art, the beneficial effects of the context example selection system based on factorial latent variables and common factors provided in this application are the same as the beneficial effects of the context example selection method based on factorial latent variables and common factors in the above embodiments, and other technical features of the context example selection system based on factorial latent variables and common factors are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.
[0144] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0145] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0146] In addition, for technical details not described in detail in this embodiment, please refer to the context example selection method based on factorial latent variables and common factors generated in any embodiment of the present invention, which will not be repeated here.
[0147] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0148] It should be noted that the user information (including but not limited to user device information, user personal information, user location information, user behavior information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0149] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0150] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0151] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for selecting context examples based on factorial latent variables and common factors, characterized in that, The method is applied to natural language text processing, and the context example selection method based on factorial latent variables and common factors includes: Task features are extracted from the original context learning task data of the large language model, and a set of factorial latent variables is defined in the continuous embedding space of the large language model to form a factorial latent state space. The original context learning task data includes the input text and the corresponding output label for context learning. The task features are mapped and decoupled into mutually independent latent factor vectors in the set of latent variables in the latent factor state space, so as to reconstruct the input text and output labels into observations generated by the combined effect of the latent factors. The latent factor vectors are used to capture feature attributes of different dimensions in the context example. A multi-head variational encoder is constructed, and the multi-head variational encoder is trained using an objective function with a fully correlated decoupling penalty term. The multi-head variational encoder is configured to receive the concatenated embedding of input text and output labels and predict the distribution parameters of different latent factor vectors through multiple output heads. A sample size calculation model is constructed based on the coverage number theory, and the total retrieval budget parameters are configured through the sample size calculation model. Based on the total retrieval budget parameters and the trained multi-head variational encoder, a retrieval module is configured. The retrieval module is used to map the test input text to the factorial latent state space and parse the posterior distribution parameter set of the test input text in each latent factor dimension. In response to receiving test input text, the test input text is input to the retrieval module to obtain a posterior distribution parameter set. Based on the posterior distribution parameter set, the target point is determined, and a complementary example set covering all target points is retrieved from the candidate example pool through a centroid approximation strategy.
2. The context example selection method based on factorial latent variables and common factors as described in claim 1, characterized in that, The step of mapping and decoupling the task features into mutually independent latent factor vectors in the set of latent variables within the latent factorial state space, so as to reconstruct the input text and output labels into observations generated by the combined effect of the latent factors, includes: The task features are mapped to the factorial latent state space, and the task features in the factorial latent state space are decoupled by dimension to obtain a subset of task features. Each task feature subset is mapped to the corresponding vector of the factorial latent variable set to obtain the initial latent factor vector; Define independent standard normal marginal distributions for the initial latent factor vectors, and take the product of the marginal distributions as the joint prior distribution of the factorial latent variable set, as shown in the following formula: in, Describes the set of latent variables for factorial. The joint prior distribution, This represents the number of latent factor vectors. Indicates the first A vector of potential factors, Indicates the first Marginal prior distributions of potential factor vectors This represents a vector with zero mean and a covariance matrix that is the identity matrix. The normal distribution; The initial latent factor vector is optimized based on the joint prior distribution to obtain mutually independent latent factor vectors in the factorial latent state space. A generative topology is established with the set of factorial latent variables as global conditions, and the input text and output labels are concatenated into a joint observation sequence; The input text and output labels are reconstructed into observations generated by the combined effect of the latent factors through the generated topology.
3. The context example selection method based on factorial latent variables and common factors as described in claim 1, characterized in that, The construction of the multi-head variational encoder, which employs a target function with a fully correlated decoupling penalty term to train the multi-head variational encoder, includes: An approximate posterior distribution controlled by preset parameters is introduced, and this approximate posterior distribution is decomposed into independent sub-distribution structures corresponding to each latent factor vector, as shown in the following formula: in, Indicates that it is based on preset parameters The approximate posterior distribution of the control. This indicates the input text. Indicates preset parameters. Indicates the output label. Represents the set of latent variables for factorial. Indicates the first The mean vector of the potential factor vectors Indicates the first The variance vector of the latent factor vectors Represents the diagonal covariance matrix; Based on the aforementioned independent sub-distribution structure, a multi-output head architecture for a multi-head variational encoder is designed, and a multi-head variational encoder is constructed. Based on the principles of Bayesian learning, an evidence lower bound optimization objective is constructed, referring to the following formula: in, This indicates the objective term for optimizing the lower bound of the evidence. This represents the expectation operation over an approximate posterior distribution. Representing a given latent variable Time observation data The conditional log-likelihood, The parameters representing the large language model. KL divergence is used to measure the difference between two probability distributions. Indicates the first Approximate posterior distribution of the potential factor vectors, Indicates the first Sampled values of a latent factor vector, Indicates the first The standard deviation of the latent factor vector Let represent the auxiliary noise variable, which follows a standard normal distribution; A fully correlated decoupling penalty term is added to the lower bound of evidence optimization objective to form the objective function, as shown in the following formula: in, This represents the total loss term of the objective function. This represents a hyperparameter that controls the strength of decoupling. This represents the fully correlated decoupling penalty term. Denotes the aggregate posterior distribution of the set of latent variables for factorial. Indicates the first Marginal posterior distribution of each latent factor vector; The multi-head variational encoder is trained using the objective function so that it can predict the distribution parameters of different latent factor vectors through multiple output heads.
4. The context example selection method based on factorial latent variables and common factors as described in claim 1, characterized in that, The process of constructing a sample size calculation model based on coverage number theory and configuring the total retrieval budget parameters through the sample size calculation model includes: A sample complexity function is constructed based on the coverage number theory. This function is used to calculate the minimum number of examples required to control the inference error within a preset threshold under a pre-set confidence level. The sample complexity function is defined by the following formula: in, Indicates the reliability of the preset settings To control the distribution error to not exceed a preset threshold Minimum required sample size Indicates the distribution error threshold. The complement of the confidence level, Indicates a constant factor. Represents the latent state space of factorial. , with norm For measurement, radius is The coverage number reflects the factorial latent state space. The geometric complexity, Describing the latent state space of factorial The measure of entropy; The parameter space features of the factorial latent state space are imported into the sample complexity function to construct a sample size calculation model; The sample size calculation model is used to calculate the sample requirement under the single-unit characterization framework, which serves as a benchmark reference value. The sample size calculation model is used to calculate the lower bound of the example under the factorial representation framework; A sample budget allocation inequality is established based on the benchmark reference value and the sample lower limit quantity, and the total retrieval budget parameters are configured based on the sample budget allocation inequality, referring to the following formula: in, This indicates the retrieval of the total budget parameter. Indicates the first The minimum number of representative examples required for each feature dimension This represents the sample requirement under the factorial representation framework. This indicates the sample requirement under the single-unit representation framework.
5. The context example selection method based on factorial latent variables and common factors as described in claim 1, characterized in that, The retrieval module configuration based on the total retrieval budget parameters and the trained multi-head variational encoder includes: Import the total search budget parameter into the search module to configure the sample search quantity constraint of the search module; Based on the trained multi-head variational encoder, a parameter extraction submodule is built in the retrieval module. This submodule is used to extract the posterior distribution parameter set of the test input text in each latent factor dimension, as shown in the following formula: in, This indicates that the test input text is at the [number]th [position]. The posterior distribution over the latent factor vectors This represents the test input text. This indicates that the test input text is at the [number]th [position]. Central demand on a potential factor vector This indicates that the test input text is at the [number]th [position]. The variance vector over the latent factor vectors This indicates that the multi-head variational encoder is used for the test input text. The output of .
6. The context example selection method based on factorial latent variables and common factors as described in claim 1, characterized in that, The step of determining the target points based on the posterior distribution parameter set and retrieving a complementary set of combined examples covering all target points from the candidate example pool using a centroid approximation strategy includes: The mean and variance parameters of each latent factor dimension are extracted from the posterior distribution parameter set and integrated to form a feature parameter set; The set of feature parameters is used as the feature identifier of the test input text in the factorial latent state space and is determined as the target point. A weighted second-order Wasserstein distance metric is defined to calculate the distributional difference between the test input text and candidate examples in the candidate example pool, as shown in the following formula: in, Indicates the test input text With candidate examples A comprehensive measure of difference in the factorial latent state space. Indicates the first The weight coefficients of each potential factor dimension. Denotes the squared second-order Wasserstein distance between two Gaussian distributions. This indicates that the test input text is at the [number]th [position]. The posterior distribution over the latent factor vectors Indicates the candidate example in the th Posterior distribution over _ potential factor vectors; An optimization function is constructed based on the first-order moment centroid approximation strategy and the second-order Wasserstein distance metric, as shown in the following formula: in, This represents the final set of complementary examples. Represents the candidate example pool, Indicates the first One candidate example, Indicates the size of the candidate example pool. This indicates that the test input text is at the [number]th [position]. The mean of the potential factor vectors Indicates the first The candidate example in the first The mean of the potential factor vectors; Based on the optimization function, candidate examples that can correct set bias are selected from the candidate example pool, and the selected candidate examples are combined to obtain a complementary example set covering all target points.
7. The context example selection method based on factorial latent variables and common factors generation as described in any one of claims 1 to 6, characterized in that, The process of extracting task features from the original context learning task data of the large language model and defining a set of factorial latent variables in the continuous embedding space of the large language model to form a factorial latent state space includes: We acquire input text and output label pairing data for context learning scenarios in large language models, and perform data cleaning and standardization to form raw context learning task data. The continuous embedding space of the large language model is determined based on the parameter matrix of the large language model, and multi-dimensional task features are extracted from the original context learning task data within the continuous embedding space. Define a set of factorial latent variables consisting of multiple vectors in the continuous embedding space, and introduce orthogonality constraints for the set of factorial latent variables. By combining the set of factorial latent variables with orthogonality constraints and the task characteristics, a factorial latent state space is formed.
8. A context example selection system based on the generation of factorial latent variables and common factors, applying the context example selection method based on the generation of factorial latent variables and common factors as described in any one of claims 1 to 7, characterized in that, The system includes: The latent space construction module is used to extract task features from the original context learning task data of the large language model and define a set of factorial latent variables in the continuous embedding space of the large language model to form a factorial latent state space. The original context learning task data includes the input text and corresponding output labels for context learning. The task reconstruction module is used to map and decouple the task features into mutually independent latent factor vectors in the set of latent variables in the latent factorial state space, so as to reconstruct the input text and output label into observations generated by the combined effect of the latent factors. The latent factor vectors are used to capture feature attributes of different dimensions in the context example. The encoding configuration module is used to construct a multi-head variational encoder. The multi-head variational encoder is trained using an objective function with a fully correlated decoupling penalty term. The multi-head variational encoder is configured to receive the concatenation and embedding of input text and output labels and predict the distribution parameters of different latent factor vectors through multiple output heads. The budget configuration module is used to construct a sample size calculation model based on the coverage number theory, and to configure the total budget parameters for retrieval through the sample size calculation model. The retrieval configuration module is used to configure the retrieval module based on the total retrieval budget parameter and the trained multi-head variational encoder. The retrieval module is used to map the test input text to the factorial latent state space and parse the test input text to obtain the posterior distribution parameter set of each latent factor dimension. The example retrieval module is used to respond to receiving test input text, input the test input text into the retrieval module, obtain a posterior distribution parameter set, determine the target point based on the posterior distribution parameter set, and retrieve a complementary example set that combines and covers all target points from the candidate example pool through a centroid approximation strategy.
9. The context example selection system based on factorial latent variables and common factors as described in claim 8, characterized in that, The task reconstruction module is further configured to map the task features to the factorial latent state space, perform dimensional decoupling processing on the task features in the factorial latent state space to obtain a subset of task features; map each subset of task features to the corresponding vector of the factorial latent variable set to obtain an initial latent factor vector; define an independent standard normal marginal distribution for the initial latent factor vector, and use the product of each marginal distribution as the joint prior distribution of the factorial latent variable set, referring to the following formula: in, Describes the set of latent variables for factorial. The joint prior distribution, This represents the number of latent factor vectors. Indicates the first A vector of potential factors, Indicates the first Marginal prior distributions of potential factor vectors This represents a vector with zero mean and a covariance matrix that is the identity matrix. The normal distribution; The initial latent factor vector is optimized based on the joint prior distribution to obtain mutually independent latent factor vectors in the factorial latent state space; a generative topology is established with the set of factorial latent variables as global conditions, and the input text and output labels are concatenated into a joint observation sequence; the input text and output labels are reconstructed into observations generated by the combined action of the latent factors through the generative topology.
10. The context example selection system based on factorial latent variables and common factors as described in claim 8, characterized in that, The encoding configuration module is further configured to introduce an approximate posterior distribution controlled by preset parameters, and decompose the approximate posterior distribution into independent sub-distribution structures corresponding to each latent factor vector, as shown in the following formula: in, Indicates that it is based on preset parameters The approximate posterior distribution of the control. This indicates the input text. Indicates preset parameters. Indicates the output label. Represents the set of latent variables for factorial. Indicates the first The mean vector of the potential factor vectors Indicates the first The variance vector of the latent factor vectors Represents the diagonal covariance matrix; Based on the aforementioned independent sub-distribution structure, a multi-output head architecture for a multi-head variational encoder is designed, and a multi-head variational encoder is constructed. An evidence lower bound optimization objective is constructed based on the Bayesian learning principle, referring to the following formula: in, This indicates the objective term for optimizing the lower bound of the evidence. This represents the expectation operation over an approximate posterior distribution. Representing a given latent variable Time observation data The conditional log-likelihood, The parameters representing the large language model. KL divergence is used to measure the difference between two probability distributions. Indicates the first Approximate posterior distribution of the potential factor vectors, Indicates the first Sampled values of a latent factor vector, Indicates the first The standard deviation of the latent factor vector Let represent the auxiliary noise variable, which follows a standard normal distribution; A fully correlated decoupling penalty term is added to the lower bound of evidence optimization objective to form the objective function, as shown in the following formula: in, This represents the total loss term of the objective function. This represents a hyperparameter that controls the strength of decoupling. This represents the fully correlated decoupling penalty term. Denotes the aggregate posterior distribution of the set of latent variables for factorial. Indicates the first Marginal posterior distribution of each latent factor vector; The multi-head variational encoder is trained using the objective function so that it can predict the distribution parameters of different latent factor vectors through multiple output heads.
Citation Information
Patent Citations
Multi-rhetorical text generation method based on BART model
CN117725964A
Context learning example selection method based on two queries and low-rank approximate rearrangement
CN119903084A
Multi-dimensional example selection method and system applied to numerical reasoning task
CN120525061A