Supervision adjustment and task output method and device for language model, equipment and medium

By acquiring historical interaction records and user preference constraints, a candidate item set is generated using a sequential prediction model, and response labels are constructed. The language model is then supervised and adjusted, solving the problem of understanding natural language task instructions in the fintech and healthcare fields of existing recommendation systems. This achieves accuracy, controllability, and diversity in the output results.

CN121860055APending Publication Date: 2026-04-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing recommendation systems in the fintech and healthcare fields struggle to understand natural language task instructions, cannot accurately parse semantic constraints, and cannot generate outputs that meet diversity and proportion control requirements. They also lack high-quality supervision signals, resulting in insufficient interpretability and consistency of recommendation results.

Method used

By acquiring historical interaction records, user preference constraints, and output conditions, a candidate item set is generated using a pre-trained sequential prediction model. Filtering or sampling is performed based on task type to construct response labels, form instruction-response pairs, supervise and adjust the language model, and finally output the item list.

Benefits of technology

It achieves accuracy, controllability, and diversity in the output results generated in the fields of fintech and healthcare, and can accurately understand natural language task instructions and generate a list of items that conform to semantic constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860055A_ABST
    Figure CN121860055A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent decision making, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a language model supervision adjustment and task output method, device, equipment and medium, and the method comprises the steps: obtaining historical interaction records, user preference constraints and output conditions, generating a candidate item set by using a pre-trained sequence prediction model based on the historical interaction record; filtering or sampling according to the task type to obtain a processed candidate item set; combining the real project and the supplementary project to generate a response label; combining an instruction containing a historical interaction record, a user preference constraint and an output condition with a response label to form an instruction response pair; supervising and adjusting the language model by using the instruction response pair; and processing the task instruction through the adjusted language model to generate an item list. According to the method, the language model is trained in combination with the user behaviors and the task semantic information, so that the output result is more accurate, controllable and diversified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, and in particular to a method, apparatus, device, and medium for supervised adjustment and task output of a language model. Background Technology

[0002] In the fintech sector, existing recommendation systems are mostly based on collaborative filtering or sequence modeling methods, relying primarily on historical data such as transaction behavior and product click records to build user profiles. While these methods can capture users' static preferences and behavioral sequence patterns, they struggle to express dynamic recommendation intentions using fixed feature templates when faced with complex product constraints in financial business scenarios (such as risk levels, maturity matching, and regulatory compliance restrictions). Furthermore, traditional models rely on precise behavioral co-occurrence relationships and lack the ability to understand natural language task descriptions, making them unable to handle complex investment needs or product exclusion conditions expressed by users in linguistic form. For example, when a user inputs "recommend low-risk financial products with a maturity of no more than one year" in natural language, existing models cannot accurately parse semantic constraints or achieve controllable screening. In addition, in data annotation scenarios within financial businesses, the lack of supervised samples with clearly defined "ideal recommendation lists" leads to insufficient interpretability and consistency in the generated project lists, failing to meet regulatory requirements for diversity and proportion control.

[0003] In the healthcare field, existing recommendation systems also have limitations. Traditional medical recommendation algorithms primarily rely on structured data such as patient history records, medication usage data, and disease feature codes to build models. While this approach can capture the time-series characteristics of medical behavior, it struggles to understand the natural language instructions from healthcare professionals or the semantic health needs expressed by patients. For example, if healthcare professionals want the system to recommend medication regimens that are "non-hormonal antihypertensive drugs that do not affect liver function," existing models cannot parse exclusion criteria and specific medication preferences at the natural language level. Furthermore, privacy restrictions and sparsity issues with medical data are more pronounced. Real-world clinical samples often lack ideal treatment plans or recommendation lists with expert annotations, resulting in a lack of high-quality supervision signals during model training. In addition, medical recommendation tasks require a balance between safety, efficacy, and diversity, making it difficult for existing systems to simultaneously achieve accuracy and controllability, easily generating repetitive, overly concentrated, or unexplained recommendation results. Summary of the Invention

[0004] The main objective of this invention is to provide a method, apparatus, device, and storage medium for supervised adjustment and task output of a language model, aiming to solve the technical problem that existing technologies cannot accurately understand natural language task instructions and generate output results that meet semantic constraints based on user historical behavior information.

[0005] To achieve the above objectives, this invention provides a method for supervised adjustment and task output of a language model, comprising: Obtain historical interaction records, user preference constraints, and output conditions; Based on the historical interaction records, a set of candidate items is generated using a pre-trained sequential prediction model; Based on the task type, the candidate item set is filtered or sampled to generate a processed candidate item set; The real project is set as the base project for the response tag, and one or more projects are selected from the processed candidate project set according to the task type, and combined with the base project to form the response tag; The instruction containing the historical interaction record, the user preference constraints, and the output conditions is combined with the response tag to form an instruction-response pair; The language model is supervised and adjusted using the instruction response pairs. The input task instructions are processed by the adjusted language model, and a list of items is output.

[0006] Furthermore, to achieve the above objectives, the present invention provides a language model supervision adjustment and task output device, comprising: The data acquisition module is used to obtain historical interaction records, user preference constraints, and output conditions; The candidate generation module is used to generate a set of candidate items based on the historical interaction records using a pre-trained sequential prediction model. The task filtering module is used to filter or sample the candidate item set according to the task type to generate a processed candidate item set. The tag building module is used to set real projects as the base projects for response tags, and select one or more projects from the processed candidate project set according to the task type, and combine them with the base projects to form response tags; The instruction generation module is used to combine the instruction containing the historical interaction record, the user preference constraints, and the output conditions with the response tag to form an instruction-response pair; The model adjustment module is used to supervise and adjust the language model using the instruction response pairs; The inference output module is used to process the input task instructions through the adjusted language model and output a list of items.

[0007] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a language model supervision and adjustment and task output program stored in the memory and executable on the processor, wherein when the language model supervision and adjustment and task output program is executed by the processor, it implements the steps of the language model supervision and adjustment and task output method as described above.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a language model supervision adjustment and task output program, wherein when the language model supervision adjustment and task output program is executed by a processor, it implements the steps of the language model supervision adjustment and task output method as described above.

[0009] Beneficial Effects: This invention relates to the field of intelligent decision-making technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for supervised adjustment and task output of a language model, comprising: acquiring historical interaction records, user preference constraints, and output conditions; generating a candidate item set based on the historical interaction records using a pre-trained sequential prediction model; filtering or sampling the candidate item set according to the task type to generate a processed candidate item set; setting real items as the base items for response labels, and selecting one or more items from the processed candidate item set to combine with the base items to generate response labels; combining instructions containing historical interaction records, user preference constraints, and output conditions with response labels to form instruction-response pairs; supervising and adjusting the language model using the instruction-response pairs; and processing task instructions and outputting an item list through the adjusted language model. This invention combines user historical behavior information with task semantic constraints to generate response labels containing real and supplementary items, and uses these to construct instruction-response pairs for supervising the language model. This allows the model to simultaneously consider historical behavior patterns and natural language intent when generating output, thereby achieving accuracy, controllability, and diversity in the output results. Attached Figure Description

[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a language model supervision adjustment and task output method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the language model supervision adjustment and task output method of the present invention; Figure 3 A schematic diagram of the functional modules of a preferred embodiment of the language model supervision, adjustment and task output device of the present invention; Figure 4This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0011] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0012] The language model supervision adjustment and task output method provided in this embodiment of the invention can be applied to, for example, Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain historical interaction records, user preference constraints, and output conditions from the client. Based on the historical interaction records, it generates a candidate item set using a pre-trained sequential prediction model. The server filters or samples the candidate item set according to the task type to generate a processed candidate item set. Real items are set as the base items for response labels, and one or more items are selected from the processed candidate item set and combined with the base items to generate response labels. Instructions containing historical interaction records, user preference constraints, and output conditions are combined with response labels to form instruction-response pairs. The instruction-response pairs are used to supervise and adjust the language model. The adjusted language model processes task instructions and outputs a list of items. This invention combines user historical behavior information with task semantic constraints to generate response labels containing real and supplementary items. Instruction-response pairs are then used to supervise the language model, enabling the model to simultaneously consider historical behavior patterns and natural language intent when generating output, thereby achieving accuracy, controllability, and diversity in the output results. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The present invention will now be described in detail through specific embodiments.

[0013] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the language model supervision adjustment and task output method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0014] like Figure 2 As shown, the language model supervision adjustment and task output method proposed in this invention includes the following steps: S10: Obtain historical interaction records, user preference constraints, and output conditions; In this embodiment, the process of acquiring historical interaction records, user preference constraints, and output conditions is used to construct an input foundation that the language model can understand, so that the subsequent generated results can reflect individual differences and consistency with task objectives. Historical interaction records are derived from past behavioral data between users and the system, which may include sequence information such as clicks, browsing, selections, rejections, or dwell times. These behaviors reflect the evolution of users' interests and potential intentions over time. To ensure the accuracy and timeliness of this data, noise reduction and deduplication operations need to be performed during the collection phase, such as filtering recently active interactions through time windows and removing duplicate records or unrepresentative behavioral points. User preference constraints are used to define the user's bias towards content, attributes, or result types in a specific context, and can be generated by fusing static profile information and dynamic conversation information. Static profile information, such as age, occupation, risk tolerance, health status category, or asset structure, can provide long-term stable constraint references; dynamic conversation information extracts instantaneous preference features by analyzing semantic input, click sequences, and feedback instructions in real time. Output conditions represent structured parameters that constrain the scale and composition of the results during the generation phase. These parameters may include limitations on the number of outputs, the range of content diversity, the proportion of items distributed, and the coverage of results. They are obtained through system strategies or user commands and are used to balance the weight of accuracy and diversity in the results. The above three types of input features together form a semantically parsable data foundation. They are fused and represented through a unified data structure to ensure that the language model can identify the consistency between historical behavioral patterns and current intent in subsequent stages.

[0015] At the implementation level, multi-source interfaces can be configured in the data acquisition module to synchronize user behavior logs with external attribute databases, and a unified data parsing component can be used to standardize formats and map fields. In the preference extraction stage, a time decay function is introduced to strengthen the weight of recent behaviors, and a semantic clustering model is applied to extract potential interest topics from multi-turn dialogue records. In the output condition parsing stage, a strategy configuration file is loaded or a natural language parsing model is used to convert input commands into parameter templates to generate structured output control parameters. The above processes can be executed in batch or real-time modes, depending on the timeliness requirements of the business processing scenario and the data flow architecture design.

[0016] In financial business scenarios, historical interaction records can be extracted from transaction records and account interaction logs. For example, analyzing the browsing sequence and operational depth of users on wealth management product pages can construct a behavioral chain reflecting changes in risk preferences. User preference constraints can be derived from wealth management questionnaires and real-time transaction intent recognition models, while output conditions are generated according to regulatory requirements and product distribution strategies to control the proportion of output items across different risk levels. In the healthcare field, historical interaction records can be extracted from the operational trajectory and health records of individual health monitoring platforms. Preference constraints are determined by individual health goals and daily interaction habits, and output conditions are configured based on medical staff recommendations or health plans to limit the number and type of output items. In implementation, time series modeling modules can be used to enhance the sequential characteristics of behavior, and graph structure modeling can be used to capture the dependencies between interactive objects. In the preference recognition part, a self-attention mechanism can be introduced to enhance contextual understanding, and output control conditions can be dynamically adjusted through configurable templates.

[0017] This embodiment achieves precise alignment between the semantic and data layers during the input stage by uniformly collecting and structurally integrating historical interaction records, user preference constraints, and output conditions. This provides a stable contextual foundation for the subsequent language model generation stage, improving the accuracy and controllability of the generated results.

[0018] S20, Based on the historical interaction records, a set of candidate items is generated using a pre-trained sequential prediction model; In this embodiment, the process of generating a candidate item set based on historical interaction records using a pre-trained sequential prediction model is used to infer potential target items or content from the user's historical behavior sequence. The historical interaction records, as input sequences, contain information such as item identifiers arranged in chronological order, operation types, context states, or attribute features. These features are encoded and input into the sequential prediction model. The sequential prediction model is a deep learning structure capable of processing time-series data. Typical implementations include recurrent neural networks, long short-term memory networks, or transformation networks based on self-attention mechanisms. It generates prediction results by capturing implicit dependencies and positional correlations in the historical sequence. During the training phase, the model learns user behavior transition patterns and co-occurrence features between items; therefore, upon receiving a new interaction sequence, it can output the probability distribution of each candidate item based on the sequence state.

[0019] In implementation, the first step is to convert the item identifiers in the historical interaction records into vectorized representations. This is typically achieved using an embedding layer, mapping each item identifier to a fixed-dimensional embedding vector. These embedding vectors can be learned during model training or shared with downstream tasks through joint training. These embeddings are then sequentially input into the model. The model calculates the relationship between each item and its historical context using time steps or attention windows, outputting a predicted score for each candidate item. The candidate pool is then ranked based on the predicted scores, and the top-ranked items are selected to form a candidate item set according to a set threshold. In some implementations, regularization strategies can be added to prevent the model from overfitting to recent behavior, or a sliding window mechanism can be used to control the sequence length, thus balancing long-term interests and short-term preferences.

[0020] In fintech scenarios, user interaction sequences on wealth management platforms can be input into a sequential prediction model, with each interaction corresponding to a specific wealth management product or risk category. The model analyzes users' historical choices regarding product type, return period, and risk level, calculates prediction scores for different wealth management products, and generates a candidate set to provide input for subsequent response label generation. In healthcare scenarios, user health management behavior sequences, such as exercise plan adjustments, nutrition intake records, and monitoring data feedback, can be used as input. A pre-trained sequential prediction model predicts potential health improvement projects or personalized planning modules that the user might be interested in, generating a candidate set for task instruction expansion. In practical implementation, multi-layer attention mechanisms can be employed based on data characteristics to enhance the model's ability to perceive key time periods. Layer normalization and residual connection structures can also be used to improve the model's training stability and inference speed.

[0021] This embodiment constructs a sequential prediction model based on historical interaction records and generates a candidate item set. This allows the system to capture user behavior patterns and potential intentions in advance before semantic modeling, enabling it to extract a high-value candidate set that can be used by the language model from historical patterns. This reduces the invalid computation of the language model in the generation stage and improves the relevance and accuracy of the results.

[0022] S30, Based on the task type, the candidate item set is filtered or sampled to generate a processed candidate item set; In this embodiment, the process of filtering or sampling the candidate item set according to the task type to generate a processed candidate item set is used to screen and reconstruct the items predicted by the model before the generation stage, ensuring that the final input candidate set meets the task objectives and user preference constraints. The task type is a behavioral category determined by system policies or external input, indicating the semantic purpose and control method of this generated task, such as personalized tasks, diverse tasks, constrained tasks, or exploratory tasks. The task type not only affects the execution path of the filtering logic but also determines the specific operation method applied to the candidate item set.

[0023] During the filtering phase, the system eliminates items that do not meet the constraint criteria by comparing the candidate item set with the user preference constraints and output conditions. User preference constraints can be expressed as attribute restrictions, category exclusion, content relevance, or risk tolerance, among other dimensions. For example, when the task type is based on personalized preferences, the system calculates the matching score between the attribute features of the candidate items and the user preference vector. If the score is lower than a set threshold, the corresponding item is removed from the set. Output conditions may include quantity limits, diversity indicators, or distribution ratio requirements. These conditions must be considered simultaneously during the filtering operation to maintain the structural rationality of the candidate set.

[0024] During the sampling phase, the system selects an appropriate sampling strategy based on the task type. For task types emphasizing diversity, a similarity-based resampling method can be used. This method calculates the Euclidean distance or cosine similarity between items in the feature space of the candidate set, selecting items with greater differences to increase the coverage of the set. For task types emphasizing accuracy, a weighted probability-based sampling method can be used. This method performs sampling with or without replacement, using the predicted score as the weight, making the set more concentrated on high-confidence items. The system can introduce random perturbation parameters during the sampling process to dynamically balance accuracy and exploration. The final processed candidate item set is not only targeted but also ensures controllable diversity and stability when input into the language model.

[0025] In the fintech business, task types can be defined as risk control, return optimization, or multi-product exposure. When the task type is risk control, the system filters the candidate project set based on the user's risk level, eliminating products with risks exceeding the tolerance range, and uses weighted sampling during the sampling phase to highlight projects with stable returns. When the task type is multi-product exposure, the system controls the diversity of candidate projects through similarity resampling, making the output results more evenly distributed across different product categories. In the healthcare field, task types can be defined as intervention recommendation, monitoring optimization, or health education. When the task type is intervention recommendation, the system removes unsuitable projects based on the user's disease stage, lifestyle habits, and contraindications during the filtering phase, and resampling based on feature distance during the sampling phase to ensure the diversity of suggested solutions. When the task type is health education, the system can set sampling strategies to cover different topic modules, ensuring the integrity of the recommended content structure.

[0026] This embodiment filters or samples the candidate item set according to the task type, enabling dynamic optimization of candidate data before it enters the language model. This ensures that the data received during the generation stage meets both user preferences and task requirements, while maintaining a reasonable balance between diversity and accuracy. This process effectively reduces irrelevant or redundant input items, improving the relevance and stability of subsequent model generation results.

[0027] S40, set the real project as the base project of the response tag, and select one or more projects from the processed candidate project set according to the task type, and combine them with the base project to form a response tag; In this embodiment, real items are set as the base items for response labels. One or more items are selected from the processed candidate item set according to the task type and combined with the base items to form response labels. This process is used to construct a target output sequence that can be used for supervised learning by the language model. Real items represent the interaction objects actually completed by the user in historical interactions and can be regarded as correct result samples of the current task. By setting them as the first item in the response labels, it can be ensured that the training data contains the directionality of real behaviors. This operation gives the model a clear target reference during training, thereby learning the association patterns between items and the transition patterns of user behavior.

[0028] At the implementation level, the system first parses historical interaction records to identify the real projects. The selection of real projects can be based on chronological order, operation type, or contextual conditions; for example, selecting the most recently confirmed or completed project. When multiple candidates exist, a weighted mechanism can be used to determine the most representative project, such as sorting by interaction frequency, dwell time, or feedback score and selecting the highest-ranked one. After determining the base project, the system removes it from the processed candidate project set to avoid duplication and generates a set of candidate projects. Subsequently, the number of supplementary projects to be combined and the selection strategy are determined based on the task type.

[0029] For tasks emphasizing accuracy, the system can directly select the items with the highest prediction scores as supplementary items. For tasks emphasizing diversity, distance constraints or diversity metrics can be introduced into the top-ranked items to ensure that the selected items are highly differentiated in the feature space. The selected supplementary items and basic items are combined in sequence to form response labels. The order can be determined based on the relevance between items, prediction confidence, or domain rules. For example, in some applications, descending order can be used to reflect recommendation priority, or contextual order can be used to preserve logical coherence. The combined response labels are used to supervise the output target during the training phase, explicitly guiding the direction of language model generation.

[0030] In the fintech sector, real-world projects represent actual investment operations completed by users, such as purchasing funds or insurance products. The system extracts the latest transaction records from historical transaction logs as base projects, and then selects projects with complementary risk levels, return cycles, or asset classes from the processed candidate project set to form response labels. These labels are used to train the model to generate output results with a multi-layered asset allocation structure. In the healthcare sector, real-world projects can be health plans actually executed by users, such as diet plans or exercise tasks. The system extracts completed plan projects from historical health logs as base projects, and then selects supplementary projects of different intensities or domains from the candidate set based on task type (e.g., rehabilitation intervention or health promotion) to form response labels. This allows the model to learn the organizational patterns of results under different health goals. In specific implementations, a multi-task learning framework can be used to share the response label generation logic for different task types to improve the model's generalization performance.

[0031] This embodiment sets real-world items as the base items for response labels and dynamically selects supplementary items based on task type. This introduces real-world behavioral samples and semantic structure constraints into the training data, enabling the language model to learn generation patterns that more closely resemble actual user behavior. This approach effectively enhances the model's interpretability and the relevance of its output, allowing subsequent generation stages to accurately reflect user intent and task requirements while maintaining reasonable diversity.

[0032] S50, the instruction containing the historical interaction record, the user preference constraint and the output condition is combined with the response tag to form an instruction-response pair; In this embodiment, the process of combining instructions containing historical interaction records, user preference constraints, and output conditions with response labels to form instruction-response pairs is a crucial step in constructing supervised learning samples. This process organizes behavioral sequences, constraints, and model output targets into input-output pairs that can be learned by the language model. Instructions represent the natural language description of the input task, and response labels represent the expected output. Through their structured combination, the input-output mapping required for model training is formed.

[0033] In the implementation process, the system first converts historical interaction records, user preference constraints, and output conditions into natural language descriptions based on preset instruction templates. The templates can define fixed language structures and placeholders, such as "Based on the following historical records… please generate output that conforms to…", to ensure consistency and scalability of input instructions across different task types. The description of historical interaction records can include information such as chronological order, operational intent, or category labels; user preference constraints are expressed through explicit conditional statements, such as "preferring to select lower-risk projects" or "prioritizing exercise recovery-related solutions"; the output conditions are used to limit the scale or distribution of the generated results, such as "generate 5 candidate options" or "cover different categories".

[0034] Subsequently, the system formats the sequence of item identifiers in the response tags into natural language form according to the response template, enabling it to correspond to the instructions in the same language space. The response template can contain sequence structures, logical connectors, or result descriptions, allowing the language model to understand the logical dependencies between input and output during learning. After formatting, the instructions and responses are concatenated according to a predetermined combination format, such as with the instructions preceding the responses, and distinguished by special identifiers, for example, by adding "<Instruction Begins>" and "<Response Begins>" tags to the combined text. This structured concatenation helps the model identify boundaries during training, thereby enhancing the correspondence between input and output. The concatenated content constitutes the instruction-response pair, which is ultimately stored as training samples for use in the supervised adjustment phase.

[0035] In fintech scenarios, the system can generate natural language instructions using instruction templates, such as: "Generating a product portfolio that meets the fund distribution conditions based on the following historical investment records and risk preferences." The historical interaction records include the user's investment timeline and fund categories, user preference constraints reflect their risk tolerance and expected returns, and output conditions limit the number and proportion of projects. The response template converts the project identifiers in the response tags into natural language descriptions, such as "Recommended products A, B, and C, covering different risk levels." These are then concatenated to form a complete training sample, used to guide the language model in learning financial product output patterns. In healthcare scenarios, the system can generate similar natural language instructions, such as "Generating an exercise and nutrition plan that meets current energy needs based on the following health records and dietary preferences." The project identifiers in the response tags are formatted as "Recommended aerobic training, strength training, and a balanced diet," thus forming a complete instruction-response pair.

[0036] In different implementations, the system can employ different template control mechanisms: one approach ensures semantic stability through fixed templates, suitable for standardized training scenarios; another approach uses dynamic templates to automatically adjust sentence structure based on task type, enhancing the diversity and coverage of generated samples. The template filling process can incorporate conditional variables or contextual information, enabling the generated instructions and responses to simultaneously express task objectives and constraints.

[0037] This embodiment organizes behavior records, preference constraints, and output conditions into natural language instructions and combines them with response labels in a structured way. This enables a unified expression of training samples, allowing the language model to understand the correspondence between input tasks and target outputs within the semantic space. This process effectively improves the model's task alignment ability during supervised learning, making the model's generated results more controllable and interpretable.

[0038] S60, using the instruction response pair, perform supervised adjustment of the language model; In this embodiment, when using instruction-response pairs to supervise and adjust the language model, the instruction-response pairs are first processed through a data pipeline. Each instruction and its corresponding response are treated as a training sample, maintaining a one-to-one correspondence. Instructions and responses are segmented and encoded separately, retaining the instruction start marker and response start marker embedded in the previous step. A separator is inserted during sample concatenation to form a single-sequence input. Attention and loss masks are constructed. The attention mask is used to limit the autoregressive dependency direction, and the loss mask is only activated after the response start marker, ensuring that loss calculation only covers the response text, preventing the language model from generating gradients in the instruction part, thus confining the learning focus to the target output. For samples exceeding the maximum length, a priority-based truncation strategy is implemented, prioritizing the retention of the response segment and adjacent constraint prompts. Samples below the required length are padded, and the attention and loss masks are updated synchronously to ensure consistent tensor shape within the batch.

[0039] In the forward computation phase of the language model, the batch input consists of encoded sequence tensors, attention masks, and loss masks. After the language model generates the prediction distribution, word-by-word cross-entropy loss is calculated according to the label alignment method, and label smoothing is combined to suppress overconfidence. To address the risk of repetition or pattern collapse in the response, a repetition penalty regularization term is introduced. A light penalty is applied to high-frequency repetitions based on the statistics of n-gram segments within the response segment, with the upper limit of the coefficient constrained by the training stability threshold. If the instruction contains explicit segments of user preference constraints or output conditions, a constraint consistency weight vector is constructed. Response words semantically aligned with these segments are assigned higher loss weights to enhance the learning of key constraints. The final loss is the sum of the cross-entropy main loss, the repetition penalty term, and the constraint consistency weighted sum, with the weights adaptively fine-tuned according to the validation set stability.

[0040] The parameter update phase employs an efficient fine-tuning approach, injecting trainable parameters only into the attention projection, feedforward intermediate channel, or normalized channel, while freezing the base weights to reduce computational and storage overhead and decrease the probability of catastrophic forgetting. The optimizer uses an adaptive method with segmented learning rate scheduling: a short warm-up phase, a long plateau phase, and cosine fallback during convergence. To ensure numerical stability, mixed precision and gradient pruning are enabled, and abnormal gradient batches are rolled back and resampled. To improve throughput and memory utilization, gradient accumulation and gradient checkpointing techniques are employed. In multi-GPU or multi-machine environments, a combination of data parallelism and tensor parallelism is used, maintaining consistency in random seeding and data partitioning to ensure reproducible experiments.

[0041] The supervised adjustment process employs periodic evaluation and early stopping strategies. A validation set is separated from the instruction-response pairs, and perplexity, constraint hit rate, and diversity metrics are monitored. If there is no significant improvement for several consecutive cycles, early stopping is triggered, and the system rolls back to the optimal checkpoint. The checkpoint includes adaptation parameters, optimizer status, learning rate scheduler status, and vocabulary version signature, facilitating regression and comparative experiments. Training logs record batch-level loss, weight norm, gradient norm, and memory utilization. Anomaly detection rules identify divergence signs and automatically reduce the learning rate or shorten sequence length for self-healing. After completing the predetermined training cycle and meeting the stopping conditions, the adjusted language model is exported, containing only the basic weights and learned adaptation parameters. A metadata list bound to the data version, template version, and tagging specification is also generated to ensure consistency between subsequent inference and retraining.

[0042] This embodiment significantly improves output alignment by calculating the loss only on the response segment and constraining gradient propagation with a mask, focusing the training objective on output quality rather than instruction restatement. It learns task-related representations by efficiently fine-tuning parameters while freezing the basic weights, reducing computational and storage consumption and mitigating the risk of forgetting. Furthermore, by constraining consistency weighting and repetition penalties, it simultaneously strengthens adherence to user preference constraints and output conditions and suppresses pattern collapse during the learning process. This allows the adjusted language model to maintain general expressive power while more stably generating a list of items that meets constraints and possesses discriminative power.

[0043] S70 processes the input task instructions using the adjusted language model and outputs a list of items.

[0044] In this embodiment, the process of processing the input task instructions and outputting an item list through the adjusted language model is the core step in the model inference stage, demonstrating the application capability of the supervised and adjusted model in real-world tasks. This process takes the user's natural language task instructions as input signals and aims to generate an item list that conforms to semantic constraints and preference conditions as the output. Information transformation and result generation are completed through forward computation, semantic parsing, and structured mapping.

[0045] In implementation, the input module first receives the task instruction text, performs word segmentation and encoding on the input content, and transforms natural language into a vector representation that the model can recognize. This representation not only contains word-level semantic information, but also retains the preference expressions, constraints, and output size requirements in the instructions. To ensure consistency between the input and training phases, the model loads the same vocabulary and encoding rules as during training, and synchronously loads the adaptation parameters generated during the supervised adjustment phase, so that the inference process uses a consistent semantic distribution space.

[0046] The forward propagation phase of the model is computationally based on a self-attention mechanism. After the input sequence is mapped through an embedding layer, it enters a multi-layer transformation structure, where the model computes the semantic contextual dependencies of the task instructions layer by layer. By adjusting the weights learned in the supervised adjustment phase, the model can more accurately capture semantic fragments in the input corresponding to historical interaction records, user preference constraints, and output conditions, thus generating output text that better conforms to the constraints in the decoding phase. During generation, the model employs an autoregressive generation mechanism to predict the probability distribution of the next word word by word, and can control the diversity and stability of the output through temperature parameters and top-k filtering. The generated natural language prediction response contains item information and logical order understood by the model based on the input task.

[0047] The response parsing phase transforms the predicted response into a structured form. The parsing module identifies item identifiers in the predicted text using regularization rules, keyword matching, or semantic tagging. For example, in a financial task, the identified identifiers might be product codes or fund categories; in a medical task, they might be health plan items or service numbers. The parsed results are formatted and mapped to a database or knowledge index, generating a complete item list by retrieving the corresponding item information. During the retrieval process, item descriptions, category tags, or attribute indicators can be added based on preset filtering parameters, making the output both readable and business-usable. The final output module converts the item list into a visual display structure, such as JSON format, a table structure, or a natural language summary, to adapt to different interactive interfaces or downstream systems.

[0048] In different implementations, the generation control strategy for model inference can vary. For example, length constraints or domain switching can be achieved by appending control flags to instructions; a dynamic decoding algorithm can also be introduced during the generation process to trigger resampling when the confidence level falls below a threshold to avoid invalid output. For multi-turn input scenarios, the system can also retain the dialogue context state and maintain context consistency in consecutive tasks through an attention caching mechanism.

[0049] This embodiment achieves a seamless mapping from natural language input to structured output by applying a language model to process task instructions and generate an item list after supervised adjustment. This allows the model to understand the user's semantic intent while also considering constraints and preferences. This process significantly improves the semantic consistency and controllability of the output, ensuring stable performance in terms of logical order, content matching, and diversity of the generated item list. By combining parsing and retrieval, the model output not only possesses natural language interpretability but can also be directly transformed into actionable structured results, thereby enabling semantically driven, high-precision item generation and result output in various business scenarios.

[0050] This invention relates to the field of intelligent decision-making technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for supervised adjustment and task output of a language model, comprising: acquiring historical interaction records, user preference constraints, and output conditions; generating a candidate item set based on the historical interaction records using a pre-trained sequential prediction model; filtering or sampling the candidate item set according to the task type to generate a processed candidate item set; setting real items as the base items for response labels, and selecting one or more items from the processed candidate item set to combine with the base items to generate response labels; combining instructions containing historical interaction records, user preference constraints, and output conditions with response labels to form instruction-response pairs; supervising and adjusting the language model using the instruction-response pairs; and processing task instructions and outputting an item list through the adjusted language model. This invention combines user historical behavior information with task semantic constraints to generate response labels containing real and supplementary items, and uses these to construct instruction-response pairs for supervising the language model. This allows the model to simultaneously consider historical behavior patterns and natural language intent when generating output, thereby achieving accuracy, controllability, and diversity in the output results.

[0051] In one embodiment, step S10 above includes: S101, Obtain raw data including user operation sequences, user static profiles and user dynamic session data from user behavior logs and user attribute databases; S102, Process the user operation sequence, extract the item identifier and arrange them in chronological order to form a historical interaction record; S103, Analyze the user static profile and the user dynamic conversation data to identify the user's tendency to express opinions on project categories or attributes, and form user preference constraints; S104: Analyze the system's preset strategy or the user's real-time request to determine the size and distribution requirements of the output item list, thus forming the output conditions.

[0052] In this embodiment, during the data collection and processing stage, raw data is synchronized from user behavior logs and the user attribute database. This raw data covers user operation sequences, static user profiles, and dynamic user session data. User behavior logs include event time, event type, item identifier, and contextual metadata. After unifying the time base, character encoding, and field naming, the data is written to the raw buffer. The user attribute database provides stable fields such as age group, geographic tags, long-term interest tags, and blacklists / whitelists. Dynamic user session data comes from interaction session records and query expressions, including text fragments, clickstreams, and page context. During synchronization, missing value processing, abnormal timestamp correction, duplicate event merging, and user identifier desensitization (such as hashing and salt concatenation) are performed. Field-level validation and partitioning for disk storage ensure efficient subsequent retrieval.

[0053] Subsequently, in the historical interaction record generation stage, only event types related to the project (such as browsing, clicking, favoriting, placing orders, etc.) are extracted from the user's operation sequence. The event payload or link parameters are parsed to obtain the project identifier. Multiple duplicate events within the same session are deduplicated and merged, and each interaction is assigned a weight and contextual features (device type, source channel, geographical granularity). To resolve conflicts between multiple source identifiers, parsing priority and consistency rules are set. For example, when both the event payload and URL parameter exist, the payload takes precedence; when multiple identifiers appear in the same time window, the one with higher confidence is used. Finally, the historical interaction record objects are arranged in chronological order. The minimum required fields include project identifier, timestamp, interaction type, interaction weight, and context fingerprint, stored in a structured table that can be retrieved by user and time partition to support sequential modeling and replay verification.

[0054] In the user preference constraint extraction stage, user static profiles and dynamic user session data are combined to identify biased expressions targeting project categories or attributes. Static profiles provide long-term, stable inclusion or exclusion lists and priority weights; dynamic session data provides short-term expressions and temporary constraints. Phrase(s) containing intentional or negative words are identified through template matching and a lightweight classifier to extract target categories, target attributes, and biases. These are then standardized and mapped against available category dictionaries and attribute enumerations. Long-term and short-term sources are merged into a user preference constraint object, including constraint type (inclusion, exclusion, priority), scope (category, attribute, tag), effective period, priority, and conflict resolution strategy. Conflict resolution employs a combination of time-based decay and a priority table to ensure that short-term strong instructions cover long-term weak instructions within their validity period and automatically fall back upon expiration. This object is exposed as a key-value structure to ensure direct integration with the filtering or sampling logic in the candidate processing stage.

[0055] In the output condition parsing stage, the scale and distribution requirements are jointly determined based on the strategy library and the user's real-time request. The strategy library records the default list length, deduplication radius, homology restriction, category ratio, and minimum difference threshold, serving as a safety baseline. The user's real-time request may contain quantifiers and ratio terms (e.g., "give a few," "distribute evenly by type," "more new products"). The parameter parser extracts the target length and ratio vectors, and performs boundary pruning and legalization. When a conflict arises between the strategy library and the real-time request, a decision is made based on priority and safety thresholds. For example, if the target length exceeds the upper limit, it falls back to the upper limit; if the ratio vector does not cover all available categories, it is supplemented with an even distribution strategy. After the decision is made, an output condition object is generated, containing fields such as target length, category or attribute ratio, homology suppression threshold, difference lower limit, and random perturbation amplitude, so that downstream candidate processing stages can perform filtering, resampling, and order control accordingly.

[0056] This embodiment achieves hierarchical collection, parsing, and standardized integration of historical interaction records, user preference constraints, and output conditions. On the input side, without altering the downstream interface, it simultaneously ensures temporal integrity, accuracy of intent expression, and executability of quantity diversity constraints. Historical interaction records provide stable temporal evidence for sequential modeling, user preference constraints express long-term and short-term intentions isomorphically to directly participate in filtering and sorting, and output conditions implement scale and distribution control through structured parameters. The combined effect of these three elements enables the subsequent generation process to accurately match individual intentions while maintaining controlled result length and category coverage, significantly reducing reliance on manual annotation and improving result consistency and interpretability.

[0057] In one embodiment, step S20 above includes: S201, convert the item identifier in the historical interaction record into the corresponding item embedding vector to form a serialized embedding representation; S202, the serialized embedding represents the input to the pre-trained sequential prediction model; S203, The sequential embedding representation is processed by the sequential prediction model to determine a corresponding prediction score for each item in the candidate pool; S204, Sort the items in the candidate pool in descending order according to the predicted scores, and select a preset number of items at the top of the sort to form a candidate item set.

[0058] In this embodiment, the processing chain for generating a candidate item set based on historical interaction records using a pre-trained sequential prediction model includes four consecutive steps: representation construction, model inference, score merging, and result selection. It revolves around the item identifiers in the historical interaction records, completing the mapping from symbols to vectors, from sequences to conditional distributions, and from distributions to ordered sets. First, the conversion from item identifiers to item embedding vectors is performed. To ensure consistency, item identifiers use a unified namespace and dictionary version, establishing a dense vector table. When an unregistered identifier appears, a backoff strategy is activated, including mapping to a dedicated unknown vector or generating an approximate vector based on the item attribute encoder, and recording the mapping source for subsequent auditing. Historical interaction records are retained in chronological order, with each interaction containing an item identifier, timestamp, and optional weights. The corresponding item embedding vectors are concatenated in chronological order to obtain a serialized embedding representation. To handle inconsistent lengths, the sequence uses a combination of fixed-length windows and sliding padding. For very long sequences, a truncation strategy based on the closest available time is used; for very short sequences, a dedicated vector is padded at the beginning position with learnable padding markers to ensure that time alignment and masking logic can be correctly recognized by the model. Serialization embeddings represent the superposition of positional encoding and session segmentation markers. Positional encoding is used to characterize relative order, while session segmentation markers are used to distinguish semantic breakpoints across sessions. Both are fed into the model's input channel.

[0059] The sequential prediction model employs pre-trained parameter states. Conditional on sequential embedding representations, the model outputs a conditional score for each item in the candidate pool. To ensure prediction consistency, a mask constraint is applied to the input side to shield information leakage after the current time step. Positive and negative sample constructions used during training are replaced during inference with either a full candidate pool scoring strategy or a segmented scoring strategy. The segmented scoring strategy divides the candidate pool into hash-based buckets, scores each batch, and merges them externally to avoid memory bottlenecks. Internal model computation includes self-attention stacking, feedforward units, and normalization units. The attention mask only allows historical positions to contribute weight to the current query position. To improve the perception of time decay, a learnable decay factor based on time difference is multiplied on the input side or attention weights, making recent interactions have a greater impact on the score. For frequently repeated item identifiers, a repetition suppression term is introduced to prevent the model from being biased due to short-term high frequencies.

[0060] The scoring merging phase extracts predicted scores from the model output. If a segmented scoring strategy is used, the predicted scores are first sorted locally in descending order within each segment, and then the top few elements from each segment are assembled into a globally ordered stream through multi-way merging. To suppress noise introduced during training or data acquisition, the predicted scores are numerically stabilized and outlier truncation is performed before merging. When the predicted scores of different items are the same or have very small differences, a deterministic scrambling rule is adopted, prioritizing the retention of items that were most recently encountered in the historical interaction records but did not reappear in the current window, in order to reduce backflow. If the historical interaction records contain explicit exclusion information, a one-time filter is applied after merging to ensure that the final sorting does not violate the explicit exclusion.

[0061] In the results selection phase, a candidate item set is generated by selecting the top-ranked elements from the globally ordered stream based on a preset number. This preset number, derived from system parameters or upstream strategies, is not changed at this stage. When the number of valid elements is insufficient, all valid elements are retained and marked "target quantity not met" in the metadata for downstream supplementation or rollback. When too many items from the same source affect diversity, a maximum percentage threshold is set for items from the same source or parent category. Items exceeding this threshold are moved backward and given priority to subsequent items from different sources without altering the overall order. The final output candidate item set is an ordered list, where each element contains at least an item identifier and a snapshot of its corresponding predicted score, along with the dictionary version, window range, and mask configuration used during generation, ensuring reproducible experiments and online playback.

[0062] This embodiment achieves stable encoding of sequence information by uniformly mapping item identifiers to item embedding vectors and constructing a serialized embedding representation. The pre-trained sequence prediction model outputs a prediction score covering the candidate pool under the constraints of masking and time decay, preserving long-term preferences while reinforcing recent intentions. Through numerical stabilization, parallel shuffling, and control of the proportion of homogeneous items, the candidate item set avoids homogenization and information leakage while adhering to the ordering credibility. This enables downstream processing to receive higher-quality, reproducible, and ordered inputs with contextual metadata within a fixed preset quantity, thereby improving overall accuracy and controllability and reducing the frequency of backtracking and manual intervention.

[0063] In one embodiment, step S30 above includes: S301, determine the current task type based on the user preference constraints and the output conditions; S302, set the candidate item set as the current processing set; S303, if the task type includes a task type based on personalized preference constraints, then remove all items that do not conform to the user preference constraints from the current processing set and update the current processing set; S304, if the task type includes a task type based on diversity constraints, then the current processing set is resampled according to the diversity index and the current processing set is updated. S305, use the updated current processing set as the candidate item set after processing.

[0064] In this embodiment, the goal is to transform a set of candidate items into a processed set of candidate items that meets constraints and is executable, given a task type, while maintaining compatibility with the preceding ranking signal. First, the task type is determined based on user preference constraints and output conditions. User preference constraints are derived from explicit inclusion intentions, exclusion intentions, attribute tendencies, threshold preferences, and other information in the user's static profile and dynamic conversations; output conditions specify scale and distribution requirements. To form a determinable task type, two types of Boolean flags are established: personalized preference constraint flags and diversity constraint flags. Personalized preference constraint flags are set to true when there is explicit inclusion or exclusion, while diversity constraint flags are set to true when there are distribution requirements such as category coverage, source proportion, similarity suppression, or time dispersion. If both flags are true, the task type is considered a combined type; if both are false, the task type remains the default type and is processed by the downstream fallback strategy. The flag decision needs to be traceable, recording the trigger field, value range, and parsing version to avoid ambiguity.

[0065] The candidate project set is then set as the current processing set, and metadata such as sorting scores, grouping labels, source information, and timestamps are copied to ensure stable order management and constraint validation during subsequent processing. To avoid introducing non-repeatable selection results in concurrent scenarios, a unique processing batch identifier is assigned to the current processing set, and the dictionary version and partition parameters are locked.

[0066] When a task type includes personalized preference constraints, a rule-based filtering process is executed. First, executable decision expressions are parsed from the user preference constraints, including explicit inclusion lists, explicit exclusion lists, attribute matching predicates, and threshold conditions. Explicit exclusions have higher priority than explicit inclusions; when the same item is both included and excluded, it is excluded and the conflict is recorded. Attribute matching predicates cover dimensions such as category, tag, source, price range, or risk level; threshold conditions are used to limit the lower limit of rating, quality, or timeliness indicators. For each element in the current processing set, it is calculated whether it satisfies inclusion or attribute matching without triggering exclusion or threshold failure. Elements that do not meet the criteria are removed from the set, and the removal reason, triggering condition, and timestamp are recorded in the metadata. The current processing set is then updated. To reduce the risk of an empty set due to excessive shrinkage, a protection lower bound is introduced. When the number of elements to be retained is lower than the size requirement given by the output conditions, only mandatory explicit exclusions and safety thresholds are retained, while other soft constraints are deferred to the sampling stage.

[0067] When the task type includes diversity constraints, a resampling process is executed to ensure the set meets distribution requirements without compromising the reliability of the original sorting. First, diversity metrics are analyzed based on the output conditions. Common metrics include category coverage, upper limit for source proportion, upper limit for similarity, and temporal dispersion. To perform verifiable resampling, three operations are defined: stratified sampling, similarity suppression, and proportion shaping. Stratified sampling divides the current processing set into strata based on category or source, assigning a target quota to each stratum. The quota is derived from distribution requirements or adaptively calculated based on the remaining size. Within each stratum, the original sorting order is maintained, selecting elements one by one until the stratum quota is reached. Similarity suppression performs neighbor elimination on the item representation vector or offline similarity matrix. The rule is that when an element is selected, subsequent elements with similarity higher than a threshold are temporarily deferred and added to a waiting list. The threshold can be set as a single value based on the output conditions or given segmented values ​​based on the stratum and window size. Proportion shaping calculates the source proportion within a stratum or globally. If the upper limit is exceeded, the excess elements are deferred, and the total size is maintained by adding elements from different sources from the waiting list. The three types of operations are executed alternately in the order of "stratified sampling first, similarity suppression in parallel, and proportion shaping convergence" until the scale requirement is met or the traversal is completed. During the process, all shuffling, deferral, and replacement adhere to deterministic rules: for scores of the same value, priority is given to those more recent in time; for scores of the same value, priority is given to those ranked higher in the global sorting, to ensure reproducibility. After processing, the current processing set is replaced with a new ordered list.

[0068] The updated current processing set is output as the processed candidate item set. The output carrier retains the order, size, triggering preference conditions, diversity indicators used, list of suppressed elements and replacement relationships, facilitating downstream combination of response tags and post-event auditing. If the final size is still insufficient, a "size not met" flag is written and alternative filling suggestions are provided, such as relaxing the similarity threshold or reducing the lower limit of the intra-layer quota, but the previous decision is not directly changed at this stage to maintain clear boundaries of responsibility.

[0069] This embodiment decouples and composably executes personalized filtering and diverse resampling by determining the task type based on user preference constraints and output conditions. This can eliminate explicit conflicts, reduce homogenization, and meet scale and distribution requirements while preserving the credibility of the original ranking. The filtering stage prioritizes explicit exclusion and threshold constraints to prevent non-compliant items from entering the downstream. The resampling stage uses stratification, similarity suppression, and proportion shaping in synergy to improve category coverage and source balance.

[0070] In one embodiment, step S40 above includes: S401, determine an item as the real item from the interaction items contained in the historical interaction record; S402, Set the real item as the first item in the response tag; S403, Remove the real project from the processed candidate project set to form a set of candidate projects; S404, Based on the task type and the preset response tag length, determine the number of supplementary items to be selected from the candidate item set; S405, according to the order of the items in the processed candidate item set, select the items with the highest number of supplementary items from the candidate item set as supplementary items; S406, combine the first item and the supplementary item in sequence to form a response tag.

[0071] In this embodiment, the goal is to use real projects as a benchmark, combined with task type and response tag length, to stably and reproducibly extract supplementary projects from the processed candidate project set, ensuring that the generated response tags maintain consistent order while meeting quantity and constraint requirements. First, real projects are determined from the interaction projects contained in historical interaction records. This can be determined by one of three factors: time window, interaction intensity, or the most recent valid interaction, with priority given to the most recent valid interaction. When there are ties, the interaction intensity takes precedence, followed by the most recent interaction. The selection criteria and judgment version are recorded for traceability. Once determined, the real project is written into a structured container and marked as the first project, along with metadata such as interaction timestamp, source channel, and context instruction summary for subsequent consistency verification and auditing.

[0072] When setting a real item as the first item in the response tag, its position index in the final sequence needs to be frozen to zero, locking the element's non-replaceable attribute, and blocking any operations that would change the zero-position element during the response tag construction process. To avoid duplicate selections, real items must be removed from the processed candidate item set, forming a shortlisted item set. The removal process employs a two-level strategy: precise identifier matching and semantic nearest neighbor joint verification. The first layer performs constant time hash removal based on the item identifier, and the second layer removes duplicates based on nearest neighbors with embedding similarity within a threshold. The threshold is derived from the similarity upper limit set in the previous filtering or sampling stage, maintaining consistency across stages. The removal results are written into the metadata with the removal reason, reference threshold, and nearest neighbor count to ensure replayability.

[0073] The number of supplementary items is calculated based on the task type and the preset response tag length. The response tag length is a fixed or range parameter of the target output size; when defined as a range, its specific value is determined by the output conditions. The number of supplementary items equals the target length minus one. If the task type includes cases where only a single real item is displayed, the number of supplementary items is rolled back to zero. For combined task types, an in-layer quota table can be introduced to constrain the proportion of different categories or sources, allocating the total number of supplementary items to each layer. The allocation rule follows the principle of rounding to the nearest whole number and maintaining total quantity consistency. If necessary, minimum cost adjustments are used to maintain consistency in total quantity. When the set of candidate items is insufficient to meet the quota, a rollback strategy is triggered, transferring the quota of the unmet portion to the remaining layers. The layer selection is based on historical coverage deficiency or global ranking density; if both are insufficient, backfilling is performed in global order.

[0074] The selection of supplementary items follows the order of the processed candidate item set, with the highest-ranked item being the sole selection criterion. In cases of ties, a deterministic decision rule is applied. The decision order prioritizes higher prediction scores, source diversity, recent publication date, and lexicographical order of identifiers, ensuring consistent results across different execution environments. Three types of lightweight checks are performed during the item-by-item selection process: First, duplication checks ensure no duplication with the first item or with already selected supplementary items; second, constraint checks check if the distribution exceeds limits for task types with diversity or percentage restrictions before adding a candidate, skipping it and adding it to the candidate queue if it does; third, availability checks skip items marked as offline or frozen. The process stops immediately when the target number of supplementary items is reached; if the main sequence is still insufficient, items are added according to the candidate queue order, with deterministic decision rule applied within the queue as well. When addition fails, an insufficient size flag and gap value are recorded, providing explicit context for downstream training sample construction.

[0075] The first item and supplementary items are combined sequentially to form response tags. During combination, a stable sequence number and corresponding metadata mapping are maintained, and the source trajectory of each element is written, including its position from the candidate item set, the selection result, the type of constraint triggered or not triggered, and the quota position. If the task type includes a fixed order requirement, only local swaps are allowed within supplementary items to mitigate minor distribution out-of-bounds errors. When the task type does not include a fixed order requirement but includes exposure balancing, limited adjacent swaps can be performed in supplementary items near the threshold boundary without changing the zero position of the first item. The swap criterion is that the benefit of out-of-bounds mitigation outweighs the ranking loss, and the ranking loss is measured by the difference in local positions. Finally, response tags and end-to-end audit records are output to ensure that subsequent instruction responses can reproduce the same tag sequence during the construction and monitoring adjustment phases.

[0076] This embodiment uses real projects as the first fixed benchmark, combined with deterministic selection based on ranking and quota allocation driven by task type, to ensure that response tags possess both stable and reproducible attributes as well as controllable and diverse attributes. Removal and nearest-neighbor deduplication reduce duplication and false differences, maintaining incremental tag information. The number and distribution of supplementary projects are calculated by explicit rules, triggering a fallback when insufficient, avoiding imbalances in scale and proportion. The entire process records the ranking, adjudication, and constraint triggering, providing an auditable and replayable data link for downstream command response pair construction and monitoring adjustments, thereby improving tag quality and constraint consistency without compromising the credibility of upstream ranking.

[0077] In one embodiment, step S50 above includes: S501, according to the preset instruction template, the historical interaction record, the user preference constraint and the output condition are combined into a natural language instruction; S502, according to the preset response template, the sequence of item identifiers in the response tag is formatted into a natural language response; S503, the natural language instruction and the natural language response are concatenated according to a preset combination format, and an instruction start marker and a response start marker are added to the concatenated content to form a structured instruction-response pair; S504, the instruction response pair is stored as a training sample.

[0078] In this embodiment, aiming at the construction of training data, historical interaction records, user preference constraints, and output conditions are reconstructed into natural language instructions through templates, and response tags are standardized into natural language responses. These are then concatenated into structured instruction-response pairs according to a preset combination format and explicit boundary markers, and finally stored as training samples with traceable metadata. First, historical interaction records are standardized, unifying the writing rules for time expressions, entity aliases, and item identifiers, removing noisy events and missing fields, and retaining key fragments that can be used to express temporal preferences. For user preference constraints, inclusion and exclusion expressions are distinguished, and categories, attributes, thresholds, and weights are mapped to phrases that can be directly referenced in template slots. For output conditions, the source and priority of values ​​for scale requirements and distribution requirements are clarified. When conflicts exist, an explicit priority table is used to adjudicate and form an adjudication record. The instruction template uses a slotted phrase set, covering temporal cue slots, inclusion constraint slots, exclusion constraint slots, quantity slots, and distribution slots. Slot filling follows a triple constraint of field existence detection, value validity verification, and length budget control. When a field is missing, a fallback phrase is triggered to maintain semantic integrity. All numerical or numbered items undergo desensitization mapping and reversible annotation dictionary binding, ensuring both privacy and replayability. After template instantiation, the natural language instruction undergoes syntactic simplification, ambiguous phrase replacement, and conjunction deduplication to ensure semantic clarity and ease of model parsing. Simultaneously, metadata corresponding to each instruction is generated, including template version, slot filling value, desensitization mapping table, and length budget allocation.

[0079] The response template is used to transcribe the sequence of item identifiers in the response tags into a natural language response, prioritizing the preservation of sort order and source clues to ensure consistency with the upstream sorting. Each entry consists of three parts: a stable separator, a position tag, and an item alias. The item alias is obtained through a dictionary mapping from identifiers to readable names. When an alias is missing, it falls back to a shorter identifier and adds the minimum identifiable attribute to resolve name conflicts. To support parsing, the separators between entries and the field separators within entries are orthogonal and globally unique. All separator character sets and escape rules are fixed in the template version and written into the metadata. If the response tag length is less than the size requirement, the actual length is retained and a size deficiency flag is added; if the length exceeds the requirement, it is truncated according to the sort order, and the truncated position and the summary of the truncated entry are recorded for auditing purposes.

[0080] The combination process employs a preset combination format, concatenating natural language commands and responses within the same sample. The combination format provides two equivalent representations: sequential text concatenation and key-value structure representation. For sequential text concatenation, command start markers and response start markers are injected before each text segment. These markers are fixed-length strings that cannot appear in the main text and include version number, paragraph type, and checksum, facilitating rapid downstream location. Key-value structure representation uses fixed key names to carry both text segments and the same set of marker fields, ensuring consistent parsing across multi-channel training pipelines. To mitigate cross-platform line break differences, line breaks and whitespace are uniformly replaced with a controlled character set before combination. Immediately after sample generation, hash checksums and deduplication fingerprints are calculated, establishing a fingerprint-based list for duplicate detection and outlier isolation. Simultaneously, sample-level metadata is written, including historical interaction record summaries, user preference constraint summaries, output condition summaries, response tag position distribution, template version, tag version, and generation time. Training samples are stored on disk using row-based indexing. The index key covers template version, time interval, and length quantiles, facilitating subsequent binning by version and difficulty. To prevent data shift, lightweight statistical slices are constructed to record instruction length distribution, response length distribution, class coverage, and distribution requirement satisfaction, and are stored in the sample bypass to support pre-training validation. The entire process ensures that the same input and template produce byte-level consistent instruction-response pairs under any environment, facilitating comparison and playback.

[0081] This embodiment obtains structurally stable and semantically clear natural language instructions by reconstructing historical interaction records, user preference constraints, and output conditions within the template system in a slot-based manner. Combined with standardized transcription and positional fidelity of response tags, it ensures consistency between text expression and upstream sorting. The introduction of explicit start markers and combination formats clarifies parsing boundaries and facilitates cross-pipeline consumption. By integrating hash fingerprint deduplication and metadata auditing mechanisms, it forms a replayable, traceable, controllable length budget, and distributed training sample set, thereby reducing semantic ambiguity and annotation noise, improving the consistency and interpretability of supervision signals, providing stable and reliable input-output pairs for subsequent parameter updates, and maintaining the determinism and auditability of the generation process even when scale and distribution requirements change.

[0082] In one embodiment, step S70 above includes: S701 receives task instructions in natural language form; S702, The task instruction is input into the adjusted language model, and the task instruction is processed by the adjusted language model to generate a predicted response in natural language form; S703, parse the structured item identifier from the predicted response; S704, Based on the project identifier, retrieve the corresponding project information and generate a project list; S705, Output the list of items.

[0083] In this embodiment, the input channel receives task instructions in natural language form, unifies character encoding and delimiter sets, removes invisible control characters and redundant whitespace, and retains line breaks and item separator semantics. Language detection and domain-specific vocabulary normalization are performed on the task instructions, using an industry dictionary to map synonymous expressions to standard field and attribute names, maintaining a one-to-one correspondence with fields in the preceding data. To facilitate subsequent parsing, length budget and termination flag configurations are injected into the task instructions, limiting generation boundaries and recording the context summary, timestamp, and call source of this interaction, forming a traceable input snapshot. If the task instructions carry explicit constraints such as user preferences, scale requirements, or distribution requirements, constraint description fragments are written according to a predetermined priority, and a constraint index is generated. The index participates in consistency verification during subsequent parsing.

[0084] When a task instruction is entered into the adjusted language model, an inference context is constructed, including the instruction text, constraint index, historical interaction summary pointer, and response format guide phrase. The decoding strategy employs one of two modes: deterministic decoding or controlled decoding. The former is used for scenarios requiring strictly reproducible output, disabling random sampling and enabling a fixed termination marker; the latter is used for scenarios requiring coverage of a broader exploration space, allowing controlled perturbations within a limited candidate set. Both modes enable length pruning, paragraph start / end detection, and invalid symbol masking to prevent out-of-bounds output and structural disruption. During response generation, intra-paragraph boundary cue words and entry guide words are inserted to enhance itemized expression and orderliness. An end-of-decoding checksum is appended at the end of decoding for subsequent checker comparison.

[0085] The parsing process extracts structured item identifiers from the predicted response. First, boundary checks are performed to confirm the consistency between the start and end check segments. Then, item segmentation is performed, using a set of delimiters agreed upon during training and item guide words to locate item boundaries. For each item, a format matcher is run, prioritizing matching standard position plus identifier expression; if a match fails, a fault-tolerant path is entered, attempting alias mapping, spelling correction, and approximate matching. When an item contains redundant descriptions or natural language embellishments, regular expression templates and a finite-state parser are used to retain identifier segments and discard noise. The parsed item identifier sequence is deduplicated while maintaining its original position, using the identifier's standard form as the primary key, retaining the first-appearing item, and recording removed items for auditing. If the number of parsed items is less than the required scale, gap markers and a summary of gap reasons are generated; if it exceeds the required scale, items are pruned by position, and a pruning list is generated.

[0086] The retrieval process uses the project identifier as the primary key to access project information storage. To ensure latency and consistency, read-only caching is prioritized, and in case of cache misses, the process reverts to the authoritative data source, including a version number to avoid reading inconsistent snapshots. Project information corresponding to each project identifier is assembled according to a field template. Basic fields include name, category, and key attributes, while extended fields are selectively added as needed. When an invalid or decommissioned identifier is encountered, a replacement strategy is triggered: if a new version identifier exists for the same semantic meaning, it is mapped to the new version and the replacement relationship is marked; if no new version exists, it is marked as invalid and removed from the final output, and recorded in the verification report. After retrieval, the project information set is rearranged according to the positional order retained during the parsing phase to ensure that the semantic order matches the text order.

[0087] The project list is constructed following a stable field order and unified rendering rules. Each project entry includes a project identifier, name, category, and a minimum set of necessary attributes. Field values ​​undergo anonymization and unit normalization before rendering to avoid display differences across environments. List-level metadata records input snapshot hashes, parsed and retrieved version information, gap markers, and alternative lists for easy offline review. The final output supports two equivalent media: a human-readable text list and a machine-readable structured object. Both media maintain consistency with the field mapping through the same order; the text media is used for presentation, and the structured media is used for subsequent linked processing. Consistency checks are performed before output to check the relationship between the number of entries and the scale requirements, deviations between the category distribution and the distribution requirements, and whether the order has been accidentally rewritten. After the check passes, the output is pushed to the caller, and an output summary is written to the log to support issue replay and quality statistics.

[0088] Example Explanation: In a fintech business scenario, the system first extracts historical operation sequences, user profiles, and session data from enterprise user behavior logs and business interaction databases. This data includes business processing instructions, account inquiries, wealth management product click records, approval operations, etc. By using a unified time index and entity identifier, scattered behavioral fragments are combined into a continuous interaction chain. Subsequently, the system analyzes the static attributes (such as enterprise size, industry classification, and credit rating) and dynamic characteristics (such as the rhythm of fund inflows and transaction risk preferences) in the user profile, thereby identifying the user's preference for financial product or service attributes in different business modules. For example, a preference for low-risk bond assets or cyclical short-term financing products. Based on real-time requests, the system determines the output scale and distribution conditions, such as limiting the proportion of products with multiple risk levels or geographical coverage in the output list, forming complete task input conditions.

[0089] During the generation phase, the system constructs project sequence embeddings using historical interaction records, mapping the identifier of each business project to an embedding vector and inputting it into the sequential prediction model. After long-term supervised adjustment, the model acquires the ability to identify transaction logic and sequential dependencies, calculating prediction scores through temporal feature weights to reflect the potential priority of target projects. The system sorts the projects in descending order of prediction scores and selects a batch of projects from the candidate pool as a candidate set for subsequent personalized filtering.

[0090] In the personalized processing stage, the system determines whether a task is based on risk preference constraints, return diversity constraints, or periodic distribution constraints, depending on the task type. When a task includes personalized constraints, the system eliminates candidate projects that do not meet the enterprise's risk tolerance range or credit threshold. When a task includes diversity constraints, the system resamples the candidate set using diversity indicators (such as industry coverage and product category differences) to ensure a balanced structural distribution among the output projects. For example, in a comprehensive fund management scenario, the system controls the proportion of bonds, notes, and short-term investment instruments to maintain a balance between liquidity and returns. The results after screening and resampling form the processed candidate set.

[0091] In the label construction phase, the system identifies real projects from historical interactions. These real projects can be high-success-rate recent business transactions or key deals completed by the enterprise. This project is set as the first element of the response label to stabilize the model training objective. The system then removes this project from the processed candidate set to avoid duplicate selection and determines the number of supplementary projects based on task type and output length. The system selects the top-ranked projects as supplementary entries and combines them sequentially with the first project to form the response label. This combination not only maintains the continuity of business logic but also ensures that real projects have a dominant weight in the training data, thus enabling the model to tend to retain output patterns that conform to financial business rules during prediction.

[0092] The system then uses preset templates to reconstruct historical interaction records, preference constraints, and output conditions into natural language instructions, such as "Generate an asset allocation plan for the enterprise based on recent fund allocation behavior and risk preferences," and formats response tags into natural language responses, such as "Recommendations: Short-term bonds, liquid notes, credit-enhanced wealth management products." By concatenating instructions and responses, the system generates structured training samples. Each sample contains a clear input context and target output label, with additional instruction and response labels to distinguish different semantic segments. These samples are stored as supervised adjustment data for subsequent language model parameter optimization, enabling the model to learn business logic, risk control, and diversity allocation under natural language conditions.

[0093] Once the language model is adjusted, the system receives natural language task instructions from financial professionals or corporate clients, such as "Generate a fund allocation list for the next week." The model generates a natural language response through contextual reasoning, including a set of projects sorted by risk level and liquidity. The system parses the response text, extracts structured project identifiers, and retrieves corresponding project information from the financial business database. Fields include project name, asset type, yield range, maturity date, and credit rating. After parsing, the system reconstructs the list into a standard project list and performs consistency checks to ensure the output meets the task's scale and distribution requirements. The final project list is displayed in a tabular view on the business side and can also serve as input for subsequent analysis and auditing.

[0094] Through the aforementioned processing chain, the entire system achieves a closed-loop process from business behavior data extraction, semantic conditional modeling, structured candidate generation, personalized filtering, to natural language parsing and output. In the fintech field, this method achieves dual-layer controllability on both the data and model sides: the data side ensures that the output meets business compliance and risk control requirements through constraint rules, while the model side enhances the accuracy and consistency of the output through language instruction learning and structured labels. The system can dynamically generate project lists based on different customers' liquidity, risk preferences, and business types, making the collaborative decision-making process between humans and models transparent, traceable, and highly controllable in the financial operating environment, thereby improving the reliability of task response and the quality of decision-making.

[0095] In healthcare scenarios, the system first extracts historical records, preference constraints, and output conditions from individual users' health interaction data sources. Health interaction data includes vital sign sequences collected by wearable devices, exercise logs, dietary input records, and individual health inquiry interaction text. The system establishes a unified time series using timestamps and device identifiers, and integrates multi-source health behavior fragments from different devices or applications on the same time dimension. Subsequently, the system analyzes static profile information (including age, weight, sleep structure, and disease risk factors) and dynamic session data (such as recently entered health goals, intake restrictions, and exercise frequency) to identify users' preference constraints in health behaviors, such as a preference for low-sugar diets, intermittent aerobic exercise, or high-fiber intake plans. It then combines this with real-time health requests to determine output conditions, such as limiting the number of solutions in the result set or adjusting the recommendation distribution to make the output more tailored to individual needs.

[0096] During the processing phase, the system maps event identifiers (such as intake type, exercise type, and sleep cycle labels) from historical health behavior sequences into embedding vectors, forming a continuous behavioral sequence representation. This sequence is input into a pre-trained sequential prediction model, which has learned time-dependent features from long-term health behavior data and can identify potential cyclical patterns and key turning points in an individual's health behavior. The model outputs a set of candidate health solutions, each with a prediction score reflecting its suitability to the user's current state. The system sorts the candidate pool in descending order and selects the top few as the candidate set, laying the foundation for subsequent personalized selection.

[0097] For personalized processing, the system determines whether to apply preference constraints or diversity constraints based on the task type. When the task is a health intervention task based on preference constraints, the system eliminates candidate solutions that conflict with the user's current contraindications or limitations, such as excluding high-sugar diet recommendations when blood sugar fluctuations are detected. When the task is a diversity-constrained task, the system resamples the candidate set based on diversity indicators (such as nutrient distribution and exercise category differences) to ensure a balanced distribution of output solutions in terms of type and effectiveness, thus guaranteeing the continuity and coverage of the overall health plan.

[0098] In the response tag construction phase, the system identifies genuine health plans from historical interaction records, uses them as the first item in the tag, and removes these plans from the processed candidate set to avoid duplication. The system determines the number of supplementary health plans based on the task type and preset length, and selects several of the top-ranked plans to combine with the first item to form a complete response tag. For example, in a long-term chronic disease management task, the first item might be a low-salt diet plan that the user recently successfully implemented, while supplementary items would come from a multi-day nutritional adjustment plan predicted by the model.

[0099] Subsequently, the system generates natural language instructions based on the template, integrating historical behavior summaries, preference constraints, and output conditions into a semantically clear task description, such as "Based on the blood glucose monitoring and exercise records of the past seven days, generate dietary and exercise adjustment suggestions for next week." Response tags are formatted into natural language responses, such as "It is recommended to consume a high-fiber breakfast daily, reduce refined carbohydrates, and combine with 20 minutes of aerobic exercise." Both are concatenated according to a preset format and start tags for the instructions and responses are inserted to form a structured training sample, used to supervise the model's learning of the semantic mapping relationship between instructions and results. This sample is stored in the training dataset, recording the source summary, template version, and field mapping table to ensure consistency and traceability in subsequent model adjustments.

[0100] After supervised tuning, the model receives natural language task instructions from users, such as "Please generate a health adjustment list for the next five days." The system inputs this instruction into the tuned language model, which combines internal parameters with semantic patterns from training samples to generate a predictive response in natural language. This response includes multi-dimensional health management plans, such as dietary recommendations, exercise frequency, and lifestyle optimization suggestions. The system parses the predictive response, extracts structured health plan identifiers, and then accesses the health data storage center to retrieve the corresponding plan details, including nutritional composition, exercise duration, applicable population, and constraints. After parsing, the system generates a list of items and performs consistency verification to ensure that the number and distribution of output plans meet user constraints. The final list is displayed through a visual interface, allowing users to directly view their daily adjustment plans, execution order, and data tracking entry points. The output results are simultaneously written to the user's health record for subsequent evaluation.

[0101] Through this complete processing flow, the system achieves a closed-loop end-to-end in the healthcare field, encompassing user data collection, behavioral feature modeling, task constraint parsing, candidate solution generation, and natural language understanding and structured output. This process ensures a precise correspondence between health behavior data and the language model, enabling the model to understand and execute complex health management requests. By introducing personalized constraints and diversity resampling mechanisms, the system ensures that the output not only meets medical standards and nutritional balance requirements but also dynamically adapts to changes in individual health status. The final health management output possesses semantic interpretability, structural controllability, and execution traceability, providing users with a continuous, intelligent, and highly personalized health support experience.

[0102] This embodiment establishes boundary markers, length estimates, order fidelity, and consistency checks at each stage of input, generation, parsing, retrieval, and rendering. Text generation and structured parsing form a closed loop, making the conversion from task instructions to a list of items traceable and reproducible. By standardizing item guide words and delimiters, it significantly reduces omissions and errors caused by parsing ambiguity and format drift. A retrieval and substitution strategy using item identifiers as primary keys reduces the impact of invalid items on result stability and maintains consistent positional information with the predicted response during order reordering. Dual-carrier output ensures synchronized updates for both human-readable and machine-readable presentations. Even with changes in scale and distribution constraints, it can quickly locate the source of deviation and replay the data, thereby improving the accuracy, consistency, and auditability of the output while maintaining controllable latency.

[0103] In one embodiment, a language model supervision adjustment and task output device is provided, which corresponds one-to-one with the language model supervision adjustment and task output method in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the language model supervision adjustment and task output device of the present invention. The module includes a data acquisition module 10, a candidate generation module 20, a task filtering module 30, a tag construction module 40, an instruction generation module 50, a model adjustment module 60, and an inference output module 70. Detailed descriptions of each functional module are as follows: Data acquisition module 10 is used to acquire historical interaction records, user preference constraints, and output conditions; The candidate generation module 20 is used to generate a set of candidate items based on the historical interaction records using a pre-trained sequential prediction model. The task filtering module 30 is used to filter or sample the candidate item set according to the task type to generate a processed candidate item set. The tag building module 40 is used to set real projects as the base projects of response tags, and select one or more projects from the processed candidate project set according to the task type, and combine them with the base projects to form response tags; The instruction generation module 50 is used to combine the instruction containing the historical interaction record, the user preference constraint and the output condition with the response tag to form an instruction response pair; The model adjustment module 60 is used to supervise and adjust the language model using the instruction response pair; The inference output module 70 is used to process the input task instructions through the adjusted language model and output a list of items.

[0104] In one embodiment, the data acquisition module 10 is specifically used for: Retrieve raw data, including user action sequences, static user profiles, and dynamic user session data, from user behavior logs and user attribute databases; The user operation sequence is processed, the item identifier is extracted and arranged in chronological order to form a historical interaction record; By analyzing the static user profile and the dynamic user conversation data, the user's tendency to express opinions on project categories or attributes is identified, thus forming user preference constraints. The system analyzes preset strategies or real-time user requests to determine the size and distribution requirements of the output item list, thus forming the output conditions.

[0105] In one embodiment, the candidate generation module 20 is specifically used for: The item identifiers in the historical interaction records are converted into corresponding item embedding vectors to form serialized embedding representations; The serialized embedding represents the input to a pre-trained sequential prediction model; The sequential prediction model processes the serialized embedding representation to determine a corresponding prediction score for each item in the candidate pool. The items in the candidate pool are sorted in descending order based on the predicted scores, and a preset number of items at the top of the sort are selected to form a candidate item set.

[0106] In one embodiment, the task filtering module 30 is specifically used for: The current task type is determined based on the user preference constraints and the output conditions; Set the candidate item set as the current processing set; If the task type includes a task type based on personalized preference constraints, then all items that do not conform to the user preference constraints are removed from the current processing set, and the current processing set is updated. If the task type includes a task type based on diversity constraints, then the current processing set is resampled according to the diversity index and the current processing set is updated. The updated current processing set is used as the candidate item set after processing.

[0107] In one embodiment, the tag building module 40 is specifically used for: One item is selected as the real item from the interaction items contained in the historical interaction record; Set the real item as the first item in the response tag; The real projects are removed from the processed candidate project set to form a set of candidate projects; Based on the task type and the preset response tag length, determine the number of supplementary items to be selected from the candidate item set; According to the order of the items in the processed candidate item set, select the items with the highest number of supplementary items from the candidate item set as supplementary items; The first item and the supplementary items are combined in sequence to form a response tag.

[0108] In one embodiment, the instruction generation module 50 is specifically used for: According to the preset instruction template, the historical interaction records, the user preference constraints, and the output conditions are combined into natural language instructions; Based on the preset response template, the sequence of item identifiers in the response tags is formatted into a natural language response; The natural language instructions and natural language responses are concatenated according to a preset combination format, and instruction start markers and response start markers are added to the concatenated content to form a structured instruction-response pair; The instruction response pairs are stored as training samples.

[0109] In one embodiment, the inference output module 70 is specifically used for: Receive task instructions in natural language; The task instruction is input into the adjusted language model, and the adjusted language model processes the task instruction to generate a predicted response in natural language form. The structured item identifier is parsed from the predicted response; Based on the project identifier, retrieve the corresponding project information and generate a project list; Output the list of items.

[0110] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides deterministic and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a language model supervision, adjustment, and task output method on the server side.

[0111] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of a language model supervision, adjustment, and task output method.

[0112] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain historical interaction records, user preference constraints, and output conditions; Based on the historical interaction records, a set of candidate items is generated using a pre-trained sequential prediction model; Based on the task type, the candidate item set is filtered or sampled to generate a processed candidate item set; The real project is set as the base project for the response tag, and one or more projects are selected from the processed candidate project set according to the task type, and combined with the base project to form the response tag; The instruction containing the historical interaction record, the user preference constraints, and the output conditions is combined with the response tag to form an instruction-response pair; The language model is supervised and adjusted using the instruction response pairs. The input task instructions are processed by the adjusted language model, and a list of items is output.

[0113] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain historical interaction records, user preference constraints, and output conditions; Based on the historical interaction records, a set of candidate items is generated using a pre-trained sequential prediction model; Based on the task type, the candidate item set is filtered or sampled to generate a processed candidate item set; The real project is set as the base project for the response tag, and one or more projects are selected from the processed candidate project set according to the task type, and combined with the base project to form the response tag; The instruction containing the historical interaction record, the user preference constraints, and the output conditions is combined with the response tag to form an instruction-response pair; The language model is supervised and adjusted using the instruction response pairs. The input task instructions are processed by the adjusted language model, and a list of items is output.

[0114] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0115] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0117] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

[0118] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.

Claims

1. A method for supervised adjustment and task output of a language model, characterized in that, Includes the following steps: Obtain historical interaction records, user preference constraints, and output conditions; Based on the historical interaction records, a set of candidate items is generated using a pre-trained sequential prediction model; Based on the task type, the candidate item set is filtered or sampled to generate a processed candidate item set; The real project is set as the base project for the response tag, and one or more projects are selected from the processed candidate project set according to the task type, and combined with the base project to form the response tag; The instruction containing the historical interaction record, the user preference constraints, and the output conditions is combined with the response tag to form an instruction-response pair; The language model is supervised and adjusted using the instruction response pairs. The input task instructions are processed by the adjusted language model, and a list of items is output.

2. The supervised adjustment and task output method for the language model as described in claim 1, characterized in that, Retrieve historical interaction records, user preference constraints, and output conditions, including: Retrieve raw data, including user action sequences, static user profiles, and dynamic user session data, from user behavior logs and user attribute databases; The user operation sequence is processed, the item identifier is extracted and arranged in chronological order to form a historical interaction record; By analyzing the static user profile and the dynamic user conversation data, the user's tendency to express opinions on project categories or attributes is identified, thus forming user preference constraints. The system analyzes preset strategies or real-time user requests to determine the size and distribution requirements of the output item list, thus forming the output conditions.

3. The supervised adjustment and task output method for the language model as described in claim 1, characterized in that, Based on the historical interaction records, a candidate item set is generated using a pre-trained sequential prediction model, including: The item identifiers in the historical interaction records are converted into corresponding item embedding vectors to form serialized embedding representations; The serialized embedding represents the input to a pre-trained sequential prediction model; The sequential prediction model processes the serialized embedding representation to determine a corresponding prediction score for each item in the candidate pool. The items in the candidate pool are sorted in descending order based on the predicted scores, and a preset number of items at the top of the sort are selected to form a candidate item set.

4. The supervised adjustment and task output method for the language model as described in claim 1, characterized in that, Based on the task type, the candidate item set is filtered or sampled to generate a processed candidate item set, including: The current task type is determined based on the user preference constraints and the output conditions; Set the candidate item set as the current processing set; If the task type includes a task type based on personalized preference constraints, then all items that do not conform to the user preference constraints are removed from the current processing set, and the current processing set is updated. If the task type includes a task type based on diversity constraints, then the current processing set is resampled according to the diversity index and the current processing set is updated. The updated current processing set is used as the candidate item set after processing.

5. The supervised adjustment and task output method for a language model as described in claim 1, characterized in that, The real project is set as the base project for the response tag, and one or more projects are selected from the processed candidate project set according to the task type, and combined with the base project to form a response tag, including: One item is selected as the real item from the interaction items contained in the historical interaction record; Set the real item as the first item in the response tag; The real projects are removed from the processed candidate project set to form a set of candidate projects; Based on the task type and the preset response tag length, determine the number of supplementary items to be selected from the candidate item set; According to the order of the items in the processed candidate item set, select the items with the highest number of supplementary items from the candidate item set as supplementary items; The first item and the supplementary items are combined in sequence to form a response tag.

6. The supervised adjustment and task output method for a language model as described in claim 1, characterized in that, The instruction containing the historical interaction record, the user preference constraints, and the output conditions is combined with the response tag to form an instruction-response pair, including: According to the preset instruction template, the historical interaction records, the user preference constraints, and the output conditions are combined into natural language instructions; Based on the preset response template, the sequence of item identifiers in the response tags is formatted into a natural language response; The natural language instructions and natural language responses are concatenated according to a preset combination format, and instruction start markers and response start markers are added to the concatenated content to form a structured instruction-response pair; The instruction response pairs are stored as training samples.

7. The supervised adjustment and task output method for a language model as described in claim 1, characterized in that, The input task instructions are processed by the adjusted language model, and a list of items is output, including: Receive task instructions in natural language; The task instruction is input into the adjusted language model, and the adjusted language model processes the task instruction to generate a predicted response in natural language form. The structured item identifier is parsed from the predicted response; Based on the project identifier, retrieve the corresponding project information and generate a project list; Output the list of items.

8. A supervised adjustment and task output device for a language model, characterized in that, The language model's supervised tuning and task output device includes: The data acquisition module is used to obtain historical interaction records, user preference constraints, and output conditions; The candidate generation module is used to generate a set of candidate items based on the historical interaction records using a pre-trained sequential prediction model. The task filtering module is used to filter or sample the candidate item set according to the task type to generate a processed candidate item set. The tag building module is used to set real projects as the base projects for response tags, and select one or more projects from the processed candidate project set according to the task type, and combine them with the base projects to form response tags; The instruction generation module is used to combine the instruction containing the historical interaction record, the user preference constraints, and the output conditions with the response tag to form an instruction-response pair; The model adjustment module is used to supervise and adjust the language model using the instruction response pairs; The inference output module is used to process the input task instructions through the adjusted language model and output a list of items.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a language model supervision and task output program stored in the memory and executable on the processor. When the language model supervision and task output program is executed by the processor, it implements the steps of the language model supervision and task output method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a language model supervision adjustment and task output program, which, when executed by a processor, implements the steps of the language model supervision adjustment and task output method as described in any one of claims 1-7.