Merchant operation model training method and device, electronic equipment and storage medium
By acquiring merchant operation task data for semantic parsing and reinforcement learning training, a target merchant operation model is generated, which solves the problem of insufficient understanding of general inference models in vertical domain tasks and realizes efficient transfer and optimization of the model in e-commerce scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-15
AI Technical Summary
General reasoning models lack an understanding of industry rules, business logic, and semantic conventions in vertical domain tasks, resulting in reasoning results that do not conform to the actual situation in the domain, making it difficult to provide users with reliable decision analysis services.
By acquiring merchant operation task data, performing semantic parsing and task type determination, generating triplet samples, and conducting reinforcement learning training, a target merchant operation model is generated, improving the model's inference accuracy and scenario adaptability in e-commerce scenarios.
It enables efficient migration and optimization of merchant operation models in e-commerce scenarios, improves the inference accuracy and scenario adaptability of the model in vertical tasks, and avoids the high cost and subjective bias of manual annotation.
Smart Images

Figure CN121456487B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of model training technology, and in particular to a method, apparatus, electronic device and storage medium for training a business operation model. Background Technology
[0002] In related technologies, general-purpose reasoning models perform well in tasks such as everyday logical reasoning, mathematical problems, and code generation. However, in specific vertical domains (such as revenue prediction, business strategy selection, and security auditing for food delivery platforms), the reasoning results often fail to reflect the realities of the domain due to a lack of understanding of industry rules, business logic, and semantic conventions. This results in weak reasoning capabilities for vertical domain tasks, making it difficult to provide users with reliable decision analysis services. Summary of the Invention
[0003] This application provides a method, apparatus, electronic device, and storage medium for training a merchant operation model, which can improve the accuracy and scenario adaptability of the merchant operation model for operational tasks in e-commerce scenarios. The above technical solution is as follows:
[0004] In a first aspect, embodiments of this application provide a method for training a merchant operation model, the method comprising:
[0005] Acquire an initial model and multiple merchant operation task data. The merchant operation task data includes operation questions and business data related to the operation questions. The multiple merchant operation task data includes merchant operation task data of the first question and answer task type and merchant operation task data of the second question and answer task type. The first question and answer task type is used to represent merchant operation questions and answers with reference answers, and the second question and answer task type is used to represent open-ended merchant operation questions and answers without reference answers.
[0006] The initial model described above is used to perform task execution processing on the operational task data of the multiple merchants mentioned above, and the operational prediction results corresponding to each merchant's operational task data in the multiple merchant operational task data are obtained.
[0007] The reward score for the corresponding operational prediction result is determined based on the task type of the above-mentioned merchant operational task data.
[0008] Based on the above merchant operation task data, the corresponding operation prediction results, and the reward scores of the above operation prediction results, a corresponding triplet sample is generated.
[0009] The triplet samples corresponding to the above-mentioned merchant operation task data are summarized to construct an operation training sample set;
[0010] Based on the aforementioned operational training sample set, the initial model is trained using reinforcement learning to generate the target merchant operational model.
[0011] In one possible implementation, the above-mentioned initial model is used to perform task execution processing on the above-mentioned multiple merchant operation task data to obtain the operation prediction results corresponding to each merchant operation task data in the above-mentioned multiple merchant operation task data, including:
[0012] Semantic analysis was performed on the operational task data of each merchant to obtain the task type of the operational task data of each merchant.
[0013] Obtain the corresponding preset paradigm information based on the above task types;
[0014] The above-mentioned merchant operation task data and the above-mentioned preset paradigm information are processed by field decoupling and semantic concatenation to generate prompt information;
[0015] The above prompts and merchant operation task data are fused and encoded, then input into the above initial model to trigger the initial model to perform inference operations and output the corresponding operation prediction results.
[0016] In one possible implementation, the reward score for determining the operational prediction result corresponding to each merchant's operational task data based on the task type of each merchant's operational task data includes:
[0017] For each merchant operation task data with the corresponding task type of the first question and answer task mentioned above, obtain the reference answer corresponding to the merchant operation task data;
[0018] Semantic relevance detection is performed on the above operational prediction results and the above reference answers to generate relevance detection results for the operational prediction results and the above reference answers;
[0019] The above correlation detection results are quantified to obtain the corresponding semantic consistency evaluation index;
[0020] Based on the aforementioned semantic consistency evaluation index, the reward score for the operational forecast result is determined from the preset reward value range.
[0021] In one possible implementation, the reward score for determining the operational prediction result corresponding to each merchant's operational task data based on the task type of each merchant's operational task data includes:
[0022] For each merchant operation task data corresponding to the above-mentioned second question and answer task type, obtain the preset reward model corresponding to the above-mentioned second question and answer task type.
[0023] After feature encoding processing of the aforementioned merchant operation task data and the corresponding aforementioned operation prediction results, the data is input into the aforementioned preset reward model. The aforementioned preset reward model is then used to generate reward scores for the aforementioned merchant operation task data and the corresponding aforementioned operation prediction results, resulting in the reward score of the aforementioned operation prediction results output by the aforementioned preset reward model.
[0024] In one possible implementation, the task types of the aforementioned multiple merchant operation task data also include at least one of the following: calculating task type, selecting task type, and determining task type.
[0025] In one possible implementation, the reward score for determining the operational prediction result corresponding to each merchant's operational task data based on the task type of each merchant's operational task data includes:
[0026] If the task type of the above-mentioned merchant operation task data is any one of the above-mentioned calculation task type, the above-mentioned selection task type, and the above-mentioned judgment task type, obtain the preset reference execution result of the above-mentioned merchant operation task data.
[0027] The operational prediction results corresponding to the above merchant operational task data are matched with the above preset reference execution results to obtain the matching quantification results;
[0028] Based on the above matching metric results, construct the corresponding execution consistency evaluation index;
[0029] Based on the aforementioned consistency evaluation indicators, the reward score for the above operational forecast results is determined from the preset reward value range.
[0030] In one possible implementation, the initial model is trained using reinforcement learning based on the aforementioned operational training sample set to generate a target merchant operational model, including:
[0031] Obtain a preset general task sample set and the first training weights corresponding to the general task sample set, and obtain the second training weights corresponding to the operation training sample set.
[0032] Based on the first and second training weights mentioned above, the general task sample set and the operational training sample set are jointly arranged to construct a joint training sample set containing multi-source task information.
[0033] Construct a loss function based on the aforementioned joint training sample set;
[0034] The initial model is subjected to gradient backpropagation and parameter iterative update using the loss function described above to generate the target merchant operation model.
[0035] Secondly, embodiments of this application provide a merchant operation model reasoning method, including:
[0036] Acquire pending merchant operation task data, which includes pending operation questions and related business data.
[0037] Input the above-mentioned merchant operation task data into the target merchant operation model to obtain the target operation prediction results output by the target merchant operation model.
[0038] The aforementioned target merchant operation model is generated by training a model using a merchant operation model training method as described in the first aspect or any possible implementation provided by the first aspect.
[0039] Thirdly, embodiments of this application provide a merchant operation model training device, the device comprising:
[0040] The first acquisition module is used to acquire the initial model and multiple merchant operation task data. The merchant operation task data includes operation questions and business data related to the operation questions. The multiple merchant operation task data includes merchant operation task data of the first question and answer task type and merchant operation task data of the second question and answer task type. The first question and answer task type is used to represent merchant operation question and answer questions with reference answers, and the second question and answer task type is used to represent open-ended merchant operation question and answer questions without reference answers.
[0041] The processing module is used to perform task execution processing on the above-mentioned multiple merchant operation task data through the above-mentioned initial model, and to obtain the operation prediction results corresponding to each merchant operation task data in the above-mentioned multiple merchant operation task data.
[0042] The determination module is used to determine the reward score of the operation prediction result corresponding to the above-mentioned merchant operation task data based on the task type of the above-mentioned merchant operation task data;
[0043] The generation module is used to generate corresponding triplet samples based on the above-mentioned merchant operation task data, the corresponding operation prediction results, and the reward scores of the above-mentioned operation prediction results.
[0044] The module is used to summarize multiple triplet samples corresponding to the above-mentioned multiple merchant operation task data and build an operation training sample set;
[0045] The training module is used to perform reinforcement learning training on the initial model based on the above-mentioned operational training sample set to generate the target merchant's operational model.
[0046] Fourthly, embodiments of this application provide a merchant operation model inference device, the device comprising:
[0047] The second acquisition module is used to acquire merchant operation task data to be processed. The merchant operation task data includes operation questions to be answered and business data related to the operation questions to be answered.
[0048] The input module is used to input the above-mentioned merchant operation task data to the target merchant operation model and obtain the target operation prediction results output by the target merchant operation model.
[0049] The aforementioned target merchant operation model is generated by training a model using a merchant operation model training method as described in the first aspect or any possible implementation provided by the first aspect.
[0050] Fifthly, embodiments of this application provide an electronic device, including: a processor and a memory;
[0051] The aforementioned memory stores a computer program adapted to be loaded by the aforementioned processor and execute the steps of the method provided by the first aspect of the embodiments of this application or any possible implementation thereof.
[0052] Sixthly, embodiments of this application provide a computer storage medium storing a plurality of instructions adapted for a processor to load and execute the steps of the method provided by the first aspect of the embodiments of this application or any possible implementation thereof.
[0053] In this embodiment, an initial model and multiple merchant operation task data are obtained. The merchant operation task data includes operational questions and related business data. The multiple merchant operation task data includes merchant operation task data of a first question-and-answer task type and merchant operation task data of a second question-and-answer task type. The first question-and-answer task type represents merchant operation questions with reference answers, and the second question-and-answer task type represents open-ended merchant operation questions without reference answers. The initial model is used to perform task execution processing on the multiple merchant operation task data to obtain the operation prediction results corresponding to each merchant operation task data. The reward score for the operation prediction results corresponding to each merchant operation task data is determined based on the task type of each merchant operation task data. Then, for each merchant operation task data, corresponding triplet samples are generated based on the merchant operation task data, the corresponding operation prediction results, and the reward scores of the operation prediction results. Multiple triplet samples corresponding to the multiple merchant operation task data are summarized to construct an operation training sample set. Finally, the initial model is trained using reinforcement learning based on the operation training sample set to generate a target merchant operation model. Therefore, by designing corresponding task execution and reward generation mechanisms for different task types, such as merchant operation Q&A questions with reference answers and open-ended merchant operation Q&A questions without reference answers, the correctness evaluation and reward score determination of the model output results can be automated, avoiding the high cost and subjective bias of manual annotation. In addition, by organizing the task input, prediction results and reward scores into triple samples and constructing them into a training sample set, and by optimizing the initial model based on reinforcement learning, the model can retain its original general reasoning ability while acquiring data understanding, logical reasoning and business decision-making capabilities for vertical domain tasks. This significantly improves the model's reasoning accuracy and scenario adaptability in vertical domain scenarios, achieving efficient migration and optimization from a general language model to a vertical domain intelligent model. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 A schematic diagram illustrating the application environment of a merchant operation model training method provided as an exemplary embodiment of this application;
[0056] Figure 2A flowchart illustrating a merchant operation model training method provided as an exemplary embodiment of this application;
[0057] Figure 3 A flowchart illustrating another merchant operation model training method provided as an exemplary embodiment of this application;
[0058] Figure 4 A flowchart illustrating another merchant operation model training method provided as an exemplary embodiment of this application;
[0059] Figure 5 A flowchart illustrating another merchant operation model training method provided as an exemplary embodiment of this application;
[0060] Figure 6 A flowchart illustrating a merchant operation model reasoning method provided as an exemplary embodiment of this application;
[0061] Figure 7 A schematic diagram of the structure of a merchant operation model training device provided as an exemplary embodiment of this application;
[0062] Figure 8 A schematic diagram of the structure of a merchant operation model inference device provided as an exemplary embodiment of this application;
[0063] Figure 9 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation
[0064] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0065] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus. The term "at least one" in this application means one or more, and the term "multiple" means two or more; for example, multiple second devices means two or more second devices. The terms "system" and "network" are often used interchangeably herein.
[0066] First, let's introduce the application scenario applicable to this embodiment: An e-commerce platform provides intelligent decision analysis services to its member stores through an e-commerce decision analysis assistant. This e-commerce decision analysis assistant can be driven by a trained target merchant operation model. This target merchant operation model can perform intelligent reasoning and decision analysis on relevant issues in the e-commerce operation process based on multi-source business data provided by the merchant (including product information, marketing activity data, user behavior data, and historical transaction data), including but not limited to activity revenue prediction, product pricing optimization, traffic allocation analysis, marketing strategy generation, and campaign effectiveness evaluation. Through the reasoning results of this target merchant operation model, the e-commerce platform can provide merchants with actionable business suggestions and strategy adjustment plans, thereby improving merchant operating efficiency and achieving intelligent e-commerce operation decision support.
[0067] The merchant operation model training method provided in this application embodiment can be applied to, for example, Figure 1 The application environment shown includes a user terminal 10, a server 20, a communication network 30, and a data storage system 40. The user terminal 10, with an online application (App) installed, communicates with the server 20 corresponding to the online App via the communication network 30. The data storage system 40 stores data that the server 20 needs to process, such as relevant business data. The data storage system 40 can be integrated onto the server 20 or located on a cloud or other network server.
[0068] In some possible embodiments, server 20 acquires an initial model and multiple merchant operation task data, including operation questions and related business data. The multiple merchant operation task data includes merchant operation task data of a first question-and-answer task type and merchant operation task data of a second question-and-answer task type. The first question-and-answer task type represents merchant operation questions with reference answers, and the second question-and-answer task type represents open-ended merchant operation questions without reference answers. The initial model performs task execution processing on the multiple merchant operation task data to obtain the operation prediction result corresponding to each merchant operation task data. Based on the task type of each merchant operation task data, the reward score of the corresponding operation prediction result is determined. For each merchant operation task data, corresponding triplet samples are generated based on the merchant operation task data, the corresponding operation prediction result, and the reward score of the operation prediction result. The multiple triplet samples corresponding to the multiple merchant operation task data are summarized to construct an operation training sample set. The initial model is trained using reinforcement learning based on the operation training sample set to generate a target merchant operation model. Furthermore, the server 20 can obtain pending merchant operation task data from the user terminal 10. The aforementioned merchant operation task data includes pending operation questions and related business data. The server 20 inputs the pending merchant operation task data into the target merchant operation model to obtain the target operation prediction result output by the target merchant operation model. The target operation prediction result can then be returned to the user terminal 10 so that the user terminal 10 can display the target operation prediction result to the merchant user.
[0069] Understandably, the aforementioned online apps can be, but are not limited to, e-commerce platform applications (such as shopping apps, food delivery apps, etc.), social e-commerce applications (which can combine content formats such as images, short videos, or live streaming to promote product sales through user interaction, sharing, comments, and recommendations), independent merchant store management applications (online store management and transaction platforms built or hosted by individual merchants), and third-party platform applications that provide marketing, promotion, placement, pricing, or data analysis services for merchants (e-commerce service platform systems for multiple merchants). These apps can intelligently analyze and reason about merchant operation problems by calling the merchant operation model deployed on server 20. User terminal 10 can be, but is not limited to, various smartphones, tablets, personal computers, laptops, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. User terminal 10 can be implemented as a single terminal device or a terminal cluster composed of multiple terminals. Server 20 can be implemented as a separate server or a server cluster composed of multiple servers.
[0070] In one embodiment, such as Figure 2 As shown, a method for training a merchant operation model is provided. This method can be applied to the aforementioned server and includes the following steps:
[0071] S201: Obtain the initial model and multiple merchant operation task data. The merchant operation task data includes operation questions and business data related to the operation questions. The multiple merchant operation task data includes merchant operation task data of the first question-and-answer task type and merchant operation task data of the second question-and-answer task type. The first question-and-answer task type is used to represent merchant operation questions and answers with reference answers, and the second question-and-answer task type is used to represent open-ended merchant operation questions and answers without reference answers.
[0072] The initial model can be a basic large language model used for intelligent training of merchant operations. This initial model can possess general language understanding, reasoning, and generation capabilities. It can serve as the training starting point for this embodiment of the application, and be optimized through subsequent reinforcement learning processes combined with e-commerce operation task data, thereby gradually acquiring reasoning and decision-making capabilities tailored to merchant operation scenarios.
[0073] Optionally, each merchant's operational task data may include an operational problem to be solved. This operational problem can be a specific task that the merchant user wants the model to address, such as predicting campaign revenue or analyzing which discount scheme has a higher conversion rate. The business data related to the operational problem can be contextual information supporting the analysis of that problem, such as product management data (e.g., product price, inventory quantity, category information, listing time, product ratings), store management data (e.g., store level, number of followers, main category, average daily visitors, store conversion rate), marketing campaign data (e.g., campaign type, discount level, budget allocation, campaign period, ad exposure), user behavior data (e.g., user browsing history, click behavior, add-to-cart history, purchase frequency, repurchase rate), transaction conversion data (e.g., order quantity, transaction amount, average order value, return rate, payment success rate), and business strategy data (e.g., pricing strategy, promotion plan, inventory management plan, advertising strategy, traffic allocation plan).
[0074] Optionally, the "calculation task type" is used to characterize the operational questions included in the corresponding merchant operation task data, which have a unique numerical result; that is, the operational question is a calculation question. For example, the operational question included in the merchant operation task data of the "calculation task type" could be: Calculate the average order value based on the following transaction data. For example, the "selection task type" is used to characterize the operational questions included in the corresponding merchant operation task data, which have multiple options and only one correct answer; that is, the operational question is a multiple-choice question. For example, the operational question included in the merchant operation task data of the "selection task type" could be: "Which of the following discount methods can bring a higher click-through rate? A. Full reduction B. Flash sale C. Discount coupon." The "Judgment Task Type" indicates that the corresponding merchant operation task data includes operational questions with binary outputs (True / False, Yes / No), meaning the operational question is a true / false question. For example, the operational question included in the merchant operation task data of the Judgment Task Type could be "Will extending the activity duration definitely improve the conversion rate, true or false?" The "First Question-Answer Task Type" indicates that the corresponding merchant operation task data includes operational questions with a reference answer, meaning the operational question is a question-and-answer question with a reference answer. For example, the operational question included in the merchant operation task data of the First Question-Answer Task Type could be "Please explain why the conversion rate of this product has decreased? (Reference answer: Affected by a recent reduction in advertising budget)." The "Second Question-Answer Task Type" indicates that the corresponding merchant operation task data includes operational questions without a reference answer. Specifically, this could be a strategy generation or creative question, meaning the operational question is an open-ended question. For example, the operational question included in the merchant operation task data of the Second Question-Answer Task Type could be "Please generate a new promotional copy for the activity to increase product exposure."
[0075] In one embodiment, a process of collecting historical operational data, generating problem templates, matching problem-data, and verifying quality can be automatically constructed to form multiple merchant operation task data covering e-commerce merchant operation scenarios, which can then be used to drive reinforcement learning training of the target merchant operation model.
[0076] S202: Using the initial model described above, perform task execution processing on the above multiple merchant operation task data to obtain the operation prediction results corresponding to each merchant operation task data in the above multiple merchant operation task data.
[0077] In this context, operational forecast results refer to the output generated by the initial model based on the input merchant operational task data. These results characterize the initial model's reasoning, predictions, or decision-making conclusions regarding operational issues based on business data. For example, operational forecast results may include, but are not limited to, predicted revenue figures, recommended pricing strategies, generated marketing plans, or optimization suggestions.
[0078] S203: Based on the task type of the above-mentioned merchant operation task data, determine the reward score of the corresponding operation prediction result of the above-mentioned merchant operation task data.
[0079] Optionally, the reward score can be a quantitative indicator used to measure the accuracy or quality of operational forecast results. For operational questions with a single standard answer (calculation questions, multiple choice questions, true / false questions), the reward score can be determined by string matching; for open-ended questions with reference answers, a semantic consistency score can be calculated using a language model discriminator; for open-ended questions, a continuous reward score can be calculated based on the quality of the answer using a relevant reward model.
[0080] S204: For the above-mentioned merchant operation task data, generate corresponding triplet samples based on the above-mentioned merchant operation task data, the corresponding operation prediction results, and the reward scores of the above-mentioned operation prediction results.
[0081] The triplet sample can be a data structure built in the training of the merchant operation model for reinforcement learning optimization, and can include three parts: input, answer and supervision signal.
[0082] Specifically, for each merchant's operational task data, the merchant's operational task data can be used as input, the operational prediction result corresponding to the above merchant's operational task data can be used as the answer, and the reward score of the operational prediction result can be used as a supervision signal to generate a triplet sample, thereby obtaining multiple triplet samples corresponding to multiple merchant operational task data.
[0083] S205: Summarize the multiple triplet samples corresponding to the above multiple merchant operation task data to construct an operation training sample set.
[0084] In one embodiment, the operational training sample set may include multiple triplet samples such as [merchant operational task data, operational prediction results, and reward scores], used for reinforcement learning training of the initial model. That is, each operational training sample may contain: operational questions and business data (input), operational prediction results output by the model (answer), and corresponding reward scores (supervision signals).
[0085] In one embodiment, multiple triplet samples corresponding to different task types can be classified, formatted, and weighted before being merged to generate an operational training sample set for reinforcement learning optimization.
[0086] S206: Based on the above operational training sample set, perform reinforcement learning training on the above initial model to generate the target merchant operational model.
[0087] Among them, the target merchant operation model is a model optimized by reinforcement learning based on the above-mentioned operation training sample set. This target merchant operation model can have intelligent analysis and decision-making capabilities in the field of e-commerce merchant operation. It can automatically reason, predict and generate strategies for different business problems. In e-commerce platforms, it can serve as a decision analysis assistant and provide merchants with data-driven business optimization solutions.
[0088] This application embodiment acquires an initial model and multiple merchant operation task data. The merchant operation task data includes operational questions and related business data. The multiple merchant operation task data includes merchant operation task data of a first question-and-answer task type and merchant operation task data of a second question-and-answer task type. The first question-and-answer task type represents merchant operation questions with reference answers, and the second question-and-answer task type represents open-ended merchant operation questions without reference answers. The initial model performs task execution processing on the multiple merchant operation task data to obtain the operation prediction result corresponding to each merchant operation task data. Based on the task type of each merchant operation task data, the reward score for the corresponding operation prediction result is determined. For each merchant operation task data, corresponding triplet samples are generated based on the merchant operation task data, the corresponding operation prediction result, and the reward score of the operation prediction result. Multiple triplet samples corresponding to the multiple merchant operation task data are summarized to construct an operation training sample set. Finally, the initial model is trained using reinforcement learning based on the operation training sample set to generate a target merchant operation model. Therefore, by designing corresponding task execution and reward generation mechanisms for different task types, such as merchant operation Q&A questions with reference answers and open-ended merchant operation Q&A questions without reference answers, the correctness evaluation and reward score determination of the model output results can be automated, avoiding the high cost and subjective bias of manual annotation. In addition, by organizing the task input, prediction results and reward scores into triple samples and constructing them into a training sample set, and by optimizing the initial model based on reinforcement learning, the model can retain its original general reasoning ability while acquiring data understanding, logical reasoning and business decision-making capabilities for vertical domain tasks. This significantly improves the model's reasoning accuracy and scenario adaptability in vertical domain scenarios, achieving efficient migration and optimization from a general language model to a vertical domain intelligent model.
[0089] In one embodiment, such as Figure 3 As shown, another method for training a merchant operation model is provided, including the following steps:
[0090] S301: Obtain the initial model and multiple merchant operation task data, which include operation issues and business data related to the operation issues.
[0091] Specifically, S301 is the same as S201, and will not be repeated here.
[0092] S302: Perform semantic parsing on the above-mentioned merchant operation task data to determine the task type of the above-mentioned merchant operation task data.
[0093] In one embodiment, semantic parsing of the aforementioned merchant operation task data may include: first, performing text preprocessing on the operation questions and business data to extract semantic feature information; then, identifying the semantic structure of the operation questions based on a preset task type discrimination model, wherein the task type discrimination model can classify the combined semantics of the operation questions and business data according to features such as whether the operation questions contain numerical solution intent, option information, logical judgment terms, reference answer fields, or open-ended descriptions; furthermore, the initial type results output by the model can be corrected by combining preset discrimination rules, for example, classifying the task as a first question-and-answer task type when a unique standard answer is detected, and classifying the task as a second question-and-answer task type when there is no reference answer and the question semantics are open-ended. Through the above methods, the task type corresponding to each merchant operation task data can be determined.
[0094] S303: Obtain the corresponding preset paradigm information based on the above task types.
[0095] Among them, the preset paradigm information can refer to the standardized input templates and reasoning format specifications defined for different task types, which can make the initial model's thinking path and output format conform to the reasoning paradigm of the original language model when performing merchant operation tasks.
[0096] For example, when the task type is a calculation task or a selection task, the corresponding preset paradigm information includes a template to guide the model to reason step by step and output a unique standard answer. For instance, adding the instruction "Let's reason step by step and place the final answer in \(\boxed{}\)" to the input allows the initial model to explicitly expand its thought process and generate a verifiable final result during reasoning. When the task type is a judgment task, the preset paradigm information constrains the model output to be "correct" or "incorrect". When the task type is a first-level question-and-answer task (with a reference answer), the preset paradigm information can prompt the initial model to perform semantic consistency analysis based on the reference answer. When the task type is a second-level question-and-answer task (without a reference answer), the preset paradigm information can guide the initial model to generate strategic answers in an open semantic space. Thus, by setting preset paradigm information for different task types, the model can migrate to merchant operation scenarios while maintaining the consistency of its original reasoning logic, achieving a structured understanding and step-by-step reasoning of e-commerce tasks.
[0097] S304: Perform field decoupling and semantic concatenation processing on the above-mentioned merchant operation task data and the above-mentioned preset paradigm information to generate prompt information.
[0098] The prompts can be natural language task instructions that can be directly input into the initial model, and can be used to combine merchant operation task data with inference paradigms to form an input context that the model can understand.
[0099] Specifically, the prompt information may include: a task description section, which is the operational question itself, such as "Please predict the expected revenue of this promotion based on the following activity data"; a reasoning guidance section, which is determined by preset paradigm information, such as "Let's reason step by step and put the final answer in \(\boxed{}\); and a contextual data section, which is business data related to the question (products, users, marketing activities, etc.).
[0100] In one embodiment, the field decoupling and semantic concatenation processing of the aforementioned merchant operation task data and the aforementioned preset paradigm information may include: First, performing structured parsing on the merchant operation task data, mapping the operational questions, product information, activity information, user behavior data, transaction data, etc., in the merchant operation task data to preset field slots respectively; then, obtaining the corresponding preset paradigm information according to the task type, and extracting template fragments used to describe the task intent, reasoning method, and output constraints, such as "Please analyze and give the final conclusion step by step based on the following business data" or "Let's reason step by step and place the final answer in (\boxed{})", etc.; then, filling the operational question text into the question slot in the paradigm template, embedding the business data related to the question into the context data slot in the order of the fields, and performing sequential concatenation and semantic connection processing on the content of each field through natural language connectors; thereby generating prompt information that conforms to the preset reasoning paradigm and contains complete context information.
[0101] S305: Input the above prompt information and the above merchant operation task data into the above initial model to obtain the operation prediction results corresponding to the above merchant operation task data output by the above initial model.
[0102] Furthermore, after receiving prompts, the initial model can proceed with step-by-step reasoning, calculation, or judgment according to the prompts, thereby ensuring consistent thinking patterns and unified output formats across different task types.
[0103] In one embodiment, the operational forecast result can refer to the task execution output generated by the initial model after receiving prompts, based on step-by-step reasoning about the operational problem and related business data. This result can be used to characterize the model's final reasoning conclusions, calculation results, or decision recommendations regarding the operational problem.
[0104] For example, for calculation tasks, the corresponding operational forecast results may include numerical results; for selection tasks, the corresponding operational forecast results may include a single option (such as "the answer is option C"); for judgment tasks, the corresponding operational forecast results may include binary results (such as "conclusion: correct"); and for question-and-answer tasks, the corresponding operational forecast results may include natural language answers (such as "it is recommended to increase the discount ratio in this event").
[0105] S306: Based on the task type of the above-mentioned merchant operation task data, determine the reward score for the corresponding operation prediction result of the above-mentioned merchant operation task data.
[0106] Specifically, S306 is the same as S203, and will not be repeated here.
[0107] S307: For the above-mentioned merchant operation task data, generate corresponding triplet samples based on the above-mentioned merchant operation task data, the corresponding operation prediction results, and the reward scores of the above-mentioned operation prediction results.
[0108] Specifically, S307 is the same as S204, and will not be repeated here.
[0109] S308: Summarize the multiple triplet samples corresponding to the above multiple merchant operation task data to construct an operation training sample set.
[0110] Specifically, S308 is the same as S205, and will not be repeated here.
[0111] S309: Based on the above operational training sample set, perform reinforcement learning training on the above initial model to generate the target merchant operational model.
[0112] Specifically, S309 is the same as S206, and will not be repeated here.
[0113] In this embodiment, by classifying merchant operation task data into task types and configuring preset paradigm information for different tasks, the model can automatically switch inference modes according to task semantics, ensuring logical consistency and unified output across different tasks. By combining operational issues, business data, and paradigm information to generate natural language prompts, the model input becomes more structured and controllable, enabling it to accurately understand task semantics and generate results along preset inference paths. Furthermore, by embedding inference guidance text in templates, the model can maintain progressive inference capabilities in e-commerce scenarios, generating interpretable and verifiable results. Further, by calculating reward scores based on different task results and constructing an operational training sample set for reinforcement learning optimization, the trained target merchant operation model possesses stronger capabilities in business problem analysis, revenue prediction, and strategy generation. This achieves efficient transfer and precise adaptation of the language model to merchant operation decision-making scenarios while maintaining general capabilities.
[0114] In one embodiment, such as Figure 4 As shown, another method for training a merchant operation model is provided, including the following steps:
[0115] S401: Obtain the initial model and multiple merchant operation task data, which include operation issues and business data related to the operation issues.
[0116] Specifically, S401 is the same as S201, and will not be repeated here.
[0117] S402: Using the initial model described above, perform task execution processing on the above multiple merchant operation task data to obtain the operation prediction results corresponding to each merchant operation task data in the above multiple merchant operation task data.
[0118] Specifically, S402 is the same as S202, and will not be repeated here.
[0119] S403: If the task type of the above-mentioned merchant operation task data is any one of the above-mentioned calculation task type, the above-mentioned selection task type, and the above-mentioned judgment task type, obtain the preset reference execution result of the above-mentioned merchant operation task data.
[0120] Among them, the preset reference execution result can refer to the standardized correct result that is stored in advance for task types (calculation questions, multiple choice questions, true / false questions) with a unique standard answer, and can be used as a comparison benchmark for the model output.
[0121] Specifically, when the task type is a calculation task, the preset reference execution result can be a pre-calculated standard numerical result; when the task type is a selection task, the preset reference execution result can be the only correct option; when the task type is a judgment task, the preset reference execution result can be "correct" or "incorrect". The preset reference execution result can be used as a standard answer during the training phase to evaluate the correctness or consistency of the model's output.
[0122] In one embodiment, the preset reference execution result can be generated manually or by an algorithm based on the question stem logic and business rules beforehand.
[0123] S404: Match the operational prediction results corresponding to the above merchant operation task data with the above preset reference execution results to obtain the matching quantification results.
[0124] Optionally, the matching quantification results can represent the degree of consistency between the operational forecast results and the preset reference execution results, and can be used to transform the judgment of "whether it is correct" into a quantifiable matching indicator.
[0125] In one embodiment, when the task type corresponds to a numerical result, the numerical error comparison method can be used to calculate the quantified matching result; when the task type is a multiple-choice question or a true / false question, Boolean judgment, character consistency comparison, or string full matching can be used to obtain the matching label; furthermore, in order to avoid misjudgment caused by format differences, the prediction result can also be standardized, such as removing spaces, unifying capitalization, and standardizing symbol format, so as to make the matching result more stable and reliable.
[0126] S405: Construct the corresponding execution consistency evaluation index based on the above matching metric results.
[0127] Among them, the consistency evaluation index can be used to quantitatively describe the degree of conformity between the operational forecast results and the preset reference execution results, so as to more finely distinguish the quality level of different outputs in subsequent reward calculations.
[0128] Specifically, when the operational forecast result corresponding to the merchant's operational task data is strictly consistent with the above-mentioned preset reference execution result, the execution consistency evaluation index can be set to 1; when the operational forecast result corresponding to the merchant's operational task data is inconsistent with the above-mentioned preset reference execution result, the execution consistency evaluation index can be set to 0.
[0129] Furthermore, for computational tasks that may involve floating-point errors, continuous evaluation values between 0 and 1 can be obtained by mapping the error magnitude to an interval; for judgment-type tasks, binary consistency labels can be directly generated. This enables a structured evaluation of the accuracy of prediction outputs under different task types.
[0130] For example, the first preset reward value can be 1, that is, when the answer is correct, the matching result indicates that the above-mentioned operational prediction result is consistent with the above-mentioned preset reference execution result, and the reward score can be 1 point.
[0131] S406: Based on the above-mentioned consistency evaluation indicators, determine the reward score of the above-mentioned operational forecast results from the preset reward value range.
[0132] The preset reward value range can be the score range used for reinforcement learning training, constraining the maximum and minimum values of the reward signal. For example, the reward range can be preset to [0,1] or [-1,1] to meet the training requirements of different types of model frameworks.
[0133] In one embodiment, when the consistency evaluation index is 1, a high reward value, such as 1, can be selected from the reward interval; when the consistency evaluation index is 0, a low reward or penalty score, such as 0, can be selected from the reward interval. Furthermore, for an embodiment that introduces continuous consistency evaluation, the consistency evaluation index can be multiplied by the maximum value of the reward interval to obtain a continuous reward score, thereby more precisely reflecting the accuracy of the prediction results.
[0134] This application's embodiments pre-generate standardized correct results for different types of deterministic tasks during the task construction phase, enabling rapid and accurate comparative evaluation of model outputs during the training phase. This avoids inconsistencies caused by manual annotation and subjective judgment. The use of string or numerical comparison to determine the consistency between operational prediction results and standard answers makes the evaluation process efficient, universal, and computationally inexpensive. Furthermore, by mapping the quantified matching results to reward scores, a clear positive incentive and negative inhibition mechanism is formed, effectively guiding the model to continuously optimize its inference path and output accuracy during reinforcement learning training. Therefore, this not only improves the model's learning efficiency and convergence speed for verifiable problems in e-commerce task scenarios but also enhances the system's stability and scalability in multi-task training.
[0135] S407: For each merchant operation task data with the corresponding task type of the first question and answer task mentioned above, obtain the reference answer corresponding to the merchant operation task data.
[0136] The reference answers for merchant operation task data can be standard answers to operational questions, specifically correct or recommended answers obtained from manual annotation, domain knowledge bases, or existing operational strategy documents. These reference answers can be used to semantically compare the operational prediction results generated by the model, thereby assessing the similarity or consistency between the model output and the ideal answer.
[0137] For example, the operational question could be "Why did activity revenue decline last week?", with a corresponding reference answer being "Mainly due to reduced advertising budget and decreased conversion rates." This reference answer serves as a benchmark for semantic evaluation during the training phase, guiding the model to understand reasonable business logic and causal relationships.
[0138] S408: Perform semantic correlation detection on the above operational prediction results and the above reference answers, and generate correlation detection results between the operational prediction results and the above reference answers.
[0139] The relevance detection result can be used to characterize the semantic relevance or semantic distance between the operational prediction result and the reference answer. It is an intermediate result after measuring the semantic similarity between two texts. The relevance detection result can be in the form of a label (such as "relevant / irrelevant") or a continuous value (such as a similarity score of 0–1). Specifically, the relevance detection result can be generated based on semantic matching algorithms, similarity models, or lightweight discriminant models.
[0140] Specifically, the operational prediction results and the reference answers can be input into the semantic encoder to obtain the corresponding semantic representation vectors. The semantic correlation between the two can be calculated based on cosine similarity, dot product similarity or other related functions to obtain the correlation detection results.
[0141] S409: Quantify the above correlation detection results to obtain the corresponding semantic consistency evaluation index.
[0142] The semantic consistency evaluation index can be used as an evaluation parameter to reflect whether the operational prediction results conform to the semantic meaning of the reference answer. This semantic consistency evaluation index can be in binary form (e.g., 1 indicates consistency, 0 indicates inconsistency) or in continuous scoring form (e.g., linearly mapped to the 0–1 interval based on semantic similarity) to adapt to the needs of different reward strategies.
[0143] Optionally, when the correlation detection result is labeled "relevant", the semantic consistency evaluation index can be set to 1; when the correlation detection result is labeled "irrelevant", the semantic consistency evaluation index can be set to 0.
[0144] Furthermore, when the correlation detection result is a continuous similarity score, the correlation detection result can be mapped to the [0,1] interval based on threshold determination or linear normalization to form a semantic consistency evaluation index for subsequent reward calculation.
[0145] S410: Based on the above semantic consistency evaluation index, determine the reward score of the operation forecast result from the preset reward value range.
[0146] The preset reward value range can be the range of reward signals used in reinforcement learning, which is used to constrain the maximum and minimum values of the reward score. For example, the reward value range can include [0,1], [-1,1], [0,2], etc., and different range settings can be adapted to different reward function forms or training frameworks.
[0147] Specifically, when the semantic consistency evaluation index is 1 (i.e. the prediction result is semantically consistent with the reference answer), a high reward value, such as 1, can be selected from the reward range; when the semantic consistency evaluation index is 0, a low reward value or a punitive reward value, such as 0, can be selected.
[0148] Furthermore, when the semantic consistency evaluation index is a continuous value, linear interpolation can be used to generate corresponding continuous reward scores within the reward interval, so that predictions with closer semantics receive higher rewards. This enables fine-grained quality assessment of open-ended question-answering tasks, allowing reward signals to effectively guide the model to optimize towards business rationality and semantic consistency during the reinforcement learning training phase.
[0149] This application's embodiments, by obtaining standardized reference answers, enable the model to accurately align business logic and industry knowledge during the training phase, thereby improving the reliability of the evaluation benchmark. Furthermore, through a semantic relevance detection mechanism, it eliminates the need for literal consistency judgments, instead comprehensively measuring the closeness between operational prediction results and reference answers in terms of business meaning, logical causality, and expression based on semantic vector representations, making the detection process more intelligent and robust. Subsequently, by quantifying the semantic relevance results into semantic consistency evaluation indicators and constructing a controllable and scalable scoring system, it can flexibly switch between binary judgment and continuous interval scoring modes to adapt to different training objectives and task requirements, thereby effectively improving the model's recognition and generalization abilities in reference answer-based question-answering tasks.
[0150] S411: For each merchant operation task data with the corresponding task type of the second question and answer task mentioned above, obtain the preset reward model corresponding to the second question and answer task type mentioned above.
[0151] The preset reward model can be a reward estimation model used to evaluate the quality of the results generated by the model in open-ended question answering tasks. It can output a continuous reward score by calculating the reasonableness, relevance and expression quality of the generated answer.
[0152] In one embodiment, the pre-defined reward model can be obtained by training a pre-trained language model, incorporating human feedback data or human evaluation samples. During the training phase, a sample set consisting of "questions, model answers, and human scores" is first constructed. Human scores can be based on dimensions such as the reasonableness of the answer, its business relevance, logical completeness, and fluency. Subsequently, the pre-trained language model is optimized using this sample set through reward learning, enabling it to learn the score distribution under different answer qualities, thereby forming a reward model capable of outputting continuous reward scores.
[0153] S412: After performing feature encoding processing on the above-mentioned merchant operation task data and the corresponding above-mentioned operation prediction results, input them into the above-mentioned preset reward model. The above-mentioned preset reward model is used to generate reward scores for the above-mentioned merchant operation task data and the corresponding above-mentioned operation prediction results, so as to obtain the reward score of the above-mentioned operation prediction results output by the above-mentioned preset reward model.
[0154] Specifically, the input to the preset reward model can include merchant operation task data and corresponding operation prediction results; the output can be a reward score for the operation prediction result, which is used to measure the quality of the model's response; the preset reward model can evaluate the input content based on logical rationality, semantic coherence, business relevance, strategic innovation, and expression completeness.
[0155] Optionally, the merchant operation task data can first be structured and parsed, converting the operation question text, related business data, and semantic context information into feature vectors of a unified dimension using a pre-defined text encoder or multimodal feature extraction module. Then, the operation prediction results generated by the initial model are vectorized using the same encoding system to obtain a feature representation of the prediction results that can be used to assess the model's prediction quality. Subsequently, the operation task feature vector and the prediction result feature vector are concatenated or interactively encoded according to a pre-defined format to form a joint feature input for quality judgment.
[0156] Furthermore, after completing the above feature construction, the joint features can be input into the preset reward model. The preset reward model, based on its internally trained reward mapping capabilities, performs a comprehensive evaluation of the input content from multiple dimensions, including logical rationality, semantic relevance, business effectiveness, and expression completeness, thereby performing reward score generation processing and finally outputting a continuous reward score for the operational prediction result.
[0157] This application's embodiments train a reward model based on a pre-trained language model, incorporating human feedback data or manual scoring samples. The constructed pre-defined reward model learns the implicit mapping relationship between answer quality and score, thus automatically calculating the model's output reward score without human intervention during actual training. This reward model comprehensively considers multiple factors such as the logical rationality, semantic coherence, business relevance, strategic innovation, and expressive completeness of the answer, generating continuous reward signals. This enables fine-grained quality assessment of open-ended natural language answers, improving the automation level and evaluation consistency of the reinforcement learning training phase.
[0158] In one embodiment, such as Figure 5 As shown, another method for training a merchant operation model is provided, including the following steps:
[0159] S501: Obtain the initial model and multiple merchant operation task data, which include operation issues and business data related to the operation issues.
[0160] Specifically, S501 is the same as S201, and will not be repeated here.
[0161] S502: Using the initial model described above, perform task execution processing on the above multiple merchant operation task data to obtain the operation prediction results corresponding to each merchant operation task data in the above multiple merchant operation task data.
[0162] Specifically, S502 is the same as S202, and will not be repeated here.
[0163] S503: Based on the task type of the above-mentioned merchant operation task data, determine the reward score for the corresponding operation prediction result of the above-mentioned merchant operation task data.
[0164] Specifically, S503 is the same as S203, and will not be repeated here.
[0165] S504: For the above-mentioned merchant operation task data, generate corresponding triplet samples based on the above-mentioned merchant operation task data, the corresponding operation prediction results, and the reward scores of the above-mentioned operation prediction results.
[0166] Specifically, S504 is the same as S204, and will not be repeated here.
[0167] S505: Summarize the multiple triplet samples corresponding to the above multiple merchant operation task data to construct an operation training sample set.
[0168] Specifically, S505 is the same as S205, and will not be repeated here.
[0169] S506: Obtain the preset general task sample set and the first training weights corresponding to the general task sample set, and obtain the second training weights corresponding to the operation training sample set.
[0170] The general task sample set can be a training data set used to maintain the model's general language understanding and basic reasoning capabilities. It can include general task data from non-e-commerce domains, such as mathematical reasoning, code generation, logical judgment, and natural language instruction understanding. This general task sample set, together with the operational training sample set from the e-commerce domain, participates in the model's reinforcement learning training to prevent the model from losing its original general capabilities after fine-tuning in the e-commerce domain.
[0171] Optionally, the first training weight can be a weight parameter assigned to the general task sample set during reinforcement learning, used to control the participation ratio of general task samples in the overall training. By setting the first training weight, the importance of general tasks and vertical tasks can be balanced during model optimization to maintain the stability of the model's language understanding and logical reasoning abilities. The second training weight can be a weight parameter assigned to the operational training sample set during reinforcement learning, used to increase the influence of e-commerce operation-related task samples in model updates. The second training weight can be dynamically adjusted according to business importance or sample size, allowing the model to more focusedly learn the reasoning patterns and strategy preferences of merchant operational tasks.
[0172] In one embodiment, the first training weights may include multiple types of general task data sub-weights. Taking a general task sample set including instruction task data, mathematical task data, and code task data as an example, each task data may have its own corresponding general task data sub-weights, used to control the participation ratio of different tasks in the reinforcement learning training phase. Specifically, sub-weights can be dynamically allocated according to the importance of each task in the model's capability dimension and experimental results. For example, a higher general task data sub-weight (e.g., 0.2) can be allocated to mathematical task data to enhance the model's numerical reasoning ability, a medium general task data sub-weight (e.g., 0.15) can be allocated to instruction task data to maintain consistency between language understanding and execution, and a moderate general task data sub-weight (e.g., 0.1) can be allocated to code task data to maintain logical reasoning and generation capabilities. By flexibly adjusting the general task data sub-weights, a dynamic balance can be achieved between the model's general language capability, logical reasoning, and specific task adaptability, thereby taking into account the synergistic improvement of general reasoning and e-commerce scenario capabilities in the overall reinforcement learning process.
[0173] In one embodiment, the second training weights may also include multiple types of merchant operation task data sub-weights, used to control the influence ratio of different vertical domain task types in the reinforcement learning training process. Specifically, taking merchant operation task data including computation tasks (i.e., merchant operation task data of computation task type), selection tasks (i.e., merchant operation task data of selection task type), judgment tasks (i.e., merchant operation task data of judgment task type), reference question answering tasks (i.e., merchant operation task data of first question answering task type), and open question answering tasks (i.e., merchant operation task data of second question answering task type) as an example, each task can correspond to different merchant operation task data sub-weights to reflect the differences in task complexity and model capability improvement requirements. For example, the sub-weight of the merchant operation task data corresponding to the calculation task can be set to 0.15 to enhance the stability of the model in numerical reasoning and revenue prediction problems; the sub-weight of the selection task can be set to 0.2 to improve the decision accuracy of the model in multi-solution comparison and strategy selection scenarios; the sub-weight of the merchant operation task data corresponding to the judgment task can be set to 0.05 to maintain the sensitivity of the model in logical judgment and risk identification; the sub-weight of the merchant operation task data corresponding to the reference question answering task can be set to 0.1 to optimize the semantic alignment ability of the model in questions with standard answers; and the sub-weight of the merchant operation task data corresponding to the open question answering task can be set to 0.05 to enhance the model's strategy generation and innovative expression ability in scenarios without reference answers.
[0174] It should be noted that the specific values of the above weights can be flexibly set according to actual needs, and this application does not impose any specific restrictions on them.
[0175] S507: Based on the first and second training weights mentioned above, the general task sample set and the operational training sample set mentioned above are jointly arranged to construct a joint training sample set containing multi-source task information.
[0176] The joint training sample set refers to the comprehensive data set obtained by fusing the general task sample set and the operational training sample set according to the first training weight and the second training weight before the parameter update. The general task sample set is used to maintain the model's general language understanding and logical reasoning ability, while the operational training sample set is used to enhance the model's business reasoning and strategy decision-making ability under e-commerce operation tasks. By fusing according to weights, a balance between the two types of abilities can be achieved during training.
[0177] Specifically, the sub-weights corresponding to each general task type in the first training weights (e.g., instruction tasks, mathematical tasks, code tasks, etc.) and the sub-weights corresponding to each merchant operation task type in the second training weights (e.g., calculation tasks, selection tasks, judgment tasks, reference question-answering tasks, and open question-answering tasks) can be read first. Then, weighted sampling can be performed from the corresponding sample sets according to the proportion of sub-weights for each task type, ensuring that the frequency of different categories of samples appearing in the joint training sample set conforms to the preset weight distribution. After forming the initial sampling results, the extracted general samples and operation samples can be further mixed and arranged at the task level, including but not limited to staggered arrangement according to proportions, reordering according to task difficulty or business importance, or splicing and combining based on a multi-source task mixing strategy. This allows the joint training sample set to simultaneously present both language reasoning ability constraints and merchant operation scenario constraints, thereby improving the stability and generalization of the model during training.
[0178] S508: Construct a loss function based on the above joint training sample set.
[0179] The loss function can refer to the objective function constructed based on the reward scores of each sample in the joint training sample set during the reinforcement learning optimization phase. It can be used to guide the model parameters to optimize in the direction of increasing the probability of high reward output by calculating the difference between the model prediction results and high reward behaviors.
[0180] In one embodiment, a loss function based on Group Relative Policy Optimization (GRPO) can be used, which utilizes the relative reward values of the output results within each group as weighting coefficients to strengthen the model's preference for high-quality outputs.
[0181] S509: Perform gradient backpropagation and parameter iterative update on the initial model using the above loss function to generate the target merchant operation model.
[0182] Understandably, the optimized parameters of the model can be obtained after backpropagation of the loss function and multiple rounds of gradient updates. These parameters integrate the learning features of general tasks and merchant operation tasks, enabling the model to maintain linguistic and logical stability while possessing the ability to predict, analyze, and optimize strategies for e-commerce scenarios, thus forming the final target merchant operation model.
[0183] This application's embodiments, by setting first training weights for the general task sample set and further dividing the general task data into sub-weights, can dynamically maintain the model's robust performance in language understanding, logical reasoning, and code generation. By setting second training weights for the operational training sample set and refining the sub-weights for merchant operational task data, differentiated reinforcement training can be achieved for various types of tasks such as calculation, selection, judgment, and question answering, thereby significantly improving the model's business understanding, revenue prediction, and strategy decision-making capabilities in e-commerce scenarios. Furthermore, by constructing a joint training sample set based on weight fusion and performing gradient backpropagation and parameter iterative updates on the initial model using the GRPO loss function, the model can continuously strengthen high-reward outputs while avoiding degradation of general capabilities, ultimately generating a target merchant operation model that combines general inference stability with intelligent decision-making capabilities for merchant operations.
[0184] In one embodiment, such as Figure 6 As shown, a merchant operation model reasoning method is provided and applied to e-commerce scenarios. The method includes the following steps:
[0185] S601: Obtain pending merchant operation task data, which includes pending operation questions and related business data.
[0186] In one embodiment, the pending merchant operation task data can refer to the data set input in real time by the merchant or system after the target merchant operation model is deployed, used for intelligent analysis and decision support. This data may contain contextual information required for the model to perform inference, describing the merchant's current operational situation. Specifically, the pending merchant operation task data may include operational questions to be answered and related business data, such as product sales data, inventory status, marketing activity information, user behavior characteristics, transaction conversion metrics, or historical operating records, to support the model's analysis and inference.
[0187] Optionally, the operational question to be answered can refer to the specific business problem or decision-making need that the merchant or system hopes the target merchant's operational model will solve. For example, a merchant might input the question, "What is the expected revenue of this promotional activity?" This question describes the operational goal that the merchant is concerned with. The model analyzes and infers by combining relevant business data to output the corresponding target operational prediction result, which is used to assist the merchant in making intelligent decisions and adjusting strategies.
[0188] Optionally, the relevant business data can refer to various e-commerce operational information directly or indirectly related to the operational question to be answered, used to support the model in reasoning and decision analysis. This business data typically originates from the merchant's real-world operational scenarios, including but not limited to: product management data (such as product prices, inventory, and new product launch schedules), store operation data (such as traffic, conversion rates, and customer service response), marketing activity data (such as advertising placement, promotional strategies, and discount strength), user behavior data (such as click-through rates, repurchase rates, and user profile characteristics), and transaction conversion data (such as order volume, transaction amount, and average order value). This data can provide the model with a basis for decision-making, enabling the model to combine historical performance with current strategies to deduce reasonable answers or optimization suggestions for the question.
[0189] S602: Input the above-mentioned merchant operation task data to be processed into the target merchant operation model to obtain the target operation prediction results output by the above-mentioned target merchant operation model.
[0190] Specifically, the target operational prediction result refers to the conclusions or predictions output by the target merchant operational model after reasoning and analysis based on the input data of the merchant's operational tasks to be processed. It represents the model's final answer to operational questions. The target operational prediction result can directly provide merchants with decision-making references, supporting various business applications such as revenue forecasting, strategy selection, activity evaluation, and risk analysis, thereby achieving intelligent operational optimization.
[0191] Among them, the target merchant operation model can adopt Figures 2-5 The merchant operation model was trained using the method shown.
[0192] This application embodiment achieves intelligent analysis and prediction of the actual business scenario of merchants by inputting pending merchant operation task data, including operational issues and related business data, into the target merchant operation model. The target merchant operation model can integrate multi-dimensional business data (including product, store, marketing, user, and transaction information) and combine existing strategies and historical performance to automatically infer the target operation prediction results corresponding to the operational issues. Through this process, intelligent decision support such as revenue prediction, strategy evaluation, risk warning, and optimization suggestions can be provided to merchants, significantly reducing the cost of manual analysis and improving decision accuracy and response efficiency.
[0193] Please refer to the following. Figure 7 This is a schematic diagram of the structure of a merchant operation model training device provided in an exemplary embodiment of this application. Figure 7 As shown, the aforementioned merchant operation model training device 700 includes:
[0194] The first acquisition module 701 is used to acquire an initial model and multiple merchant operation task data. The merchant operation task data includes operation questions and business data related to the operation questions. The multiple merchant operation task data includes merchant operation task data of the first question and answer task type and merchant operation task data of the second question and answer task type. The first question and answer task type is used to represent merchant operation question and answer questions with reference answers, and the second question and answer task type is used to represent open-ended merchant operation question and answer questions without reference answers.
[0195] Processing module 702 is used to perform task execution processing on the above-mentioned multiple merchant operation task data through the above-mentioned initial model, and obtain the operation prediction results corresponding to each merchant operation task data in the above-mentioned multiple merchant operation task data.
[0196] The determination module 703 is used to determine the reward score of the operation prediction result corresponding to the above-mentioned merchant operation task data based on the task type of the above-mentioned merchant operation task data.
[0197] The generation module 704 is used to generate corresponding triplet samples based on the above-mentioned merchant operation task data, the operation prediction results corresponding to the above-mentioned merchant operation task data, and the reward scores of the above-mentioned operation prediction results.
[0198] Module 705 is used to summarize multiple triplet samples corresponding to the above-mentioned multiple merchant operation task data and construct an operation training sample set;
[0199] Training module 706 is used to perform reinforcement learning training on the initial model based on the above-mentioned operational training sample set to generate the target merchant operational model.
[0200] In one embodiment, the processing module 702 is specifically used to: perform semantic parsing on the merchant operation task data to obtain the task type of the merchant operation task data; obtain the corresponding preset paradigm information according to the task type; perform field decoupling and semantic concatenation processing on the merchant operation task data and the preset paradigm information to generate prompt information; and input the prompt information and the merchant operation task data into the initial model after fusion encoding to trigger the initial model to perform inference calculation and output the corresponding operation prediction result.
[0201] In one embodiment, the determining module 703 is specifically configured to: obtain reference answers corresponding to the merchant operation task data for each merchant operation task data whose corresponding task type is the first question-and-answer task type; perform semantic correlation detection on the operation prediction result and the reference answer to generate correlation detection results between the operation prediction result and the reference answer; quantify the correlation detection results to obtain the corresponding semantic consistency evaluation index; and determine the reward score of the operation prediction result from a preset reward value range based on the semantic consistency evaluation index.
[0202] In one embodiment, the determining module 703 is specifically used to: obtain a preset reward model corresponding to the second question-and-answer task type for each merchant operation task data whose corresponding task type is the second question-and-answer task type; input the merchant operation task data and the corresponding operation prediction results into the preset reward model after feature encoding processing; and generate reward scores for the merchant operation task data and the corresponding operation prediction results through the preset reward model to obtain the reward score of the operation prediction results output by the preset reward model.
[0203] In one embodiment, the task types of the aforementioned multiple merchant operation task data may include at least one of the following: calculating task type, selecting task type, and determining task type.
[0204] In one embodiment, the determining module 703 is specifically configured to: obtain a preset reference execution result of the merchant operation task data when the task type of the merchant operation task data is any one of the calculation task type, the selection task type, and the judgment task type; match the operation prediction result corresponding to the merchant operation task data with the preset reference execution result to obtain a matching metric result; construct a corresponding execution consistency evaluation index based on the matching metric result; and determine the reward score of the operation prediction result from a preset reward value range according to the execution consistency evaluation index.
[0205] In one embodiment, the training module 705 is specifically configured to: obtain a preset general task sample set and a first training weight corresponding to the general task sample set, and obtain a second training weight corresponding to the operation training sample set; based on the first training weight and the second training weight, jointly arrange the general task sample set and the operation training sample set to construct a joint training sample set containing multi-source task information; construct a loss function based on the joint training sample set; and perform gradient backpropagation and parameter iterative update on the initial model through the loss function to generate a target merchant operation model.
[0206] Please refer to the following. Figure 8 This is a schematic diagram of the structure of a merchant operation model inference device provided in an exemplary embodiment of this application. Figure 8 As shown, the aforementioned merchant operation model inference device 800 includes:
[0207] The second acquisition module 801 is used to acquire merchant operation task data to be processed. The merchant operation task data includes operation questions to be answered and business data related to the operation questions to be answered.
[0208] The input module 802 is used to input the above-mentioned merchant operation task data to be processed into the target merchant operation model, and obtain the target operation prediction result output by the target merchant operation model.
[0209] The aforementioned target merchant operation model is generated by training the aforementioned merchant operation model using the aforementioned merchant operation model training method.
[0210] Each module in the aforementioned merchant operation model training device 700 and merchant operation model inference device 800 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0211] This application also provides an electronic device, which can be a server, and its internal structure diagram can be as follows: Figure 9 As shown. Please see below. Figure 9The electronic device 900 includes a processor 910, a memory 920, an input / output interface (I / O) 930, and a communication interface 940. The processor 910, memory 920, and I / O interface 930 are connected via a system bus 950, and the communication interface 940 is connected to the system bus 950 via the I / O interface 930. The processor 910 provides computing and control capabilities. The memory 920 includes a non-volatile storage medium 921 and internal memory 922. The non-volatile storage medium 921 stores an operating system 9211, computer programs 9212, and a database 9213. The internal memory 922 provides an environment for the operation of the operating system 9211 and computer programs 9212 stored in the non-volatile storage medium 921. The database 9213 stores business status characteristics, etc. The I / O interface 930 is used for exchanging information between the processor 910 and external devices. The communication interface 940 of the electronic device 900 is used to communicate with external terminals via a network. The processor 910 of the electronic device 900 executes a computer program 9212 to implement a merchant operation model training method or a merchant operation model inference method.
[0212] This application also provides a computer storage medium storing instructions that, when executed on a computer or processor 910, cause the computer or processor 910 to perform one or more steps in the above embodiments. If the constituent modules of the above-described electronic device 900 are implemented as software functional units and sold or used as independent products, they can be stored in the above-described computer-readable storage medium.
[0213] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted through the computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0214] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.
[0215] The embodiments described above are merely preferred embodiments of this application and are not intended to limit the scope of this application. Any modifications and improvements made to the technical solutions of this application by those skilled in the art without departing from the spirit of this application should fall within the protection scope defined by the claims.
[0216] The information, data, and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0217] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A method for training a merchant operation model, characterized in that, include: Acquire an initial model and multiple merchant operation task data, wherein the merchant operation task data includes operation questions and business data related to the operation questions, and the multiple merchant operation task data includes merchant operation task data of a first question and answer task type and merchant operation task data of a second question and answer task type, wherein the first question and answer task type is used to represent merchant operation question and answer questions with reference answers, and the second question and answer task type is used to represent open-ended merchant operation question and answer questions without reference answers; The initial model is used to perform task execution processing on the multiple merchant operation task data to obtain the operation prediction results corresponding to each merchant operation task data in the multiple merchant operation task data. The reward score for the operational prediction result corresponding to each merchant's operational task data is determined based on the task type of each merchant's operational task data; For each merchant's operational task data, a corresponding triplet sample is generated based on the merchant's operational task data, the operational prediction results corresponding to the merchant's operational task data, and the reward score of the operational prediction results; The multiple triplet samples corresponding to the multiple merchant operation task data are summarized to construct an operation training sample set; The initial model is trained using reinforcement learning based on the operational training sample set to generate a target merchant operational model. The step of processing the multiple merchant operation task data using the initial model to obtain the operation prediction result corresponding to each merchant operation task data includes: performing semantic parsing on each merchant operation task data to obtain the task type of each merchant operation task data; obtaining the corresponding preset paradigm information according to the task type; performing field decoupling and semantic concatenation processing on the merchant operation task data and the preset paradigm information to generate prompt information; and inputting the prompt information and the merchant operation task data into the initial model after fusion encoding to trigger the initial model to perform inference calculations and output the corresponding operation prediction result. The process of determining the reward score for the operational prediction result corresponding to each merchant's operational task data based on the task type of each merchant's operational task data includes: For each merchant operation task data corresponding to the first question-and-answer task type, obtain the reference answer corresponding to the merchant operation task data; perform semantic correlation detection on the operation prediction result and the reference answer to generate a correlation detection result between the operation prediction result and the reference answer; quantify the correlation detection result to obtain the corresponding semantic consistency evaluation index; determine the reward score of the operation prediction result from a preset reward value range based on the semantic consistency evaluation index; and / or, For each merchant operation task data corresponding to the second question-and-answer task type, a preset reward model corresponding to the second question-and-answer task type is obtained; after feature encoding processing of the merchant operation task data and the corresponding operation prediction results, the data is input into the preset reward model; the preset reward model is used to generate reward scores for the merchant operation task data and the corresponding operation prediction results, and the reward score of the operation prediction results output by the preset reward model is obtained.
2. The method as described in claim 1, characterized in that, The task types of the multiple merchant operation task data also include at least one of the following: calculating task type, selecting task type, and judging task type.
3. The method as described in claim 2, characterized in that, The process of determining the reward score for the operational prediction result corresponding to each merchant's operational task data based on the task type of each merchant's operational task data includes: If the task type of the merchant operation task data is any one of the calculation task type, the selection task type, and the judgment task type, obtain the preset reference execution result of the merchant operation task data; The operational prediction results corresponding to the merchant's operational task data are matched with the preset reference execution results to obtain a matching quantification result; Construct corresponding execution consistency evaluation indicators based on the matching metric results; Based on the aforementioned consistency evaluation index, the reward score for the operational forecast result is determined from a preset reward value range.
4. The method as described in claim 1, characterized in that, The step of training the initial model using reinforcement learning based on the operational training sample set to generate a target merchant operational model includes: Obtain a preset general task sample set and the first training weights corresponding to the general task sample set, and obtain the second training weights corresponding to the operation training sample set; Based on the first training weight and the second training weight, the general task sample set and the operation training sample set are jointly arranged to construct a joint training sample set containing multi-source task information. Construct a loss function based on the joint training sample set; The initial model is subjected to gradient backpropagation and parameter iterative update using the loss function to generate the target merchant operation model.
5. A reasoning method for a merchant operation model, characterized in that, include: Acquire pending merchant operation task data, which includes pending operation questions and related business data. The data of the merchant operation tasks to be processed is input into the target merchant operation model to obtain the target operation prediction results output by the target merchant operation model; The target merchant operation model is generated by training the merchant operation model using the merchant operation model training method as described in any one of claims 1 to 4.
6. A merchant operation model training device, characterized in that, include: The first acquisition module is used to acquire an initial model and multiple merchant operation task data. The merchant operation task data includes operation questions and business data related to the operation questions. The multiple merchant operation task data includes merchant operation task data of a first question-and-answer task type and merchant operation task data of a second question-and-answer task type. The first question-and-answer task type is used to represent merchant operation questions and answers with reference answers, and the second question-and-answer task type is used to represent open-ended merchant operation questions and answers without reference answers. The processing module is used to perform task execution processing on the multiple merchant operation task data through the initial model to obtain the operation prediction results corresponding to each merchant operation task data in the multiple merchant operation task data; The determination module is used to determine the reward score of the operation prediction result corresponding to each merchant's operation task data based on the task type of each merchant's operation task data; The generation module is used to generate corresponding triplet samples based on the merchant operation task data, the operation prediction results corresponding to the merchant operation task data, and the reward scores of the operation prediction results for each merchant operation task data. The construction module is used to summarize multiple triplet samples corresponding to the multiple merchant operation task data and construct an operation training sample set; The training module is used to perform reinforcement learning training on the initial model based on the operational training sample set to generate a target merchant operational model; The processing module is specifically used for: performing semantic parsing on the operational task data of each merchant to obtain the task type of each merchant operational task data; obtaining the corresponding preset paradigm information according to the task type; performing field decoupling and semantic concatenation processing on the merchant operational task data and the preset paradigm information to generate prompt information; and inputting the prompt information and the merchant operational task data into the initial model after fusion encoding to trigger the initial model to perform inference operations and output the corresponding operational prediction results. The determining module is specifically used for: obtaining reference answers corresponding to merchant operation task data for each merchant operation task data whose corresponding task type is the first question-and-answer task type; performing semantic correlation detection on the operation prediction result and the reference answer to generate a correlation detection result between the operation prediction result and the reference answer; quantifying the correlation detection result to obtain a corresponding semantic consistency evaluation index; determining the reward score of the operation prediction result from a preset reward value range based on the semantic consistency evaluation index; and / or, For each merchant operation task data corresponding to the second question-and-answer task type, a preset reward model corresponding to the second question-and-answer task type is obtained; after feature encoding processing of the merchant operation task data and the corresponding operation prediction results, the data is input into the preset reward model; the preset reward model is used to generate reward scores for the merchant operation task data and the corresponding operation prediction results, and the reward score of the operation prediction results output by the preset reward model is obtained.
7. A business operation model reasoning device, characterized in that, include: The second acquisition module is used to acquire merchant operation task data to be processed, the merchant operation task data including operation questions to be answered and business data related to the operation questions to be answered; The input module is used to input the data of the merchant operation task to be processed into the target merchant operation model, and obtain the target operation prediction result output by the target merchant operation model; The target merchant operation model is generated by training the merchant operation model using the merchant operation model training method as described in any one of claims 1 to 4.
8. An electronic device, characterized in that, include: Processor and memory; The memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as described in any one of claims 1 to 5.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the steps of the method as described in any one of claims 1 to 5.