Data processing model training method and device, data processing method and title generation method
By constructing a data processing model that combines object attributes and related attribute information, and utilizing multi-dimensional scoring weights and reinforcement learning, the multi-objective conflict problem in e-commerce title optimization is solved, generating high-quality, personalized e-commerce titles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-27
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional e-commerce title optimization methods struggle to simultaneously address search exposure, user conversion, and content quality, leading to conflicting objectives and a lack of in-depth personalization and project understanding.
By constructing a data processing model, the attributes and associated attribute information of the target object are determined. Multi-dimensional scoring weights are used to generate prompt text, and the model is adjusted through reinforcement learning algorithms to generate e-commerce titles that conform to multi-objective preferences.
It enables the generation of high-quality e-commerce titles in multi-objective scenarios, taking into account exposure, conversion and content quality, adapting to different project preferences, and improving the accuracy and personalization of title generation.
Smart Images

Figure CN121746044A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a data processing model training method, a data method, and a title generation method. Background Technology
[0002] E-commerce titles are a key factor affecting product search ranking and conversion rates, directly impacting product exposure, click-through rates, and final transaction performance.
[0003] However, traditional title optimization faces the challenge of conflicting objectives: increasing keyword density to boost search exposure often reduces users' willingness to convert, and excessive keyword stuffing can affect title readability; while strengthening marketing terms may affect search exposure and title quality. Summary of the Invention
[0004] In view of the above, embodiments of this specification provide a data processing model training method. One or more embodiments of this specification also relate to a data processing model training apparatus, a data processing method, a title generation method, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a data processing model training method is provided, comprising: Determine the object attribute information and associated attribute information of the target object, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension; The object attribute information, the associated attribute information, and the prompt text are input into the data processing model to generate the target object text; Based on the rating weights of each rating dimension, the target object text is rated according to at least one rating dimension to obtain the text rating of the target object text. The text score is determined as a reward signal, and the data processing model is adjusted using a reinforcement learning algorithm based on the reward signal to obtain an adjusted text generation model.
[0006] According to a second aspect of the embodiments of this specification, a data processing model training apparatus is provided, comprising: The determination module is configured to determine the object attribute information and associated attribute information of the target object, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension; The generation module is configured to input the object attribute information, the associated attribute information, and the prompt text into a data processing model to generate target object text. The scoring module is configured to score the target object text based on the scoring weights of each scoring dimension, thereby obtaining a text score for the target object text. The training module is configured to determine the text score as a reward signal and use the reward signal to adjust the data processing model through a reinforcement learning algorithm to obtain an adjusted text generation model.
[0007] According to a third aspect of the embodiments of this specification, a data processing method is provided, comprising: Determine the object attribute information and associated attribute information of the target object, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension; The object attribute information, the associated attribute information, and the prompt text are input into the text generation model to obtain the target object text output by the text generation model, wherein the text generation model is obtained based on the data processing model training method described above.
[0008] According to a fourth aspect of the embodiments of this specification, a title generation method is provided, comprising: Determine the product attribute information and related attribute information of the target product, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension; The product attribute information, the associated attribute information, and the prompt text are input into the text generation model to obtain the target product title output by the text generation model, wherein the text generation model is obtained based on the data processing model training method described above.
[0009] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps of the above-mentioned data processing method, title generation method, and data processing model training method.
[0010] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the above-described data processing method, title generation method, and data processing model training method.
[0011] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method, title generation method, and data processing model training method.
[0012] This specification provides a data processing model training method in one embodiment, which determines the object attribute information and associated attribute information of the target object to provide a data foundation for generating target object text. At least one rating dimension and the rating weight of each rating dimension provide a reference for evaluating the target object text. Based on at least one rating dimension and the rating weight of each rating dimension, prompt text is constructed to provide a clear generation preference direction for generating target object text. The object attribute information, associated attribute information, and prompt text are input into the data processing model to generate target object text. Based on the rating weight of each rating dimension, the target object text is quantitatively evaluated in at least one rating dimension to obtain a comprehensive text score of the target object text. The text score is determined as a reward signal, and the data processing model is adjusted using a reinforcement learning algorithm using the reward signal to obtain an adjusted text generation model. This enables the text generation model to learn to generate high-quality text that meets the corresponding preference requirements under specific rating weights. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating a data processing model training method provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating a data processing method provided in one embodiment of this specification; Figure 3 This is a flowchart illustrating a title generation method provided in one embodiment of this specification; Figure 4 This is a schematic diagram illustrating a data processing model training method provided in one embodiment of this specification; Figure 5 This is a schematic diagram of the structure of a data processing model training device provided in one embodiment of this specification; Figure 6 This is a schematic diagram of the structure of a data processing device provided in one embodiment of this specification; Figure 7 This is a flowchart of a title generation apparatus provided in one embodiment of this specification; Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0019] Multi-objective reinforcement learning: a reinforcement learning method capable of simultaneously optimizing multiple potentially conflicting objectives. PCGrad: Projecting Conflicting Gradients, is an optimization technique used to resolve gradient conflicts in multi-objective training.
[0020] Prompt Augmentation: A technique that uses structured information to enhance input prompts and control the generation behavior of large language models.
[0021] Pareto Optimal: In multi-objective optimization, the concept of the optimal solution refers to a solution that cannot further improve any objective without compromising other objectives.
[0022] Three-dimensional scoring system: a title evaluation system based on three dimensions: exposure, conversion, and quality.
[0023] Hard constraints: Constraints that must be strictly met, such as title length and compliance.
[0024] Dirichlet distribution: A continuous multivariate probability distribution used to generate probability vectors.
[0025] Conditional generation: a technique for generating text based on specific conditions or control signals.
[0026] E-commerce titles are a key factor in improving product search rankings and conversion rates, directly determining product exposure, click-through rates, and even overall sales. However, traditional methods in title optimization consistently face the fundamental challenge of conflicting objectives. For example, keyword stuffing strategies adopted to increase search exposure often weaken title readability, thereby reducing user click-through and conversion intentions; while marketing expressions emphasizing conversion effects may affect search weight allocation or content quality scores. Furthermore, different product categories and project scenarios (such as daily sales versus major promotions) have significantly different emphases on titles, making it difficult to reconcile personalized needs with standardized generation.
[0027] Current title optimization solutions can be divided into two categories: First, tools built into e-commerce platforms, based on rules or simple machine learning models. These typically optimize around a single objective (such as keyword density or click-through rate prediction) and struggle to meet multi-dimensional comprehensive requirements. Second, title templates and keyword suggestions provided by third-party services lack in-depth project understanding and customization capabilities. Currently available title optimization solutions cannot handle complex multi-objective trade-offs, such as the complex relationship between exposure and conversion, or conversion and quality, and offer low personalization.
[0028] With the development of artificial intelligence technology, reinforcement learning has provided new technical approaches for title optimization in the field of text generation. However, existing research mainly focuses on optimizing a single reward function, with significant shortcomings in exploration and application in multi-objective scenarios. E-commerce title generation is essentially a typical multi-objective optimization problem, requiring simultaneous consideration of three major objectives: search exposure (affecting the probability of discovery), user conversion (affecting clicks and purchase behavior), and content quality (affecting user experience and platform regulations), while strictly meeting hard constraints such as title length and compliance. Traditional technical approaches, such as weighted summation methods with pre-defined weights, are prone to getting stuck in local optima, while converting some objectives into constraints can lead to the loss of important optimization information. Although large language models possess powerful natural language generation capabilities, they lack precise multi-objective control and complex constraint handling mechanisms.
[0029] This specification provides a method for training a data processing model. One or more embodiments of this specification also relate to a data processing model training apparatus, a data processing method, a title generation method, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0030] See Figure 1 , Figure 1 A flowchart of a data processing model training method according to an embodiment of this specification is shown.
[0031] Step 102: Determine the object attribute information and associated attribute information of the target object, and construct the prompt text based on at least one rating dimension and the rating weight of each rating dimension.
[0032] The target object can be understood as the entity to be described in text. For example, in an e-commerce scenario, the target object can be the product for which a title is to be generated, while in a content creation scenario, the target object can be the video or article for which a title is to be generated for publication. The object attribute information can be understood as the data used to describe or characterize the features of the target object. For example, in an e-commerce scenario, the object attribute information can be the brand, category, function, specifications, target audience, and other attribute parameters of the product; while in a content creation scenario, the object attribute information can be the theme, sentiment, keywords, audience tags, and other content features of the video or article.
[0033] Related attribute information can be understood as extended information related to the attributes of a target object, extracted from historical behavior or contextual data. It typically consists of data describing the characteristics of other objects that are associated with the target object. For example, in e-commerce, related attribute information could be high-value features in the titles of similar products corresponding to the target object, such as high-exposure keywords or high-conversion tags. In content creation, related attribute information could be popular title keywords for videos on the same topic. For instance, if a medicine (i.e., the target object) has the attribute "cough-relieving effect," the related attribute information obtained from historical data could include high-exposure keywords such as "nighttime cough relief" and "for children," and high-conversion keywords such as "fast-acting."
[0034] Scoring dimensions can be understood as specific standards or perspectives used to measure or evaluate the quality of generated text. These dimensions are related to the optimization direction of the target text. For example, in the scenario of generating product titles, scoring dimensions may include exposure, conversion, and quality dimensions, while in the scenario of generating article summaries, scoring dimensions may include completeness and logical coherence. The scoring dimensions differ in different scenarios and are not limited here. Scoring weights can be understood as quantified numerical parameters of importance corresponding to each scoring dimension. Hint text can be understood as guiding text input into the data processing model, integrating task descriptions, background information, and generation constraints.
[0035] Specifically, the target object to be described in the text is determined, and the object attribute information corresponding to the characteristics of the target object is obtained. For example, if the target object is a medicine, and a product recommendation title is generated for the medicine, the relevant attribute features of the medicine (i.e., object attribute information) are obtained. For example, the object attribute information includes brand name, product name, efficacy (such as pain relief, cough relief), specifications, and other information.
[0036] When generating text for a target object, the associated attribute information of similar or related objects can be obtained from historical data. This associated attribute information provides guidance for the generation of the target object's text. At least one rating dimension in the prompt text can evaluate the target object text generated by the data processing model from multiple perspectives, thus achieving comprehensive text evaluation.
[0037] In one or more embodiments of this specification, when a target object is determined, the object attribute information of the target object can be determined from an object attribute library containing a large amount of attribute information, and the associated attribute information that is related to the object attribute information and has an impact on the target task can be determined from historical data. Specific implementation methods are described below: Determine the object attribute information and associated attribute information of the target object, including: Identify the target object and determine the object attribute information corresponding to the target object from the object attribute library; Based on the object attribute information, related attribute information is determined from historical data.
[0038] The object attribute library can be understood as a library of object attributes containing a large amount of structured attribute information that has undergone semantic unification processing. Historical data can be understood as data generated by textual descriptions of related objects (objects that belong to the same type as or are similar to the target object). This historical data contains attribute information that affects the target task. For example, when the target task is to generate a high-exposure title, the related attribute information is determined from the high-exposure attribute information covered by the historical data. When the target task is to generate a high-conversion title, the related attribute information is determined from the high-conversion attribute information covered by the historical data.
[0039] Specifically, the target object is identified, and text descriptions and user interaction records of similar or related objects in historical data are filtered based on object attribute information (such as category and efficacy). By analyzing the target task performance of text descriptions in historical data, relevant attribute information that affects the target task is extracted from the text descriptions. For example, if the target task is product conversion, the exposure and conversion performance of attribute information in historical data can be analyzed to mine and extract relevant attribute information, such as the highly exposed keyword "nighttime cough relief" associated with drug efficacy.
[0040] The data processing model training method provided in the embodiments of this specification extracts related attribute information from historical data to provide data basis for subsequent guidance of target object text, ensuring that the generated target object text is consistent with the target object and closely related to historical successful experience.
[0041] Step 104: Input the object attribute information, the associated attribute information, and the prompt text into the data processing model to generate the target object text.
[0042] Among them, the data processing model is an artificial intelligence model with text generation capabilities; the target object text can be understood as the text content output by the data processing model that describes the target object, such as product titles or video titles.
[0043] Specifically, the determined object attribute information, associated attribute information, and the constructed prompt text are input into the data processing model. When the data processing model generates the object text of the target object based on the object attribute information and associated attribute information, it will consider the rating weight of each rating dimension in the prompt text. The rating weight of each rating dimension represents the preference for object text generation. Therefore, when guiding the data processing model to generate object text based on the prompt text, it guides the data processing model to generate target object text that conforms to the preference.
[0044] For example, in the scenario of generating drug titles, the object attribute information includes "Brand: A; Product Name: Ganmaoling Granules; Efficacy: Antipyretic and Analgesic". The prompt text contains each scoring dimension and its corresponding scoring weight, such as "Exposure Optimization: 0.5, Conversion Optimization: 0.3, Quality Optimization: 0.2". This scoring weight guides the data processing model to make more use of high-exposure words when generating target object text based on object attribute information and related attribute information.
[0045] By combining object attribute information, associated attribute information, and structured prompt text as input, the data processing model is guided to produce target object text that conforms to both object characteristics and project preferences.
[0046] Step 106: Based on the rating weights of each rating dimension, score the target object text according to at least one rating dimension to obtain the text score of the target object text.
[0047] Among them, text scoring can be understood as the comprehensive quantitative score of the target text, which is used to evaluate the overall performance level of the target text across multiple scoring dimensions.
[0048] Specifically, when generating target text, the target text is scored one by one according to the preset scoring dimensions, and the overall text score of the target text is calculated based on the scoring weight of each scoring dimension.
[0049] In one or more embodiments of this specification, the at least one rating dimension includes a first rating dimension and a second rating dimension; in the first rating dimension, the target object text is rated using associated attribute information to obtain a first rating result; in the second rating dimension, the target object text is rated using object attribute information to obtain a second rating result. Specific implementation methods are as follows: Based on the rating weights of each rating dimension, the target object text is rated according to at least one rating dimension to obtain a text rating of the target object text, including: Based on the associated attribute information, the target object text is scored according to the first scoring dimension to obtain the first scoring result of the target object text; Based on the object attribute information, the target object text is scored according to the second scoring dimension to obtain the second scoring result of the target object text; The text score of the target object text is obtained based on the first score result, the second score result, and the score weights of each score dimension.
[0050] The first and second scoring dimensions are used to evaluate different indicators of the target object text. In this embodiment, the first scoring dimension can be understood as the dimension for evaluating the target object text by combining the associated attribute information of historical data, and the second scoring dimension can be understood as the dimension for evaluating the target object text based on the target object's own attributes (i.e., object attribute information). The first scoring result and the second scoring result can be understood as the quantitative scores of the target object text calculated on the first and second scoring dimensions, respectively.
[0051] Specifically, when generating target object text, the target object text is evaluated using a first scoring dimension and a second scoring dimension. In the first scoring dimension, the coverage of related attribute information (such as historically high-exposure words and historically high-conversion words) by the target object text is analyzed to calculate the first scoring result of the target object text. In the second scoring dimension, the target object text is mainly based on the object attribute information itself. The second scoring result of the target object text is calculated by evaluating the completeness, accuracy, and logicality of the object attribute information.
[0052] By utilizing the scoring weights of each scoring dimension, the first and second scoring results of the target text are merged to obtain the final text score of the target text.
[0053] In practical applications, when the scoring dimensions include a first scoring dimension based on external historical data and a second scoring dimension based on the object's own attributes, a more refined and diverse quantitative evaluation of the target object's text can be achieved by conducting independent evaluations and fusing the results on the two different scoring dimensions.
[0054] The data processing model training method provided in the embodiments of this specification, based on the first scoring dimension and the second scoring dimension, can not only take into account the degree to which the target text conforms to historical trends, but also verify the fidelity and richness of the target text to the essential characteristics of the target object.
[0055] In one or more embodiments of this specification, the first scoring dimension includes an exposure dimension and a conversion dimension, and the second scoring dimension includes a quality dimension. The performance of the target object text in terms of exposure and conversion can be evaluated using the associated attribute information in historical data, and the completeness and richness of the target object text can be verified using object attribute information. Specific implementation methods are as follows: Based on the associated attribute information, the target object text is scored according to the first scoring dimension to obtain a first scoring result for the target object text, including: Exposure attribute information and conversion attribute information are determined from the associated attribute information; The target object text is scored using the exposure attribute information to determine the exposure score of the target object text in the exposure dimension, wherein the exposure score is used to evaluate the degree to which the target object text covers the exposure attribute information; The conversion attribute information is used to score the target object text to determine the conversion score of the target object text in the conversion dimension, wherein the conversion score is used to evaluate the degree to which the target object text covers the conversion attribute information; Based on the exposure score and conversion score of the target text, the first score result of the target text is obtained.
[0056] Based on the object attribute information, the target object text is scored according to the second scoring dimension to obtain a second scoring result for the target object text, including: The target object text is scored using the object attribute information to determine the quality score of the target object text in the quality dimension, wherein the quality score is used to evaluate the completeness of the object attribute information in the target object text; Based on the quality score of the target object text, a second score result for the target object text is obtained.
[0057] Exposure attribute information can be understood as vocabulary and phrase features extracted from historical data that are highly relevant to display opportunities; conversion attribute information can be understood as vocabulary or phrases extracted from historical data that can effectively promote user conversion behaviors such as clicks, purchases, and registrations; exposure score is used to quantify the degree to which the target object text covers exposure attribute information, and a higher degree of coverage means a greater chance of displaying the target object text in search scenarios; conversion score is used to quantify the degree to which the target object text covers conversion attribute information, and a higher degree of coverage means a greater ability to stimulate user interaction or conversion intentions. Quality score is used to quantify the completeness and accuracy of the target object text's coverage of object attribute information.
[0058] Specifically, exposure attribute information (such as high-search-volume keywords) used to assess exposure potential and conversion attribute information (such as high-conversion keywords) used to assess conversion intention are filtered from the associated attribute information. When generating target text, the exposure score of the target text in the exposure dimension is determined by calculating the coverage of the exposure attribute information by the target text; the conversion score of the target text in the conversion dimension is determined by calculating the usage of the conversion attribute information by the target text.
[0059] By analyzing whether the target object text comprehensively and accurately contains the given object attribute information (such as brand, category, efficacy), the quality score of the target object text in the quality dimension is calculated.
[0060] When the prompt text contains the rating weights for each rating dimension, the rating weights for exposure, conversion, and quality dimensions can be determined separately. Based on the rating weights on these three dimensions, the exposure rating, conversion rating, and quality rating of the target text can be weighted and summed to calculate the text rating of the target text.
[0061] For example, when generating a drug title for "Vitamin C Effervescent Tablets", the exposure attribute information determined from the associated attribute information includes the high-frequency search term "immunity", the conversion attribute information includes high-conversion words such as "sweet and sour taste" and "family staple", and the target attribute information is determined to include "Brand P, Vitamin C Effervescent Tablets, 20 tablets / box, suitable for adults".
[0062] The target text "Brand P Vitamin C Effervescent Tablets Boost Immunity" covers "immunity" and has a high exposure score, for example, an exposure score of 8; however, it does not cover high-conversion keywords such as "sweet and sour and delicious" and has a lower conversion score, for example, a conversion score of 6. When candidate title A fully includes the brand, generic name and efficacy, it has a good quality score, for example, a quality score of 9.
[0063] The data processing model training method provided in the embodiments of this specification, when the scoring dimensions include exposure, conversion, quality, etc., can accurately reflect the performance potential of the target object text in the project objectives based on the exposure and conversion dimensions, and ensure the quality of the target object text itself based on the quality dimension; and by decomposing the comprehensive text scoring into multiple quantitative scoring dimensions such as exposure, conversion, and quality for calculation, the evaluation process becomes more refined and objective.
[0064] In one or more embodiments of this specification, when the scoring weights of each scoring dimension are integrated into the prompt text, the exposure score, conversion score, and quality score obtained above are weighted and summed based on the scoring weights in the prompt text to obtain the final text score. Specific implementation methods are described below: Based on the first scoring result, the second scoring result, and the scoring weights of each scoring dimension, a text score for the target object text is obtained, including: Based on the scoring weights of the exposure dimension, the conversion dimension, and the quality dimension, the exposure score, conversion score, and quality score of the target object text are weighted and summed to obtain the text score of the target object text.
[0065] Specifically, based on the exposure score, conversion score, and quality score of each target text, the three score results of the target text are weighted and summed according to the score weight of each score dimension (e.g., exposure weight is 0.5, conversion weight is 0.3, and quality weight is 0.2) to calculate the text score of the target text.
[0066] It should be noted that the rating weights for each rating dimension are obtained through dynamic sampling from the weight distribution. The weight distribution can be understood as the probability distribution or range of possible values for each rating weight; dynamic sampling can be understood as the process of extracting a set of specific weight values from a preset weight distribution based on strategy, scenario, or randomness requirements.
[0067] During the model training phase, a set of specific scoring weights for each scoring dimension (exposure, conversion, quality) is obtained by dynamically sampling from a preset weight distribution (such as the Dirichlet distribution). After obtaining the scoring results for each scoring dimension, the exposure score, conversion score, and quality score of the target text are weighted and summed using these dynamically sampled scoring weights to determine the final text score of the target text.
[0068] The rating weights for each rating dimension indicate the project preference for that dimension. The rating results for each dimension are then fused based on these weights to ensure that the results are integrated according to the set project preferences. This allows the final text rating to accurately reflect the preference orientation under a specific project scenario. Furthermore, by dynamically sampling different rating weights using a Dirichlet distribution, the data processing model ensures effective operation under various project preferences. For example, it can effectively determine the target text from extreme preferences (such as exposure-oriented [0.9, 0.05, 0.05]) to balanced strategies (such as a balanced configuration [0.33, 0.33, 0.34]).
[0069] Using the previous example, the target text has an exposure score of 8, a conversion score of 6, and a quality score of 9. A set of score weights dynamically sampled from the preset weight distribution is: exposure weight (i.e., the score weight of the exposure dimension) 0.6, conversion weight (i.e., the score weight of the conversion dimension) 0.25, and quality weight (i.e., the score weight of the quality dimension) 0.15. Then, the final text score of the target text is 8×0.6+6×0.25+9×0.15=4.8+1.5+1.35=7.65 points.
[0070] The data processing model training method provided in the embodiments of this specification achieves the flexibility and diversity of multi-dimensional evaluation by introducing dynamic sampling scoring weights. The dynamic weights break the rigidity of generation that may be caused by fixed weights, enabling the model to adapt to different generation strategies (such as focusing on exposure during major promotions and focusing on conversion in daily operations), thereby improving the breadth of application scenarios.
[0071] Step 108: The text score is determined as a reward signal, and the data processing model is adjusted using a reinforcement learning algorithm based on the reward signal to obtain the adjusted text generation model.
[0072] The reward signal can be understood as a feedback signal used to evaluate the quality of the text generated by the data processing model, which can guide the learning direction of the data processing model; the text generation model can be understood as an artificial intelligence model with better generation performance obtained after the data processing model has been adjusted through reinforcement learning.
[0073] Specifically, the text scores corresponding to each target object text are used as reward signals during reinforcement learning training. These reward signals are then used to update the policy gradient and adjust the parameters of the data processing model under training using the selected reinforcement learning algorithm. In effect, the text scores measure whether the data processing model has generated the corresponding target object text according to the required item preferences. During training, different combinations of score weights are dynamically sampled (e.g., using a Dirichlet distribution to generate multiple score weight combinations) to simulate different item preferences, guiding the model to learn to generate text with higher scores under various preference configurations.
[0074] In practical applications, gradient conflict resolution techniques (such as PCGrad) are introduced during model training to handle gradient conflicts that may arise from rewards of different dimensions, ensuring stable convergence of the training process and coordinated optimization of multiple objectives. Through iterative optimization, a text generation model is obtained that can better understand and respond to project preferences, generating text with superior overall performance.
[0075] In this embodiment, the logic of reinforcement learning is as follows: score the target object text to obtain a reward signal, and update the data processing model according to the magnitude of the reward signal (a high reward indicates that the target object text generated by the data processing model conforms to the project preference, and encourages the text generation model to generate the target object text under the corresponding scoring weight).
[0076] For example, in a title generation scenario, the data processing model generates target text A with a score of 7. During reinforcement learning training, this score of 7 is used as the reward signal for the generation action. Through numerous similar interactions, the data processing model learns which word combinations and sentence structures receive higher rewards under specific weights. During training, the scoring weights are dynamically changed through sampling; for example, one sampling might have a scoring weight of [exposure weight: 0.8, conversion weight: 0.1, quality weight: 0.1], emphasizing exposure; the next sampling might have a scoring weight of [exposure weight: 0.2, conversion weight: 0.7, quality weight: 0.1], emphasizing conversion. The data processing model needs to adapt to these changes, and PCGrad technology ensures that the model does not oscillate due to conflicting gradient directions when simultaneously optimizing exposure and conversion goals. Ultimately, the adjusted text generation model will be able to generate target text that performs well in the corresponding emphasis under different project preferences (reflected in the weight configuration). For example, it can skillfully use high-traffic keywords when emphasizing exposure, and effectively embed promotional information when emphasizing conversion, while always ensuring basic information quality and compliance.
[0077] The data processing model training method provided in the embodiments of this specification transforms the text scores determined based on the three-dimensional scoring system into reward signals for reinforcement learning, enabling the model to surpass learning based on fixed rules or static examples and learn better text generation strategies through reward signals. Furthermore, by introducing dynamically sampled scoring weights and gradient conflict handling techniques under multi-objective optimization, the text generation model has the ability to flexibly adapt to diverse project preferences.
[0078] See Figure 2 , Figure 2 A flowchart of a data processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0079] Step 202: Determine the object attribute information and associated attribute information of the target object, and construct the prompt text based on at least one rating dimension and the rating weight of each rating dimension.
[0080] Step 204: Input the object attribute information, the associated attribute information, and the prompt text into the text generation model to obtain the target object text output by the text generation model, wherein the text generation model is obtained based on the data processing model training method described above.
[0081] The target object can be understood as any entity that needs to generate a text description, such as a product, article, video, or marketing material; the object attribute information can be understood as attribute data used to characterize the target object itself, such as the brand and category of a product, the theme of an article, or the style tags of a video; the associated attribute information can be understood as words and sentence structures mined from historical data that are related to the performance of similar entities (such as exposure and conversion).
[0082] Specifically, the target object for which text descriptions need to be generated is identified, and its object attribute information is extracted. Simultaneously, based on the target object's category or characteristics, related attribute information is retrieved from historical data and determined. A prompt text is then constructed based on at least one rating dimension and each rating dimension. The object attribute information, related attribute information, and the constructed prompt text are input into the text generation model to obtain the target object text as output. When the text generation model is trained using the aforementioned data processing model training method, it can generate target object text that conforms to project preferences based on set rating weights.
[0083] In practical applications, object attribute information includes the attributes of the target object itself, while associated attribute information includes high-exposure and high-conversion keywords related to the target object. When a user needs to generate target text corresponding to the target object, their preference requirements can be clearly defined. For example, if they want to increase the use of exposure keywords in the title, they can increase the score of the exposure dimension. By setting the scoring weights for each scoring dimension, the text generation model can be guided to adopt certain preferences (e.g., exposure weight 0.6, conversion weight 0.3, and quality weight 0.1). At this point, based on the suggested text, the text generation model will use more high-exposure keywords based on the input object attribute information and the detailed associated attribute information (high-exposure keywords, high-conversion keywords).
[0084] It should be noted that during the model training phase, a multi-objective learning approach is adopted. Therefore, the various scoring dimensions are not isolated. If there is an attribute that can simultaneously improve exposure, conversion, and increase attribute type (quality), the text generation model will prioritize adding this attribute to the target object text.
[0085] In one or more embodiments of this specification, users can configure corresponding rating weights for each rating dimension. Specifically, users can set different rating weights according to their actual needs, thereby generating target object text tailored to their project preferences. Detailed implementation methods are described below: Before constructing the prompt text based on at least one rating dimension and the rating weights of each rating dimension, the following steps are also included: In response to a weight configuration request sent by the client, determine the rating weights for each rating dimension carried in the weight configuration request.
[0086] The weight configuration request can be understood as a request sent by the user through the client to configure the corresponding weights of each rating dimension. In practical applications, the weight configuration template can be displayed on the user interface of the client. The weight configuration template contains each rating dimension and the rating weights of each rating dimension to be filled. Based on the weight configuration template, the conversion from fuzzy semantic control to precise numerical control can be realized. Subsequently, the text generation model can dynamically adjust the generation strategy according to the rating weights input by the user.
[0087] Specifically, the system receives and parses the weight configuration request sent by the client, and obtains the specific rating weight values set by the user for each preset rating dimension from the weight configuration request. Then, it combines at least one preset rating dimension and the rating weights of each rating dimension obtained above with the predefined structured template to construct a prompt text, which is used by the text generation model to perform precise conditional control when generating target object text.
[0088] For example, in a product title generation scenario, the client sends a weight configuration request, specifying that the weight for "exposure optimization" is 0.5, "conversion optimization" is 0.3, and "quality optimization" is 0.2 in this generation task. Therefore, based on the received weight configuration request, the structured weight configuration is parsed as {"exposure weight": 0.5, "conversion weight": 0.3, "quality weight": 0.2}. Based on each rating dimension and its weight, the final prompt text is generated according to a preset structured template, thus providing clear and quantifiable generation instructions for the text generation model.
[0089] The data processing method provided in this specification integrates the dynamically configured scoring weights and scoring dimensions on the client in a structured manner, thereby enabling the text generation process to be guided by project preferences in a precise numerical way. By adjusting the weight configuration, the text generation strategy can be flexibly switched, allowing the text generation model to adapt to diverse project scenarios and needs.
[0090] See Figure 3 , Figure 3 A flowchart of a title generation method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0091] Step 302: Determine the product attribute information and related attribute information of the target product, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension.
[0092] Step 304: Input the product attribute information, the associated attribute information, and the prompt text into the text generation model to obtain the target product title output by the text generation model, wherein the text generation model is obtained based on the data processing model training method described above.
[0093] Among them, product attribute information can be understood as parameter data used to describe the characteristics of the target product, such as brand, category, specifications, efficacy, target audience, appearance, etc.; related attribute information can be understood as words or phrases related to the potential exposure and conversion effect of the product, mined from historical data, such as popular search terms, high click-through rate phrases, etc.; scoring dimensions represent different indicator directions for evaluating the effectiveness of the title, such as exposure potential, conversion guidance, information completeness, etc.; prompt text can be understood as a structured instruction text that integrates the generation task description, related attribute information, and scoring dimensions, used to guide the generation direction of the text generation model.
[0094] Specifically, the target product for which the title is to be generated is identified and its product attribute information is obtained. Related attribute information related to the product is extracted from historical data. In practice, historical data includes historical search data and historical conversion rate data related to the product attribute information corresponding to the target product. Exposure attribute information is extracted from historical search data, and conversion attribute information is extracted from historical conversion rate data.
[0095] Based on at least one preset rating dimension and the rating weight of each rating dimension, a structured prompt text is constructed. The product attribute information, related attribute information and the constructed prompt text are input into the text generation model, and the text generation model generates the target product title based on this information.
[0096] For specific implementation details, please refer to the above embodiments, which will not be repeated here.
[0097] The title generation method provided in this specification combines structured product attribute information, related attribute information, configurable multiple rating dimensions, and rating weights to construct precise generation instructions and evaluation standards, realizing an intelligent and standardized process for product title generation. Based on product attribute information, it ensures that the generated title accurately reflects the core selling points of the product; based on related attribute information, it effectively incorporates historically proven and effective marketing elements; and through quantitative multi-dimensional scoring, it objectively selects target product titles, thereby improving the click-through rate, conversion rate, and content quality of the target product.
[0098] See Figure 4 , Figure 4 This diagram illustrates a data processing model training method provided in one embodiment of this specification.
[0099] In terms of multi-objective optimization, traditional multi-objective evolutionary algorithms (such as NSGA-II and MOEA / D) perform well in continuous optimization problems, but are mainly applicable to numerical optimization tasks and have extremely limited applicability in discrete text generation tasks. Furthermore, traditional methods such as weighted summation require pre-determining weights and cannot dynamically adapt to the differentiated needs of different categories and stages in specific scenarios, making them prone to getting trapped in local optima.
[0100] Existing text generation control technologies, such as prompt engineering based on natural language descriptions, suffer from insufficient control precision. For example, the text instruction "Please improve the conversion effect of the title" is too vague and cannot achieve the quantifiable dynamic allocation of multi-objective weights required by the project. While reinforcement learning techniques based on human feedback have achieved success in model alignment, they often lead to training instability due to gradient conflicts in multi-objective scenarios and lack a structured project preference injection mechanism.
[0101] The data processing model training method provided in the embodiments of this specification adopts an end-to-end optimization framework that can effectively solve the multi-objective conflict problem. Specifically, taking the product title generation scenario as an example, the data processing model training method is described in detail.
[0102] Identify the target product and its attribute information. Specifically, the target product is analyzed in depth through a pre-built product attribute library to accurately extract structured product attribute features. For example, attribute information is extracted from six dimensions: brand name, product name, specifications, efficacy, target audience, and appearance, to form a structured product profile corresponding to the target object.
[0103] The system retrieves historical search data and conversion rate data related to the target product from historical data (which includes existing titles and their corresponding search rates and conversion rates). Specifically, based on a large amount of historical search and user behavior data, it automatically identifies and extracts high-value keywords related to the target product's attributes (high-value keywords include high-exposure keywords identified from historical search data and high-conversion keywords identified from conversion rate data).
[0104] By obtaining a JSON data structure containing high-value vocabulary and product attribute information, a data foundation is provided for subsequent accurate scoring.
[0105] A three-dimensional scoring system is constructed, specifically using an exposure-conversion-quality scoring framework to quantify the overall performance of subsequently generated text. Exposure scoring assesses search exposure potential by calculating the title's coverage of high-exposure keywords; conversion scoring predicts user conversion intent by analyzing the usage of high-conversion keywords; and quality scoring measures the title's information richness through attribute completeness. The system also sets hard constraints such as title length and compliance rules to ensure that generated titles conform to project specifications.
[0106] In practical implementation, during the model training phase, a Dirichlet distribution can be used to dynamically sample different rating weights to ensure that the model works effectively under various project preferences. For example, if the rating weights for exposure, conversion, and quality are set to 0.7, 0.2, and 0.1 respectively (i.e., w=[0.7, 0.2, 0.1]), the final output title will be a title with high exposure preference; if the rating weights for exposure, conversion, and quality are set to 0.2, 0.7, and 0.1 respectively (i.e., w=[0.2, 0.7, 0.1]), The final output title is a high-conversion-preference title; when the scoring weights for exposure, conversion, and quality are set to 0.2, 0.1, and 0.7 respectively (i.e., w=[0.2, 0.1, 0.7]), the final output title is a high-quality-preference title; and when the scoring weights for exposure, conversion, and quality are set to 0.3, 0.3, and 0.4 respectively (i.e., w=[0.3, 0.3, 0.4]), the final output title is a balanced-preference title; of course, titles with custom preferences can be generated through custom weight configurations.
[0107] During the model application phase, users can configure the weights of each scoring dimension according to their custom preferences through the weight configuration template on the client. Based on the weight configuration template, structured scoring weights can be obtained and embedded into the prompt text, avoiding the uncertainty of vague descriptions such as "please optimize the conversion effect". This enables the text generation model to accurately understand the goal orientation of the current optimization task and link it with the training strategy, realizing the transformation from fuzzy semantic control to precise numerical control.
[0108] In practical applications, merchants can set corresponding marketing strategies according to their actual needs (for example, focusing on conversion during major promotions and focusing on brand and quality in daily operations), and flexibly set the scoring weights of the above three scoring dimensions through the weight configuration template on the client user interface.
[0109] The product attribute information, related attribute information, and prompt text (including each rating dimension and its corresponding rating weight) are input into the data processing model for conditional generation to obtain the target product title. Of course, if the target product has an original product title, the original product title can also be input into the text generation model, and the original product title can be optimized through the configured rating weight, product attribute information, and related attribute information.
[0110] In fact, the data processing model provided in this embodiment consists of a data input layer, a vocabulary acquisition module, a three-dimensional scoring module, a multi-objective optimization module, and an intelligent output layer. Based on the data input layer, the original product title, historical search data, conversion rate data, product attribute library, and weight configuration can be input into the data processing model. The vocabulary acquisition module then obtains the aforementioned high-value vocabulary (including high-exposure vocabulary and high-conversion vocabulary) and product attribute information.
[0111] During the training of the data processing model, the multi-objective optimization module calls the three-dimensional scoring module to calculate scores for the generated target text across three dimensions: search exposure, conversion attraction, and content quality. These scores are then combined into a vectorized reward signal, which is fed back to the data processing model to guide its learning and adjustments. Through multi-objective fusion and gradient conflict handling techniques, the multi-objective optimization module enables the data processing model to stably coordinate the relationships between different optimization objectives, ensuring that, while meeting strict constraints, the final title accurately matches the user's project preferences.
[0112] The intelligent output layer can output titles with corresponding preferences based on different weight configurations. As described in the above embodiment, when w=[0.7, 0.2, 0.1], a high exposure preference title is output; when w=[0.2, 0.7, 0.1], a high conversion preference title is output; when w=[0.2, 0.1, 0.7], a high quality preference title is output; and when w=[0.3, 0.3, 0.4], a balanced preference title is output. Furthermore, titles with corresponding custom preferences can be generated through custom weight configurations.
[0113] The final product title is the Pareto optimal solution found under specified project preferences, while meeting hard constraints such as length and compliance requirements. The entire process achieves end-to-end automation from "product attribute information" to "high-quality product title". Users can then obtain product titles that are both standardized and personalized by setting project preferences.
[0114] This data processing model training method addresses the inherent shortcomings of existing technologies in the field of intelligent title generation. Through a multi-objective reinforcement learning mechanism driven by a multi-dimensional scoring system, it constructs a precisely quantifiable three-dimensional scoring system for the three dimensions of exposure, conversion, and quality, transforming these into quantifiable reward signals to drive multi-objective reinforcement learning. Specifically, the exposure score calculates the exposure value of keywords based on historical search and browsing data; the conversion score evaluates the conversion effectiveness of keywords based on click-through rate data; and the quality score measures the information richness of the title through the completeness of attribute coverage. Furthermore, hard constraints such as title length and platform compliance are incorporated into the reward calculation through an indicator function, ensuring that the generated results strictly comply with project specifications.
[0115] In terms of multi-objective reinforcement learning strategies, a Dirichlet distribution is used to dynamically sample diverse scoring weights, enabling the model to adapt to various project scenarios ranging from extreme preferences to balanced strategies. By introducing PCGrad gradient conflict handling technology, the mutual interference of gradient directions in multi-objective training is effectively eliminated, ensuring stable convergence of the training process. This allows the trained text generation model to find a better solution on the Pareto front under any given project weight configuration, thus solving the multi-objective balance problem in the title generation process.
[0116] By constructing a structured Prompt Augmentation (PAR) preference control, the problem of insufficient control precision in conditional generation of large language models is addressed. Specifically, during multi-objective reinforcement learning, item preferences are embedded into the prompt text of the text generation model in a precise numerical form through the structured configuration of the prompt text. In practical applications, by injecting different weight configurations into the prompt text, the model's generation strategy can be dynamically adjusted, enabling a single model to flexibly master multiple optimization strategies, significantly improving the control precision and adaptability.
[0117] The data processing model training method provided in this specification combines a multi-objective scoring system based on accurate quantification of actual project data with structured instruction control to construct an end-to-end title generation method. This method not only applies multi-objective reinforcement learning to discrete text generation tasks, resolving the inherent conflicts in title optimization, but also dynamically embeds project preferences into the model context and links training strategies through a structured weight configuration method. This achieves flexible control of preferences, enhances the text generation model's understanding and response to project preferences, and enables the text generation model to generate titles that perform well in multiple dimensions such as search exposure, user conversion, and content quality, based on dynamic project needs.
[0118] Corresponding to the above method embodiments, this specification also provides embodiments of a data processing model training device. Figure 5A schematic diagram of a data processing model training apparatus according to one embodiment of this specification is shown. Figure 5 As shown, the device includes: The determination module 502 is configured to determine the object attribute information and associated attribute information of the target object, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension; The generation module 504 is configured to input the object attribute information, the associated attribute information and the prompt text into the data processing model to generate the target object text; The scoring module 506 is configured to score the target object text based on the scoring weights of each scoring dimension, thereby obtaining a text score for the target object text. Training module 508 is configured to determine the text score as a reward signal and use the reward signal to adjust the data processing model through a reinforcement learning algorithm to obtain an adjusted text generation model.
[0119] Optionally, the determining module 502 is further configured to: Identify the target object and determine the object attribute information corresponding to the target object from the object attribute library; Based on the object attribute information, related attribute information is determined from historical data.
[0120] Optionally, the scoring module 506 is further configured to: Based on the associated attribute information, the target object text is scored according to the first scoring dimension to obtain the first scoring result of the target object text; Based on the object attribute information, the target object text is scored according to the second scoring dimension to obtain the second scoring result of the target object text; The text score of the target object text is obtained based on the first score result, the second score result, and the score weights of each score dimension.
[0121] Optionally, the scoring module 506 is further configured to: Exposure attribute information and conversion attribute information are determined from the associated attribute information; The target object text is scored using the exposure attribute information to determine the exposure score of the target object text in the exposure dimension, wherein the exposure score is used to evaluate the degree to which the target object text covers the exposure attribute information; The conversion attribute information is used to score the target object text to determine the conversion score of the target object text in the conversion dimension, wherein the conversion score is used to evaluate the degree to which the target object text covers the conversion attribute information; Based on the exposure score and conversion score of the target text, the first score result of the target text is obtained.
[0122] Optionally, the scoring module 506 is further configured to: The target object text is scored using the object attribute information to determine the quality score of the target object text in the quality dimension, wherein the quality score is used to evaluate the completeness of the object attribute information in the target object text; Based on the quality score of the target object text, a second score result for the target object text is obtained.
[0123] Optionally, the scoring module 506 is further configured to: Based on the scoring weights of the exposure dimension, the conversion dimension, and the quality dimension, the exposure score, conversion score, and quality score of the target object text are weighted and summed to obtain the text score of the target object text.
[0124] The above is an illustrative scheme of a data processing model training device according to this embodiment. It should be noted that the technical solution of this data processing model training device and the technical solution of the data processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the data processing model training device, please refer to the description of the technical solution of the data processing model training method described above.
[0125] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 6 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 6 As shown, the device includes: The determination module 602 is configured to determine the object attribute information and associated attribute information of the target object, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension; The generation module 604 is configured to input the object attribute information, the associated attribute information, and the prompt text into the text generation model to obtain the target object text output by the text generation model, wherein the text generation model is obtained by the above-mentioned data processing model training method.
[0126] The device further includes: The response module is configured to respond to a weight configuration request sent by the client and determine the rating weights of each rating dimension carried in the weight configuration request.
[0127] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.
[0128] Corresponding to the above method embodiments, this specification also provides embodiments of a title generation apparatus. Figure 7 A schematic diagram of a title generation apparatus according to one embodiment of this specification is shown. Figure 7 As shown, the device includes: The determination module 702 is configured to determine the product attribute information and associated attribute information of the target product, and construct prompt text based on at least one rating dimension and the rating weight of each rating dimension; The generation module 704 is configured to input the product attribute information, the associated attribute information, and the prompt text into the text generation model to obtain the target product title output by the text generation model, wherein the text generation model is obtained by the above-mentioned data processing model training method.
[0129] The above is a schematic scheme of a title generation device according to this embodiment. It should be noted that the technical solution of this title generation device and the technical solution of the title generation method described above belong to the same concept. For details not described in detail in the technical solution of the title generation device, please refer to the description of the technical solution of the title generation method described above.
[0130] Figure 8 A structural block diagram of a computing device 800 according to one embodiment of this specification is shown. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.
[0131] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0132] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0133] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.
[0134] The processor 820 is used to execute the following computer program / instruction, which, when executed by the processor, implements the steps of the above-mentioned data processing method, title generation method, and data processing model training method.
[0135] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the computing device embodiments are basically similar to the data processing method, title generation method, and data processing model training method embodiments, so the description is relatively simple. Relevant parts can be referred to the descriptions of the data processing method, title generation method, and data processing model training method embodiments.
[0136] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method, title generation method, and data processing model training method.
[0137] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the computer-readable storage medium embodiments are relatively simple in description because they are fundamentally similar to the data processing method, title generation method, and data processing model training method embodiments. Relevant parts can be referred to the descriptions in the data processing method, title generation method, and data processing model training method embodiments.
[0138] An embodiment of this specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described data processing method, title generation method, and data processing model training method.
[0139] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the data processing method, title generation method, and data processing model training method described above. For details not described in detail in the technical solution of the computer program product, please refer to the descriptions of the technical solutions of the data processing method, title generation method, and data processing model training method described above.
[0140] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0141] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0142] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0143] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0144] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing model training method, comprising: determining object attribute information of a target object and associated attribute information, and constructing prompt text according to at least one scoring dimension and scoring weights of each scoring dimension; inputting the object attribute information, the associated attribute information and the prompt text into a data processing model to generate target object text; scoring the target object text in the at least one scoring dimension based on the scoring weights of each scoring dimension to obtain a text score of the target object text; determining the text score as a reward signal, and adjusting the data processing model by a reinforcement learning algorithm using the reward signal to obtain an adjusted text generation model.
2. The method of claim 1, wherein determining object attribute information of a target object and associated attribute information comprises: determining a target object, and determining object attribute information corresponding to the target object from an object attribute library; determining associated attribute information related to the object attribute information from historical data according to the object attribute information.
3. The method of claim 1, wherein the at least one scoring dimension comprises a first scoring dimension and a second scoring dimension; scoring the target object text in the at least one scoring dimension based on the scoring weights of each scoring dimension to obtain a text score of the target object text comprises: scoring the target object text in the first scoring dimension according to the associated attribute information to obtain a first scoring result of the target object text; scoring the target object text in the second scoring dimension according to the object attribute information to obtain a second scoring result of the target object text; obtaining a text score of the target object text according to the first scoring result, the second scoring result and the scoring weights of each scoring dimension.
4. The method of claim 3, wherein the first scoring dimension comprises an exposure dimension and a conversion dimension; scoring the target object text in the first scoring dimension according to the associated attribute information to obtain a first scoring result of the target object text comprises: determining exposure attribute information and conversion attribute information from the associated attribute information; scoring the target object text using the exposure attribute information to determine an exposure score of the target object text in the exposure dimension, wherein the exposure score is used to evaluate the coverage of the target object text on the exposure attribute information; scoring the target object text using the conversion attribute information to determine a conversion score of the target object text in the conversion dimension, wherein the conversion score is used to evaluate the coverage of the target object text on the conversion attribute information; obtaining a first scoring result of the target object text according to the exposure score and the conversion score of the target object text.
5. The method of claim 4, wherein the second scoring dimension comprises a quality dimension; scoring the target object text in the second scoring dimension according to the object attribute information to obtain a second scoring result of the target object text comprises: score the target object text according to the object attribute information, to determine a quality score of the target object text in the quality dimension, wherein the quality score is used to evaluate the completeness of the object attribute information in the target object text; obtain a second score result of the target object text according to the quality score of the target object text.
6. The method of claim 5, wherein a text score of the target object text is obtained according to the first score result, the second score result, and score weights of the score dimensions, comprising: performing weighted summation on the exposure score, the conversion score, and the quality score of the target object text according to the score weights of the exposure dimension, the conversion dimension, and the quality dimension, to obtain the text score of the target object text.
7. The method of any one of claims 1-6, wherein the score weights of the score dimensions are obtained by dynamic sampling in a weight distribution.
8. A data processing method, comprising: determining object attribute information and associated attribute information of a target object, and constructing a prompt text according to at least one score dimension and score weights of the score dimensions; inputting the object attribute information, the associated attribute information, and the prompt text into a text generation model, to obtain a target object text output by the text generation model, wherein the text generation model is obtained based on the method of any one of claims 1-7.
9. The method of claim 8, further comprising, before constructing the prompt text according to the at least one score dimension and the score weights of the score dimensions: determining the score weights of the score dimensions carried in a weight configuration request in response to the weight configuration request sent by a client.
10. A title generation method, comprising: determining product attribute information and associated attribute information of a target product, and constructing a prompt text according to at least one score dimension and score weights of the score dimensions; inputting the product attribute information, the associated attribute information, and the prompt text into a text generation model, to obtain a target product title output by the text generation model, wherein the text generation model is obtained based on the method of any one of claims 1-7.
11. A data processing model training apparatus, comprising: a determination module configured to determine object attribute information and associated attribute information of a target object, and construct a prompt text according to at least one score dimension and score weights of the score dimensions; a generation module configured to input the object attribute information, the associated attribute information, and the prompt text into a data processing model, to generate a target object text; a score module configured to score the target object text in the at least one score dimension based on the score weights of the score dimensions, to obtain a text score of the target object text; a training module configured to determine the text score as a reward signal, and adjust the data processing model by a reinforcement learning algorithm using the reward signal, to obtain an adjusted text generation model.
12. A computing device, comprising: a memory and a processor; The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, and the computer programs / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1 to 10.
13. A computer readable storage medium storing computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
14. A computer program product comprising computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN120975225A
Title generation model training method and device, title generation method and device, electronic equipment, medium and program product
CN121189308A
Systems and methods for contextual machine learning prompt generation
US20250078454A1