Information processing method and apparatus

CN122838697APending Publication Date: 2026-09-29SWEET POTATO TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611069678.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,由于发布平台上的标签较多且相关业务人员对标签的理解缺乏全面的认知,因此导致在进行筛选时,存在同质化严重,以及筛选效果不理想的问题

Benefits of technology

[0011]根据本说明书提供的信息处理方法,根据目标筛选需求信息,确定目标约束条件,使得在明确具体筛选需求的同时,以筛选需求为基准,确定目标约束,进一步约束了对后续筛选出的对象的准确度,避免了仅根据标签进行筛选时所导致的同质化严重的问题。在此基础上,确定目标筛选需求信息对应的至少一个候选筛选策略,扩大了后续确定目标筛选策略的筛选范围,并避免直接确定目标筛选策略而导致的对目标筛选对象的不准确的情况。进一步的,根据各候选筛选策略以及目标约束条件确定各候选筛选对应的策略调整属性信息,为调整候选筛选策略提供依据,以使得调整后的候选筛选策略更加满足目标筛选需求。在此基础上,将满足更新停止条件的候选筛选策略确定为目标筛选策略,在提升筛选质量的同时,提升了对筛选目标对象的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838697A_ABST
    Figure CN122838697A_ABST
Patent Text Reader

Abstract

The present specification provides an information processing method and device, wherein the information processing method comprises: obtaining target screening requirement information, and determining a target screening semantic corresponding to the target screening requirement information; determining a target constraint condition corresponding to the target screening requirement information based on the target screening semantic; generating at least one candidate screening strategy based on the target screening requirement information; determining a strategy adjustment attribute information corresponding to each candidate screening strategy based on each candidate screening strategy and the target constraint condition; updating each candidate screening strategy based on each strategy adjustment attribute information until a stop condition of updating is met, and determining at least one target screening strategy; and determining a target screening object based on each target screening strategy. According to the method provided in the present specification, the accuracy of the circle selection of the target circle selection object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to information processing methods. This specification also relates to information processing apparatus, a computing device, a computer-readable storage medium, and computer program products. Background Technology

[0002] With the continuous development of artificial intelligence technology, people have increasingly higher requirements for the convenience and accuracy of filtering objects on publishing platforms.

[0003] Typically, filtering is based on the desired criteria combined with existing tags for the target audience on the publishing platform. However, due to the large number of tags on the platform and the lack of comprehensive understanding of these tags among relevant business personnel, filtering often suffers from severe homogenization and unsatisfactory results. Furthermore, the lack of proprietary knowledge in the filtering strategy model leads to the "illusion problem," further reducing the quality of the filtering.

[0004] Therefore, how to efficiently and accurately screen target objects has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of this specification provide an information processing method. This specification also relates to an information processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the aforementioned problems existing in the prior art.

[0006] According to a first aspect of the embodiments of this specification, an information processing method is provided, comprising: Obtain target filtering requirement information and determine the target filtering semantics corresponding to the target filtering requirement information; Based on the target filtering semantics, determine the target constraints corresponding to the target filtering requirement information; Based on the target screening requirements information, at least one candidate screening strategy is generated; Based on each candidate screening strategy and the target constraints, determine the strategy adjustment attribute information corresponding to each candidate screening strategy; Based on the attribute information of each strategy, update each candidate screening strategy until the update stop condition is met, and determine at least one target screening strategy. Based on the screening strategies for each target, the target screening objects are determined.

[0007] According to a second aspect of the embodiments of this specification, an information processing apparatus is provided, comprising: The acquisition unit is configured to acquire target screening requirement information and determine the target constraints corresponding to the target screening requirement information. The generation unit is configured to generate at least one candidate filtering strategy based on the target filtering requirement information; The judgment unit is configured to input each candidate screening strategy and the target constraint into the strategy discrimination model, and obtain the strategy adjustment attribute information corresponding to each candidate screening strategy output by the strategy discrimination model; The processing unit is configured to update each candidate filtering strategy based on the attribute information of each strategy until the update stop condition is met, and to determine at least one target filtering strategy; and to determine the target filtering object based on each target filtering strategy.

[0008] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described information processing method.

[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the information processing method described above.

[0010] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the information processing method described above.

[0011] According to the information processing method provided in this specification, target constraints are determined based on the target screening requirements. This clarifies the specific screening needs and, using these requirements as a benchmark, determines the target constraints, further constraining the accuracy of the subsequently screened objects and avoiding the severe homogenization problem caused by screening solely based on tags. Furthermore, at least one candidate screening strategy corresponding to the target screening requirements is determined, expanding the screening scope for subsequent target screening strategies and avoiding inaccuracies in the target screening objects caused by directly determining the target screening strategy. Further, based on each candidate screening strategy and the target constraints, strategy adjustment attribute information corresponding to each candidate screening strategy is determined, providing a basis for adjusting the candidate screening strategies to better meet the target screening requirements. Finally, candidate screening strategies that meet the update stopping condition are determined as the target screening strategies, improving both the screening quality and the accuracy of the screened target objects. Attached Figure Description

[0012] Figure 1 A flowchart of an information processing method according to an embodiment of this specification is shown; Figure 2 A flowchart illustrating the training process of a policy discrimination model provided in one embodiment of this specification is shown; Figure 3 A flowchart illustrating an information processing method for selecting a target population, provided in one embodiment of this specification, is shown. Figure 4 shows a selection architecture diagram provided in one embodiment of this specification; Figure 4A This specification illustrates a schematic diagram of a knowledge base construction method provided in one embodiment. Figure 4B This diagram illustrates the construction of a supervised training dataset according to an embodiment of this specification. Figure 4C This specification illustrates a schematic diagram of an embodiment of constructing direct preference data and policy discrimination model data. Figure 4D A schematic diagram illustrating the training of a policy discrimination model provided in one embodiment of this specification is shown. Figure 4E A schematic diagram of a policy generation model for direct preference optimization provided in an embodiment of this specification is shown; Figure 4F A flowchart illustrating the self-evolution of an intelligent processing unit according to an embodiment of this specification is shown; Figure 5 This specification shows a schematic diagram of the structure of an information processing apparatus according to an embodiment of the present specification; Figure 6 This specification shows an architecture diagram of an information processing system provided in one embodiment; Figure 7 A structural block diagram of a computing device provided according to an embodiment of this specification is shown. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0017] The information processing methods provided in this manual can be applied to scenarios involving information filtering. For example, in scenarios where target objects are selected on a content application platform, these target objects could be users carrying user attribute information or published content containing feature tags, etc.

[0018] It is important to understand that the filtering described in this manual can also be a type of selection operation. That is, this selection operation is a filtering or clustering operation based on specific characteristics of the objects to be selected. For ease of understanding, this manual uses the selection of target users on a content application platform as an example to explain the information processing method.

[0019] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0020] Content Application Platform: This is an application platform that integrates various types of content. The content application platform provides users with a rich variety of multimedia content, including but not limited to live streaming, on-demand video, audio, text and image information, social interaction, shopping, etc.

[0021] Published Content: This refers to the content service provided to the content application platform. Published content consists of any content pre-published by the recommending account on the platform. For example, the published content could be a product recommendation note published by the recommending account on the platform. This product recommendation note can include product images, videos, descriptions, advantages and disadvantages, etc. The product recommendation note can be text and image content, text information, videos, etc. The content platform's server stores the published content from each recommending account and sends it to browsing accounts on the content platform through a content distribution algorithm deployed on the server.

[0022] "Responding to": "Responding to" generally refers to a reaction or response to a situation, problem, request, or event. In the technical field, it can refer to a system's reaction to input information or events. It can be understood as the system or software's perception and recognition of user input or operations, and the resulting corresponding actions or behaviors.

[0023] Triggered actions refer to actions such as clicking, swiping, or other interactive methods on content application platforms, published content, live streaming pages, and other display interfaces. These actions can trigger corresponding events and display the page or information associated with the triggered action. For example, clicking the recommendation information block on a product details page can display a list of recommended accounts.

[0024] Controls are objects in forms, reports, or data access pages used to display data, perform operations, or serve as decoration. They are integral parts of software; for example, tables, reports, and network communication elements frequently used in software can all directly utilize controls.

[0025] Token: The basic discrete unit for the model to process and measure text, which may correspond to a single character, subword, punctuation mark, or special symbol. The input text is segmented into a token sequence by a token segmenter, and the model encodes, predicts, and generates tokens at the token level. Context length, billing, and computational overhead are usually measured in terms of the number of tokens.

[0026] Attention mechanism: A computational mechanism used in sequence modeling to dynamically assign "attention weights," enabling the model to aggregate information from different positions in the sequence based on relevance when processing the representation of a particular position. It typically manifests as a weighted aggregation of a set of candidate contexts, thereby highlighting content more relevant to the current task and suppressing irrelevant content.

[0027] Transformer: A type of neural network model architecture with attention mechanism at its core, commonly used for sequence modeling tasks such as natural language processing. It models the dependencies between positions in a sequence in parallel through self-attention, and excels at capturing long-range context. A typical structure consists of multiple stacked multi-head self-attention and feedforward networks, combined with residual connections and layer normalization to stabilize training. It usually also introduces positional encoding to represent the order information of tokens.

[0028] With the continuous development of artificial intelligence (AI) technology, it is rapidly iterating and becoming increasingly widespread. More and more online operations are leveraging AI capabilities to conduct refined user operations. In business operation scenarios, operators often combine specific rules and pre-set assumptions of various business scenarios to accurately match platform user groups with corresponding product services and content resources. Based on this, they promote further personalized and precise user operations, comprehensively improving the targeting and effectiveness of operations. However, many problems urgently need to be addressed during the operation process.

[0029] Different business operation scenarios involve a large number and variety of user filtering tags, covering multiple dimensions such as basic user attributes, behavioral habits, consumption preferences, and activity status. However, operations personnel lack a systematic, comprehensive, and clear understanding of the specific definition, filtering logic, applicable scenarios, and correct usage of each sub-tag, making it difficult to accurately distinguish the applicable scope and value of different tags. Therefore, in the process of batch selecting tags and filtering target users, in order to avoid errors and simplify operations, they often only choose general tags that are universal and applicable to all scenarios. This leads to serious homogenization among the target users filtered by different business scenarios, and each business line is unable to filter out precise user groups that match the characteristics of its own scenario. This makes subsequent operations such as tiered operations, targeted push notifications, and activity outreach lose their specificity, significantly reducing the overall operational effectiveness and failing to achieve the goal of refined and differentiated user operations.

[0030] Furthermore, intelligent processing units can be used to filter based on specific filtering needs and tags corresponding to different business scenarios. However, when intelligent processing units break down user filtering needs and formulate corresponding filtering strategies, they are prone to problems such as arbitrary logical constructs and subjective, divergent judgments, failing to strictly align with the user's actual needs. Specifically, in the process of automatically generating user filtering strategies, the intelligent model generates a large number of irrelevant additional associations, adding many filtering conditions that do not conform to the core requirements of this filtering or the business scenario, resulting in redundant, chaotic, and significantly reduced accuracy in the final generated filtering strategies. On the other hand, the model often fails to fully break down the needs, omitting many key filtering conditions and core requirements from the user's input. The final generated strategies cannot fully match the user's actual filtering needs and are unlikely to achieve the expected filtering effect. Therefore, the filtering operation performed by the intelligent processing unit, and the determination of filtering strategies based on filtering needs information, is prone to unrealistic assumptions.

[0031] Furthermore, when generating relevant strategies, on the one hand, additional associations may arise, leading to some conditions in the filtering strategies that do not meet the filtering requirements; on the other hand, some important conditions in the input of the filtering requirements may be overlooked, failing to meet the filtering needs. Moreover, during the generation of relevant strategies, it is impossible to judge the quality of a generated strategy, resulting in inconsistent quality of the final generated strategies, requiring users to confirm them again. In addition, with continuous use, on the one hand, new tags may be added; on the other hand, over time, after repeated use, the old filtering strategies learned by the large model become too homogeneous in the target audience, reducing the effectiveness of targeted advertising, but it is not yet able to update its filtering strategies automatically.

[0032] In view of this, this specification provides an information processing method that determines the corresponding target constraints and at least one candidate screening strategy based on target screening requirement information. Based on this, it analyzes each candidate screening strategy using the target constraints according to a strategy discrimination model to determine the strategy adjustment attribute information corresponding to each candidate screening strategy. Based on the strategy adjustment attribute information, it updates each candidate screening strategy to obtain the target screening strategy, and then filters the target object according to the target screening strategy, thereby improving the efficiency and accuracy of target object screening. This specification also relates to an information processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0033] Figure 1 A flowchart of an information processing method according to an embodiment of this specification is shown, which specifically includes the following steps: Step 102: Obtain target screening requirement information and determine the target screening semantics corresponding to the target screening requirement information; based on the target screening semantics, determine the target constraints corresponding to the target screening requirement information.

[0034] In this context, target screening requirement information can be understood as a natural language description representing the characteristics of the objects to be selected. For example, when the selection object is a target population, the target screening requirement information could be the characteristics of the target population. It should be understood that the target screening requirement information mentioned in this specification can also be understood as a task requirement information, specifically a selection task. That is, target screening requirement information can also be understood as target selection requirement information; the two are used interchangeably, but their meanings should be understood to be consistent.

[0035] Target constraints can be understood as structured limitations extracted from the target selection requirements. They include at least one constraint dimension, with varying degrees of constraint strength. For example, they may include strong constraints (must) that must be met, soft constraints (should) that allow for optional preferences, and mandatory constraints (forbidden).

[0036] In one specific embodiment provided in this specification, textual description information is obtained and used as target filtering requirement information; alternatively, voice information is obtained and converted into textual description information, which is then used as target filtering requirement information. Based on this, semantic parsing is performed on the textual description information to extract the entity information and intent information contained therein, and then target constraints are generated based on the entity information and intent information.

[0037] To facilitate understanding, this manual explains the method for determining the target constraints corresponding to the target selection requirements in the following way.

[0038] In one specific embodiment provided in this specification, the target constraints corresponding to the target filtering requirement information are determined based on the target filtering semantics, including: Based on the target filtering semantics, extract constraint features of at least one constraint dimension; Based on the characteristics of each constraint, the target constraint conditions are determined.

[0039] In this context, target filtering semantics can be understood as the semantic content extracted from filtering requirement information. This semantic content includes entity information and intent information. Examples include user intent recognition, attribute condition recognition, and time / spatial range recognition.

[0040] Constraint dimensions can be understood as information used to characterize the strength of constraints. For example, they can include strong constraints (must) that must be met, soft constraints (should) that allow for optional preferences, and mandatory constraints (forbidden) that prohibit certain conditions.

[0041] In one specific implementation, strong constraint features can be understood as necessary conditions explicitly defined by the user. Soft constraint features can be understood as the user's expressed tendencies and preferences. Prohibited mandatory constraint conditions can be understood as the user's explicit or implicit scope boundaries. Furthermore, the constraint features for each dimension can be determined based on keywords in the extracted semantic content. For example, restrictions containing words such as "only" or "limited to" are identified as prohibited mandatory constraint conditions.

[0042] Constraint features can be understood as features extracted from the filtering semantics, with different constraint dimensions corresponding to different constraint features.

[0043] In one specific embodiment provided in this specification, target filtering requirement information is obtained, and the target filtering semantics corresponding to the target filtering requirement information are determined. Entity information and intent information are extracted from the target filtering semantics to facilitate the determination of user intent recognition, attribute condition recognition, and time / space range recognition, etc. Based on the extracted target filtering semantics, constraint features are determined, including strong constraints (must) that must be met, soft constraints (should) that include optional preference conditions, and mandatory constraints (forbidden) that include prohibited conditions. Target constraint conditions are determined based on each constraint feature.

[0044] According to a specific implementation method provided in this specification, by performing structured parsing of the screening requirements, ambiguous natural language is converted into constraints, laying the foundation for subsequent strategy generation and discrimination, and avoiding semantic ambiguity caused by directly using the original text. Furthermore, the constraints are refined and classified according to different constraint dimensions, allowing for differentiated processing of constraint features of different natures in subsequent processes.

[0045] Furthermore, the target constraints corresponding to the target selection requirement information defined in this specification also include: The target screening requirements are input into the feature extraction model; Based on the feature extraction model, constraint features of at least one constraint dimension corresponding to the target screening requirement information are extracted, and target constraint conditions are generated based on each constraint feature.

[0046] The feature extraction model can be understood as a large language model trained under supervised fine-tuning, used to extract constraints of at least one constraint dimension from the screening requirements. It's important to understand that the feature extraction model possesses vertical domain semantic understanding capabilities.

[0047] Optionally, the feature extraction model can be trained by acquiring a filtering strategy dataset and extracting the corresponding object labels and enumeration values ​​from each filtering strategy based on the dataset.

[0048] The filtering strategy data includes descriptions of real users on the platform and their corresponding filtering strategies. Furthermore, the filtering strategy dataset is an updatable dataset, determined based on existing data in the current application scenario and periodically updated data. It should be noted that the data in the current application scenario can be updated at set intervals.

[0049] Based on this, the annotators assign constraint features of three dimensions to each extracted object label or condition, according to semantic strength, logical necessity, and scope boundaries: "must" (strong constraint): The conditions that this group of people must meet; omitting them would completely violate the user's intent.

[0050] Should (soft constraint): Conditions that can be prioritized but are not mandatory, allowing for reasonable supplementation or adjustment within the strategy.

[0051] Forbidden (prohibits externalization): The boundaries that users explicitly or implicitly prohibit from being extended.

[0052] Based on the above, the user description, the extracted enumeration values, and the constraint features of the three constraint dimensions are used as the filtering strategy dataset to fine-tune the general large model, resulting in the feature extraction model.

[0053] In one specific embodiment provided in this specification, based on the trained feature extraction model, the target screening requirement information is input into the feature extraction model so that, according to the feature extraction model, at least one constraint feature corresponding to the target screening requirement information is extracted, and target constraint conditions are generated based on each constraint feature.

[0054] In one specific implementation, taking the execution of a selection operation as an example, the target selection requirement information is input into the feature extraction model, and the feature extraction model performs the extraction of constraint features and the generation of target constraint conditions.

[0055] The feature extraction model was pre-obtained through supervised fine-tuning (SFT). Specifically, historical datasets from real-world user segmentation scenarios within the platform were acquired. Each dataset contained sample segmentation requirements and their corresponding segmentation strategies. Labels and enumeration values ​​were extracted from the segmentation strategies in each sample dataset. Annotators then applied ternary annotations to the extracted labels and enumeration values, assigning them strong constraints (must), soft constraints (should), or forbidden extrapolation based on their semantic role in the user requirements. Using the sample segmentation requirements as input and the annotated labels and enumeration values ​​along with their respective constraint dimensions as output targets, a general language model was subjected to supervised fine-tuning to obtain the feature extraction model.

[0056] According to a specific implementation method provided in this specification, a trained feature extraction model is used for constraint extraction, avoiding the inefficiency and incompleteness of manually written rules. Furthermore, the feature extraction model is supervised and fine-tuned based on real user descriptions and corresponding filtering strategies within the platform, enabling it to understand the unique filtering business context and tagging system in different business scenarios, thus making the extracted constraints more closely aligned with actual business needs.

[0057] Step 104: Based on the target screening requirement information, generate at least one candidate screening strategy.

[0058] Here, the candidate filtering strategy can be understood as a specific filtering logic scheme, which can be composed of object labels, enumeration values, and logical operators. The candidate filtering strategy can be extracted based on keywords corresponding to the target filtering requirements, or it can be generated based on a strategy generation model. It should be noted that, in order to balance the accuracy of the target filtering strategy with the efficiency of filtering strategy generation, this specification can generate 3-5 candidate filtering strategies.

[0059] In one specific implementation provided in this specification, at least one candidate screening strategy is determined based on the target screening requirements information and in conjunction with reference documents.

[0060] In one specific embodiment provided in this specification, at least one candidate filtering strategy is generated based on the target filtering requirement information, including: The target screening requirements information is input into the strategy generation model; The target filtering semantics are determined in the strategy generation model, and at least one candidate filtering strategy output by the strategy generation model is generated based on the target filtering semantics.

[0061] The strategy generation model can be understood as a large language model trained in two stages: Supervised Fine-tuning (SFT) and Direct Preference Optimization (DPO). This model is used to generate filtering strategies based on filtering requirements. The SFT stage enables the model to learn the filtering business logic corresponding to different business scenarios, while the DPO stage enables it to internalize user preferences.

[0062] For ease of understanding, the training process of the strategy generation model is described in the following manner in this manual.

[0063] In one specific implementation, the initial policy generation model is supervised and fine-tuned using various screening requirement information to enable it to possess semantic understanding capabilities. The supervised and fine-tuned initial policy generation model is then further trained using direct preference optimization, employing positive and negative sample pairs to enable the model to generate the final screening policy.

[0064] In one specific implementation provided in this specification, target filtering requirement information is input into the strategy generation model. The strategy generation model first performs semantic understanding to identify user intent. Then, combining the filtering business knowledge and preferences learned by the strategy generation model through SFT and DPO, candidate filtering strategies are generated.

[0065] In one specific implementation, a preset tag knowledge base is obtained, wherein the preset tag knowledge base contains all tag information currently valid on the platform. It should be understood that this preset tag knowledge base can be the same as the selection strategy dataset.

[0066] In response to the strategy generation model generating strategy condition combinations, a pre-defined tag knowledge base is invoked to anchor the tags of each selection condition in the strategy condition combination, limiting each selection condition to existing tag identifiers in the pre-defined tag knowledge base. The target selection requirement information is input into the strategy generation model, which outputs at least one strategy condition combination, and each strategy condition combination is treated as a candidate selection strategy. The selection strategy generation model is built based on a large language model and is used to parse the target selection requirement information into a structured strategy combination containing at least one selection condition and its logical relationships.

[0067] Furthermore, it is also possible to obtain a set of historical selection strategies and calculate the matching degree between the target selection requirement information and each historical selection strategy. The at least one historical selection strategy with the highest matching degree is determined as a candidate selection strategy, or the at least one historical selection strategy with the highest matching degree is combined and mutated with conditions to become a candidate selection strategy.

[0068] In one specific embodiment provided in this specification, at least one candidate selection strategy is generated based on target selection requirement information. Specifically, the target selection requirement information is input into a pre-trained strategy generation model, which then performs semantic understanding and strategy generation operations.

[0069] The strategy generation model is pre-trained through a two-stage process. The first stage is Supervised Fine-Tuning (SFT): Historical data from real-world user targeting scenarios within the platform are collected to construct a training dataset. Each data point contains a sample targeting requirement and its corresponding manually labeled or historically adopted targeting strategy. This training set is used to supervise the fine-tuning of the general language model, enabling the model to learn the platform's unique user targeting business knowledge, tag usage logic, and strategy generation paradigm. The second stage is Direct Preference Optimization (DPO): Based on the same batch of data, candidate strategies are generated from the SFT-adapted model and a DPO training set is constructed using preference labeling. The model is then further trained using the Direct Preference Optimization algorithm, internalizing user quality preferences for targeting strategies into the model parameters.

[0070] In the actual reasoning process, after obtaining the target selection requirement information input by the user, this information is input into the policy generation model in the form of natural language text. Optionally, before inputting into the policy generation model, relevant tag information from a preset tag knowledge base, a list of currently valid tag enumeration values ​​on the platform, and a reference document consisting of historical successful selection examples are associated and combined with the target selection requirement information to form a policy generation input containing constraint guidance information. This preset tag knowledge base maintains its timeliness by synchronizing with the latest tag information from the platform daily, ensuring that all tags used by the policy generation model are currently existing and valid tags on the platform.

[0071] After receiving the target selection request information, the strategy generation model first performs contextual semantic understanding on the input information through an encoding layer to determine the target selection semantics corresponding to the target selection request information. The target selection semantics includes: identifying the user's core selection intent, extracting the attribute conditions explicitly mentioned by the user, identifying time range information and spatial range information, as well as the logical relationships between the conditions.

[0072] Based on this, the strategy generation model performs inference and generation operations based on the determined target selection semantics, combined with the platform's user segmentation business knowledge learned in the SFT stage, user preferences internalized in the DPO stage, and the tag range constraints and historical example reference guidance from the pre-set tag knowledge base in the Reference document. Specifically, in its internal decoding layer, the strategy generation model maps each condition extracted from the target selection semantics to the corresponding tags and enumeration values ​​in the platform's tag system, and combines them into a structured selection strategy expression according to logical relationships, ultimately outputting at least one candidate selection strategy.

[0073] In a preferred embodiment of this implementation, the policy generation model is configured to generate multiple candidate selection policies at once during inference, for example, 3 to 5 candidate selection policies, to provide a candidate policy pool for subsequent selection and iterative optimization. Specifically, during the generation process, the policy generation model obtains multiple different policy outputs by sampling from the decoded probability distribution by setting different temperature parameters or using a beam search strategy. Each candidate selection policy maintains the core selection intention... Figure 1 Under the premise of consistency, reasonable differences exist in label selection, condition combination methods, or condition leniency to provide diverse selection schemes for subsequent discrimination model evaluation.

[0074] In another optional implementation, if the strategy generation model identifies that certain conditions in the target selection requirements do not have exact corresponding tags in the preset tag knowledge base during the generation of candidate selection strategies, the strategy generation model, guided by the historical selection examples, selects semantically similar available tags for substitution and annotates this substitution in the generated candidate selection strategy so that the user or the discrimination model can perform targeted evaluation of the substitution operation in subsequent processes. This implementation can effectively reduce the problem of strategy unavailability caused by knowledge gaps, while maintaining the transparency and interpretability of the strategy generation process.

[0075] In another optional implementation, if the target selection requirement information received by the strategy generation model is text converted from speech input, text correction and semantic completion operations are performed before inputting it into the strategy generation model to improve the model's tolerance to non-standard input. Specifically, a preset text correction module corrects typos generated by speech recognition, and a semantic completion module supplements any missing implicit subjects or default conditions in the speech input, forming complete target selection requirement information that meets the input format requirements of the strategy generation model, which is then input into the strategy generation model for subsequent processing.

[0076] Through the above implementation methods, the strategy generation model can accurately transform fuzzy natural language selection requirements into structured candidate selection strategies that conform to the platform's labeling system and are practically executable. This not only ensures the professionalism and usability of the strategy but also provides diverse initial solutions for subsequent strategy discrimination and iterative updates.

[0077] According to a specific implementation method provided in this specification, the policy generation model undergoes two-stage training: SFT and DPO. This not only possesses the semantic understanding capabilities of a general-purpose large-scale model but also internalizes business knowledge and user preferences from different business scenarios. Simultaneously, DPO training enables the policy generation model to distinguish between good and bad policies, reducing the number of subsequent iterations. Furthermore, the policy generation model can generate multiple candidate policies in a single inference, providing a comparison space for subsequent discrimination and selection.

[0078] Furthermore, given the candidate screening strategies and target constraints, they are input into the strategy discrimination model in parallel, so that the accuracy and optimization direction of each candidate screening strategy can be determined based on the strategy discrimination model.

[0079] Step 106: Based on each candidate screening strategy and the target constraints, determine the strategy adjustment attribute information corresponding to each candidate screening strategy.

[0080] In one specific embodiment provided in this specification, based on each candidate screening strategy and the target constraint, the strategy adjustment attribute information corresponding to each candidate screening strategy is determined, including: Determine an initial candidate screening strategy, wherein the initial candidate screening strategy is any one of the candidate screening strategies; Input prompt words are constructed based on the initial candidate selection strategy and the target constraints; The input prompts are fed into the policy discrimination model to obtain the policy adjustment attribute information corresponding to each candidate filtering policy output by the policy discrimination model.

[0081] The policy discrimination model can be understood as a model trained on a large language model to evaluate the quality of candidate selection policies. The input to this model is the candidate policy and the target constraints, and the output is policy adjustment attribute information. It is important to understand that the policy discrimination model described in this specification has a dual-head architecture, including a shared semantic encoder, a text generation module, and multiple scoring task modules. Specifically, this dual-head architecture refers to the parallel processing architecture between the text generation module and the scoring task modules.

[0082] For ease of understanding, the training process of the policy discrimination model is explained in the following manner in this manual.

[0083] In one specific embodiment provided in this specification, at least one filtering strategy output by the strategy generation model, the target constraints output by the feature extraction model, and the data in the preset label knowledge base are used as inputs to the initial strategy discrimination model.

[0084] In the initial policy discrimination model, each selection policy is semantically encoded and aligned with the target constraints using a semantic encoder to obtain the policy features corresponding to each selection policy.

[0085] Using the text generation module, natural language text corresponding to each filtering strategy is generated based on the strategy characteristics of each filtering strategy. The content includes: what specific problems exist, the location of the problems, and clear modification suggestions.

[0086] Simultaneously, the strategy features corresponding to each filtering strategy are input into the scoring task module to determine the strategy score for each filtering strategy. Specifically, based on the intent drift detection head (Head 1), it is calculated whether the filtering strategy has added additional conditions outside the forbidden list, or whether semantic extrapolation has been performed on the must condition. The intent drift probability is output.

[0087] Based on the knowledge gap detection head (Head 2), each tag / enumeration value in the filtering strategy is hard-matched against a preset tag knowledge base. The knowledge gap probability is output based on whether the filtering strategy references a tag that does not exist in the knowledge base.

[0088] Based on the constraint missing detection header (Head 3), check one by one whether all constraints in the must list appear in the filtering strategy; check whether forbidden constraints are violated. Output the constraint missing probability according to the severity of the missing / violation.

[0089] The probability of intent drift, the probability of knowledge gaps, and the probability of constraint loss corresponding to each screening strategy are weighted and fused to obtain the strategy score corresponding to each screening strategy.

[0090] Strategy adjustment attribute information can be understood as evaluation information corresponding to candidate screening strategies. Based on this information, the optimization direction of the candidate screening strategy can be clarified, ensuring that the optimized strategy further meets the target screening requirements. Furthermore, strategy adjustment attribute information can include a textual description of the problem and modification suggestions, or it can include strategy scores corresponding to each constraint dimension in the target constraints.

[0091] In one specific embodiment provided in this specification, the candidate filtering strategies corresponding to the target filtering requirement information and the target constraints corresponding to the target filtering requirement information are input into the strategy discrimination model. In the strategy discrimination model, based on a shared semantic encoder, a text generation module, and multiple scoring task modules, the text information and strategy score corresponding to each candidate filtering strategy are determined. The text information and strategy score corresponding to each candidate filtering strategy are then defined as the strategy adjustment attribute information corresponding to each candidate filtering strategy.

[0092] To facilitate understanding, this manual explains the process of determining the strategy adjustment attribute information corresponding to each candidate screening strategy based on the strategy discrimination model in the following manner.

[0093] In one specific embodiment provided in this specification, the policy discrimination model includes a semantic encoder, a text generation module, and at least one scoring task module.

[0094] The semantic encoder can be understood as the shared backbone part of the policy discrimination model, typically a Transformer encoder, responsible for jointly semantically encoding the input candidate selection policies and constraints. Its output is a shared policy feature vector used by downstream tasks, possessing the ability to understand the semantics of the judgment task.

[0095] The text generation module can be understood as the LM Head (Language Model Head) of the policy discrimination model, a decoder that generates natural language text based on policy features. Its output includes: a description of the specific problems with the policy; and targeted modification suggestions. Its generation process is influenced by the autoregressive sampling mechanism of the large model and has a certain degree of randomness.

[0096] The scoring task module can be understood as the MLP Heads (Multilayer Perceptron Task Heads) of the policy discrimination model, used to output probability scores for each constraint dimension based on policy features. Each task head is an independent MLP network, trained independently, with the policy feature vector as input and continuous probability values ​​from 0 to 1 as output. Specifically, it includes an intent drift detection head, a knowledge gap detection head, and a constraint missing detection head. Its score is output through a regression path, without sampling randomness.

[0097] Each candidate selection strategy and the target constraint are input into the strategy discrimination model to obtain the strategy adjustment attribute information corresponding to each candidate selection strategy output by the strategy discrimination model, including S1062-S10610: S1062. Determine the initial candidate selection strategy, wherein the initial candidate selection strategy is any one of the candidate selection strategies.

[0098] The initial candidate selection strategy can be understood as any one of the candidate selection strategies. It should be understood that while this specification allows for parallel processing of the candidate selection strategies, for ease of explanation, only one strategy is used as an example here.

[0099] S1064. Input the initial candidate selection strategy and the target constraint into the semantic encoder for feature extraction to obtain the initial strategy features output by the semantic encoder.

[0100] The initial policy feature can be understood as the vector representation output by the semantic encoder after jointly encoding the input. This feature is the common input of the text generation module and the scoring task module. The two modules process the same feature obtained from the same forward computation in parallel, thereby ensuring the semantic consistency between text generation and scoring.

[0101] In one specific implementation provided in this specification, the complete logical expression of the initial candidate strategy and the complete Anchor Set of the target constraints are input to the semantic encoder. The semantic encoder performs joint semantic encoding and alignment on the two to understand which constraint in the anchor point each condition in the strategy corresponds to, identify their semantic relationship and logical consistency, and thus output the initial strategy features. It should be noted that the initial strategy features contain all semantic association information between the initial selection strategy and the constraints.

[0102] S1066. Input the initial strategy features into the text generation module, and generate initial text information containing problem description and modification suggestions based on the text generation module.

[0103] In one specific implementation provided in this specification, the initial policy features are fed into the LM Head, and natural language text is generated word by word through an autoregressive method. Modification suggestions corresponding to different constraint dimensions output by the LM Head, as well as problem description information, are obtained, thereby obtaining the initial text information.

[0104] S1068. Input the initial policy features into each scoring task module, and determine the policy score for each constraint dimension based on each scoring task module.

[0105] In one specific embodiment provided in this specification, initial policy features are fed into multiple independent MLP task heads. Each independent MLP task head is executed in parallel, outputting continuous probability values ​​through a regression path, without sampling randomness. The probabilities output by different independent MLP task heads are obtained, and thus the policy score is derived.

[0106] In one specific embodiment provided in this specification, the scoring task module is any one of the intent drift detection head, knowledge gap detection head, and constraint missing detection head.

[0107] The intent drift detection header can be understood as detecting whether a candidate policy has experienced "strategy drift". In this specification, the intent drift detection header can be used to detect whether the policy has added an additional condition of prohibition (forbidden), and also to detect whether the policy has performed semantic extrapolation on the must condition. The output P(intent drift) ∈ [0,1], with higher values ​​indicating more severe intent drift.

[0108] The knowledge void detection head can be understood as detecting whether a strategy references tags or enumeration values ​​that do not exist on the content publishing platform ("knowledge void" / illusion). In this specification, the knowledge void detection head can perform a hard match between each tag name used in the strategy and a preset tag knowledge base. The output P(knowledge void) ∈ [0,1], with higher values ​​indicating more severe knowledge voids.

[0109] The constraint missing detection head can be understood as detecting whether the strategy fully covers the must constraints and whether the forbidden constraints are violated. In this specification, the constraint missing detection head can detect whether each constraint in the must list appears in the strategy, and it can also detect whether the strategy contains any prohibited conditions listed in the forbidden list. The output P(constraint missing) ∈ [0,1], with higher values ​​indicating more severe constraint missing.

[0110] Optionally, the policy score for each constraint dimension is determined based on each scoring task module, including: Based on the intent drift detection head, it detects whether the initial candidate screening strategy meets the set semantic boundary and outputs the intent drift probability, wherein the set semantic boundary is determined based on the constraint features corresponding to each constraint dimension in the target constraint conditions.

[0111] The semantic boundary setting can be understood as the range of legal policy generation defined by the target constraints (Anchor Set). Must-class constraints define the minimum requirements that a policy must meet; forbidden-class constraints define the red line that a policy must not cross. Policies exceeding these boundaries are considered to have "intent drift." This boundary is used as a scoring benchmark for the intent drift detection head through constraint feature vectorization.

[0112] In one specific implementation provided in this specification, for the initial policy features, it is compared whether the conditions corresponding to the initial policy features have corresponding prohibited items in the forbidden list. The must class is checked to see if it is a bottom-line condition. The overall semantics of the initial candidate selection policy are evaluated to see if it deviates from the semantic boundary corresponding to the target constraint (Anchor Set). Based on this, the intent drift probability is obtained.

[0113] Based on the knowledge void detection head, it detects whether the initial candidate screening strategy meets the preset label knowledge base and outputs the knowledge void probability.

[0114] The preset tag knowledge base can be understood as containing all tags on the platform, tag descriptions, and enumerated values, and is automatically updated through a daily scheduled task. This knowledge base is used for knowledge gap detection; tags appearing in candidate strategies that do not belong to the preset tag knowledge base are judged as "knowledge gaps".

[0115] In one specific implementation provided in this specification, all used tag names and enumeration values ​​are parsed from the initial screening strategy; each tag is matched precisely or synonymously with entries in a preset tag knowledge base; if a tag does not exist in the preset tag knowledge base, the knowledge void confidence is accumulated. Based on this, the knowledge void probability is obtained.

[0116] Based on the constraint missing detection head, it detects whether the initial candidate screening strategy meets the target constraint conditions and outputs the constraint missing probability.

[0117] In one specific implementation provided in this specification, the initial policy features and target constraints are input into the constraint missing detection head. The head checks each constraint in the `must` list to see if it is covered by the policy. If any `must` conditions are missing, their severity is assessed and the degree of absence is quantified. Simultaneously, the head checks for prohibited conditions in the `forbidden` list within the policy. Based on this, the constraint missing probability is obtained.

[0118] It is important to understand that the constraint missing detection head and the intent drift detection function overlap but have different focuses: the constraint missing detection head focuses on "omissions", while the intent drift detection head focuses on "illegal additions".

[0119] The strategy score is determined based on the intent drift probability, the knowledge gap probability, and the constraint missing probability.

[0120] The strategy score can be understood as a comprehensive score for each candidate selection strategy obtained after evaluating each candidate strategy according to different constraint dimensions. For example, a comprehensive score determined based on three probability values. The specific values ​​of each weight can be preset according to actual business needs, or obtained through posterior data statistical learning. The lower the score, the higher the strategy quality; the higher the score, the more / more serious the problems.

[0121] According to a specific implementation method provided in this specification, three independent detection heads are used to detect three different types of policy problems, enabling precise problem localization and allowing modification suggestions to clearly pinpoint the existing issues. Furthermore, each MLP task head is trained independently, requiring only 0 / 1 binary labels for annotation data, eliminating the need for complex multi-class annotation or refined scoring annotation, thus reducing annotation costs.

[0122] S10610. The text information and the strategy score are determined as the strategy adjustment attribute information.

[0123] The strategy score can be understood as the collective probability score output by each scoring task module, which can be further weighted and averaged to obtain the comprehensive score after the output of each module.

[0124] In one specific implementation provided in this specification, the text information generated by the LM Head and the probability scores output by the MLP Heads are combined to obtain complete policy adjustment attribute information.

[0125] According to a specific implementation provided in this specification, a semantic encoder is used for feature extraction to obtain an initial candidate selection strategy and target constraints that are semantically aligned. Based on this, text generation and scoring share the same forward computation of semantic encoding, outputting in parallel to avoid redundant computation, thereby reducing inference latency and computational resource consumption. Furthermore, a strategy score is obtained based on the scoring task model, improving the credibility of the score. Building upon the above, modification suggestions output by the text generation module are used to facilitate precise modifications in subsequent implementations.

[0126] Step 108: Update each candidate filtering strategy based on the attribute information of each strategy until the update stop condition is met, and determine at least one target filtering strategy.

[0127] The update stopping condition can be understood as a criterion used to determine whether the iterative update process of each candidate screening strategy should be terminated. In this specification, the update stopping condition includes at least: the change in the overall score is less than or equal to a preset threshold; and the strategy adjustment attribute information output by the discriminant model indicates that the current strategy no longer has a problem. It should be understood that when this condition is met, the current candidate strategy can be determined as the target screening strategy.

[0128] The target selection strategy can be understood as a selection strategy that satisfies the update stopping condition after one or more rounds of iterative updates to the candidate selection strategy based on the update stopping condition.

[0129] In one specific embodiment provided in this specification, any candidate screening strategy is selected as the candidate screening strategy to be processed from among the candidate screening strategies, in order to facilitate explanation.

[0130] Based on the candidate selection strategies to be processed and their corresponding strategy adjustment attribute information, the candidate selection strategies are updated. The update method involves the strategy generation model making targeted corrections to the strategies based on the text modification suggestions from the discriminant model, such as inserting "must" missing conditions, deleting "forbidden" violations, and replacing empty labels. Furthermore, the corrected strategies are re-submitted to the discriminant model for scoring, and the difference between the scores before and after correction is calculated. The difference and the problem status determine whether to continue iterating or stop. Candidate selection strategies that meet the update stopping condition are identified as target selection strategies.

[0131] For ease of understanding, step 108 is explained in this specification in the following manner.

[0132] In one specific embodiment provided in this specification, each candidate filtering strategy is updated based on the attribute information of each strategy until the update stop condition is met, and at least one target filtering strategy is determined, including: Determine the candidate filtering strategy to be processed and the corresponding adjustment attribute information of the candidate filtering strategy to be processed, wherein the candidate filtering strategy to be processed is any one of the candidate filtering strategies; The candidate filtering strategy to be processed and the adjustment attribute information of the strategy to be processed are input into the strategy generation model to obtain the updated candidate filtering strategy output by the strategy generation model. The updated candidate filtering strategy is input into the strategy discrimination model to obtain the updated strategy adjustment attribute information corresponding to the updated candidate filtering strategy output by the strategy discrimination model; When the updated attribute information meets the update stop condition, the target filtering strategy is obtained.

[0133] Here, the candidate selection strategy to be processed can be understood as any one of the candidate selection strategies, which is the object of processing in this round of update operation. It should be noted that this specification can process each candidate selection strategy in parallel, but for the sake of explanation, any one candidate selection strategy is used as an example here.

[0134] Updating the candidate selection strategy can be understood as the strategy generation model adjusting the attribute information of the strategy to be processed to make targeted corrections to the strategy and outputting a new strategy. In this specification, correction methods include, but are not limited to, inserting missing must conditions, deleting forbidden conditions, replacing empty tags with valid tags existing in the knowledge base, and adjusting the time window.

[0135] Updating the policy adjustment attribute information can be understood as obtaining new policy adjustment attribute information after updating the candidate policy and re-feeding it into the policy discrimination model. This information includes a new problem description, modification suggestions, and new scores for each dimension and a comprehensive score. This information is used to compare with the information from the previous round to determine whether the update stopping condition is met.

[0136] The update stopping condition can be understood as a criterion used to terminate the iteration. In this specification, the update stopping condition can be understood as follows: when the absolute value of the difference in the comprehensive score is less than or equal to a preset threshold, it is determined that the strategy quality has converged (or reached the improvement bottleneck), the iteration is stopped, and the current strategy is output as the target screening strategy.

[0137] In one specific embodiment provided in this specification, one policy is selected from the candidate policy set as the target for processing in this round, and its policy adjustment attribute information generated by the policy discrimination model is obtained. The candidate policy to be processed and its policy adjustment attribute information are input into the policy generation model. The policy generation model understands the semantic intent of the modification suggestion and locates the corresponding position in the candidate policy that needs modification. Insertion, deletion, or replacement operations are then performed accordingly. Updated policy adjustment attribute information is obtained.

[0138] Based on this, it is determined whether the attribute information of the update strategy adjustment meets the update stop condition. If the attribute information of the update strategy adjustment meets the update stop condition, it is determined as the target filtering strategy.

[0139] It should be understood that, in determining the candidate selection strategy, this manual may prioritize the candidate with the highest comprehensive score or process them sequentially according to the generation order.

[0140] In one specific implementation, after the strategy discrimination model outputs the strategy adjustment attribute information corresponding to each candidate selection strategy, the candidate selection strategies are iteratively updated based on the strategy adjustment attribute information until the update stop condition is met, thereby determining at least one target selection strategy. For ease of explanation, any one of the candidate selection strategies is used as the candidate selection strategy to be processed here. In engineering implementation, each candidate selection strategy can be processed in parallel or serially in the same way.

[0141] In one specific embodiment provided in this specification, the candidate selection strategy to be processed and the corresponding adjustment attribute information of the candidate selection strategy are first determined. The candidate selection strategy to be processed is any one of the candidate selection strategies, for example, one of the 3 to 5 candidate selection strategies generated by the strategy generation model in the initial inference stage. The adjustment attribute information of the candidate selection strategy is the complete set of attribute information output by the strategy discrimination model in the previous round of discrimination, including text information containing a problem description and modification suggestions, probability scores for each constraint dimension, and a comprehensive score. It should be noted that in the first iteration, the adjustment attribute information of the candidate selection strategy is the strategy adjustment attribute information output by the strategy discrimination model when it first reviews the candidate strategy; in subsequent iterations, the adjustment attribute information of the candidate selection strategy is the updated strategy adjustment attribute information output by the strategy discrimination model at the end of the previous iteration.

[0142] In one specific embodiment provided in this specification, after determining the candidate selection strategy to be processed and its corresponding adjustment attribute information, the system inputs the candidate selection strategy to be processed and the adjustment attribute information to be processed into the strategy generation model to drive the strategy generation model to perform strategy update operation.

[0143] Optionally, the text information portion of the attribute information for the strategy to be processed includes a specific problem description and modification suggestions generated by the text generation module (LM Head) of the strategy discrimination model. When input into the strategy generation model, the modification suggestions are used as correction instructions and combined with the original logical expression of the candidate selection strategy to be processed, forming an updated input containing the "original strategy and modification instructions." Based on its semantic understanding of user segmentation business learned through the SFT stage and its user preference alignment capabilities internalized in the DPO stage, the strategy generation model performs semantic parsing and intent understanding of the modification instructions, locates the corresponding position in the original strategy that needs modification, and performs the corresponding insertion, deletion, or replacement operations. After the modification operation is completed, the strategy generation model outputs the updated candidate selection strategy.

[0144] In one specific embodiment provided in this specification, the updated candidate selection strategy output by the strategy generation model is re-inputted into the strategy discrimination model. The strategy discrimination model then reviews the updated candidate selection strategy according to the same discrimination process described above, obtaining the updated strategy adjustment attribute information corresponding to the updated candidate selection strategy. Specifically, the semantic encoder of the strategy discrimination model jointly encodes the updated candidate strategy and the target constraints, the text generation module outputs the updated problem description and modification suggestion text, and each scoring task module outputs the updated probability scores for each dimension and the updated comprehensive score. The updated comprehensive score in the updated strategy adjustment attribute information is denoted as score_t, while the comprehensive score obtained before this iteration, i.e., in the previous discrimination round, is denoted as score_{t-1}, which serves as the benchmark value for judging the subsequent update stopping conditions.

[0145] In one specific implementation provided in this specification, it is determined whether the update strategy adjustment attribute information meets the update stop condition.

[0146] Optionally, obtain the updated comprehensive score `score_t` from the updated policy adjustment attribute information and obtain the comprehensive score `score_{t-1}` from the policy adjustment attribute information to be processed, and calculate the difference between the two. The overall score is a weighted average of the probability scores for each dimension output by the three scoring task modules, calculated using preset weights. A lower overall score indicates higher policy quality. Therefore, a negative Δscore indicates that the policy quality has improved in this round of updates, while a positive Δscore indicates that the policy quality has decreased.

[0147] The absolute value of the calculated difference is compared with a preset threshold. The preset threshold can be configured according to business needs, for example, it can be set to 0.05. If it is greater than the preset threshold, it indicates that the current iteration has significantly changed the strategy quality. At this time, the system further determines whether the text information in the updated strategy adjustment attribute information indicates that the strategy still has problems: if problems still exist, the system returns to the previous step of inputting the updated candidate strategy and the updated strategy adjustment attribute information into the strategy generation model, and continues iterative optimization; if no problems exist, the update stopping condition is met, and the current updated candidate strategy is determined as the target strategy. If it is less than or equal to the preset threshold, it indicates that the change in strategy quality brought about by the current iteration has become gradual. At this time, regardless of whether the updated strategy adjustment attribute information indicates that the strategy still has problems, the update stopping condition is met.

[0148] If the strategy no longer has problems, it will converge naturally. If problems still exist, it will be determined that the self-evolution bottleneck has been reached, and the current candidate selection strategy will be determined as the target selection strategy. At the same time, if the bottleneck is reached, a prompt message will be output through the user interface to wait for manual intervention from the user.

[0149] In another specific embodiment provided in this specification, when multiple candidate selection strategies are independently iterated and updated, at least two candidate selection strategies may meet the update stopping condition. In this case, the comprehensive strategy score corresponding to each candidate selection strategy is obtained, and the candidate selection strategy with the lowest comprehensive strategy score (i.e., the best strategy quality) is selected as the target selection strategy. In another optional embodiment where multiple strategies need to be output for users to conduct A / B testing or manual review, the system sorts the candidate strategies from low to high according to their comprehensive strategy scores, selects the top two or three strategies as the target selection strategies, and determines the final target selection strategy used for the selection operation based on the user's selection or system configuration. The above-mentioned comparison and selection operation of multiple strategies further improves the flexibility of system output and the upper limit guarantee of strategy quality.

[0150] According to a specific implementation provided in this specification, precise corrections are made based on the specific modification suggestions given by the policy discrimination model, rather than random regeneration. This reduces the number of iterations from aimless retries to targeted adjustments, improving convergence speed. Furthermore, each candidate policy can be iteratively updated independently without interference, facilitating parallel computing and multi-policy comparison and selection, while also preventing modifications to one policy from affecting others. Moreover, the update termination condition ensures that iteration terminates within a reasonable number of rounds, keeping computational resource consumption within a preset range and preventing infinite loops caused by abnormal situations.

[0151] In one specific embodiment provided in this specification, the method for determining whether the update strategy adjustment attribute information meets the update stop condition includes: Obtain the update strategy score corresponding to the update strategy adjustment attribute information, and obtain the strategy score corresponding to the strategy adjustment attribute information to be processed; Based on the update strategy score and the difference between the strategy scores, it is determined whether the update strategy adjustment attribute information meets the update stop condition.

[0152] The updated policy score can be understood as the comprehensive score output by the policy discrimination model after updating the candidate policy, i.e., the score value of this iteration. The lower the value, the higher the quality of the policy (because the lower the probability value output by each detection head, the fewer problems there are).

[0153] The update stopping condition can be understood as an operation used to stop the loop iteration. In one case, the difference is large and a problem still exists, so iteration continues. In another case, the difference is large and there is no longer a problem, so it stops and outputs the result. In yet another case, the difference is small and there is no longer a problem; this can be understood as natural convergence, so it stops and outputs the result. In yet another case, the difference is small but a problem still exists; at this point, a bottleneck has been reached, so it stops and waits for user feedback.

[0154] In one specific embodiment provided in this specification, after the strategy generation model updates the candidate selection strategy based on the text modification suggestions in the strategy adjustment attribute information, it determines whether the updated strategy meets the update stopping condition to decide whether to continue iterative optimization or terminate the update and output the current strategy. The following provides a detailed explanation of this determination process.

[0155] In a specific embodiment provided in this specification, the update strategy score corresponding to the update strategy adjustment attribute information is obtained, and the strategy score corresponding to the strategy adjustment attribute information to be processed is also obtained. As can be seen from the above, after reviewing any candidate selected strategy, the strategy discrimination model's three independent MLP scoring task modules will output the probability scores of the strategy on each constraint dimension, namely, the intent drift probability P(drift), the knowledge void probability P(void), and the constraint missing probability P(missing). The three are then weighted and averaged to obtain the comprehensive score of the strategy. The preset weights corresponding to each dimension can be flexibly configured according to the tolerance of different business scenarios for the three types of problems.

[0156] A lower overall score indicates fewer problems and higher overall quality across all dimensions; a higher overall score indicates more or more serious problems. Based on this scoring mechanism, the overall score corresponding to the updated candidate selection strategy is extracted from the updated strategy adjustment attribute information and denoted as score_t; simultaneously, the overall score corresponding to the candidate selection strategy before the update is extracted from the unprocessed strategy adjustment attribute information and denoted as score_{t-1}.

[0157] In one specific embodiment provided in this specification, after obtaining the two policy scores mentioned above, the difference between the updated policy score and the policy score to be processed is calculated, i.e. This difference quantitatively reflects the actual direction and magnitude of the changes in strategy quality brought about by this round of strategy adjustments.

[0158] In one scenario, when Δscore is negative, it indicates that the overall score after the update is lower than before the update, meaning that the policy quality has been improved.

[0159] In another scenario, when Δscore is positive, it indicates that the updated overall score has increased compared to the previous score, implying a decrease in policy quality.

[0160] In another scenario, when Δscore is zero or extremely small, it indicates that the current update has had little impact on strategy quality. Based on the absolute value and direction of this difference, combined with a preset threshold, the system determines whether the updated strategy adjustment attribute information meets the update stop condition. This preset threshold can be configured according to actual business needs. For scenarios with high strategy quality requirements, a smaller threshold (e.g., 0.02) can be set to pursue more sufficient quality convergence; for scenarios with high efficiency requirements, a larger threshold (e.g., 0.08) can be set to complete optimization within fewer iterations.

[0161] In another specific embodiment provided in this specification, the difference is smoothed before calculation, for example, by using an exponential moving average to smooth the score changes over multiple consecutive rounds, in order to avoid misjudgments caused by abnormal fluctuations in a single strategy adjustment. Specifically, the system records the score differences of the most recent three rounds and performs a weighted average to reduce the impact of outliers in a single round on the termination judgment, further improving the stability and reliability of updating the stop condition judgment.

[0162] In summary, through the aforementioned judgment mechanism based on the difference between the updated policy score and the score of the policy to be processed, the system can accurately determine the optimal termination time for policy iteration in an objective and quantitative manner. This avoids outputting insufficiently optimized, low-quality policies due to premature termination, and also prevents wasting computational resources due to excessively late termination, effectively ensuring a good balance between optimal quality and computational resource consumption in the final target selection policy. Furthermore, this judgment mechanism combines dual verification of the difference magnitude and the problem's existence state, further improving the accuracy and robustness of the termination decision.

[0163] According to a specific implementation method provided in this specification, the decision to stop is made by using the score difference rather than subjective judgment or a fixed number of iterations, which ensures the repeatability and interpretability of the termination decision and avoids quality fluctuations caused by inconsistent human judgment standards.

[0164] In one specific embodiment provided in this specification, after judging each candidate screening strategy, at least one target screening strategy is determined in the following manner: If at least two candidate screening strategies meet the update stopping condition, determine the strategy score corresponding to each candidate screening strategy. Based on the scores of each strategy, at least one target screening strategy is determined among the candidate screening strategies.

[0165] In one specific embodiment provided in this specification, multiple candidate strategies have all independently completed iterative updates and each has reached the update termination condition. In this case, the final comprehensive score is extracted from the final strategy adjustment attribute information of each candidate strategy. In one scenario, if a single strategy needs to be output: compare the comprehensive scores of each candidate strategy and select the one with the lowest score (best quality) as the final target selection strategy. In another scenario, if multiple strategies need to be output, sort them from lowest to highest score and select the Top N (N can be configured by the user) as the target selection strategies.

[0166] In one specific implementation, when configured to output only one final strategy, the comprehensive scores of each candidate selection strategy are compared, and the candidate selection strategy with the lowest comprehensive score (i.e., the best strategy quality) is selected as the target selection strategy. This optimization mechanism ensures that the final output strategy is the best-quality solution after multiple candidate competition, rather than simply adopting the first generated strategy or randomly selecting a strategy.

[0167] In another optional implementation, when the configuration or user requirement necessitates providing multiple strategies for manual review, A / B testing, or differentiated multi-channel deployment, the candidate strategies are ranked from lowest to highest based on their overall scores, and at least two of the top-ranked strategies are selected as the target strategies. In this case, along with the strategy details and scores, auxiliary information such as the predicted audience size and tag coverage for each strategy is also presented to the user, allowing for final manual decision-making and flexible selection based on actual business needs. For example, during the campaign performance testing phase, operations personnel can select the two strategies with the highest and second-highest scores to compare their effects in two small traffic groups. After further validating the strategy effectiveness based on actual deployment data, a full rollout is then conducted. This approach, based on the system's automatic selection, further incorporates human experience and actual data verification, further improving the final deployment effectiveness of the selected strategies. This multi-strategy output mode is particularly suitable for application scenarios with complex business scenarios, high requirements for strategy diversity, or the need to verify strategy effectiveness through small-scale experiments.

[0168] According to a specific implementation method provided in this specification, when multiple available strategies exist, the strategy with the best score is selected through objective score ranking, ensuring that the final output is the highest quality among all candidates, rather than being randomly selected or selected according to the generation order. If configured to output multiple target strategies, they can be flexibly selected according to the actual needs of different business scenarios, enhancing practicality.

[0169] Step 110: Determine the target selection objects based on each target selection strategy.

[0170] In this context, the target selection object can be understood as the specific object obtained by applying the finalized target selection strategy to a database corresponding to a specific business scenario. For example, it could be the specific user group obtained by selecting from the platform's user database.

[0171] In one specific implementation provided in this specification, a target filtering strategy is determined, each target filtering strategy is converted into an executable query statement, and a filtering and matching operation is performed in the target database / data warehouse to obtain the filtered target objects.

[0172] Furthermore, the target objects obtained through the above screening process can be further applied to downstream business operations such as precise advertising, personalized content delivery, remote device control, and product recommendation ranking.

[0173] According to a specific implementation method provided in this specification, by explicitly extracting constraint features from three constraint dimensions and using these features as the scoring benchmark for the discriminative model, policy drift is effectively prevented, ensuring that the final policy faithfully reflects the user's original intent. Furthermore, a closed-loop iterative mechanism—based on the generation of candidate policies, discrimination, and adjustments based on the discrimination results—continuously improves policy quality until convergence, avoiding a situation where a single generation cannot be improved. Moreover, by addressing intent drift, knowledge gaps, and constraint deficiencies, policy quality is transformed from subjective perception into objective numerical values, allowing operators to intuitively understand the specific performance of the policy across each dimension. Therefore, the information generation method described above in this specification improves the efficiency and accuracy of target object selection.

[0174] To facilitate understanding, this manual uses the example of target user segmentation on a content application platform to explain the information processing methods mentioned above.

[0175] It should be clarified that the target user in this example is the user to be selected, and is not the same as the user shown in steps 102-110 above in this specification. The user shown in steps 102-110 can be understood as the operations personnel in this example.

[0176] In one specific implementation provided in this specification, a preset tag library is constructed.

[0177] Specifically, since the labels used by the strategy generation model during the selection process must be strictly limited to the existing labels on the platform, it is necessary to collect and organize the existing labels, label descriptions, and enumeration values ​​on the platform to construct a pre-defined label knowledge base (i.e., the Agent label knowledge base). It is important to understand that the pre-defined label knowledge base serves as the verification benchmark for the knowledge gap detection head in the subsequent strategy discrimination model, and also as the source of labels used by the strategy generation model when generating candidate selection strategies.

[0178] In addition, platform tags may be added, modified, or removed. The latest tag information is automatically retrieved from the platform and synchronized to the preset tag knowledge base through a daily scheduled task, ensuring the timeliness of the preset tag knowledge base and enabling the knowledge gap detection head to accurately verify the candidate selection strategy based on the current tag data.

[0179] Furthermore, during the execution of the strategy generation model, it is necessary to pre-define the execution process specifications and typical selection examples, forming a Reference document to prevent the strategy generation model from experiencing execution illusions and failing to complete the selection task. In subsequent selection processes, the examples in the Reference document can be automatically updated based on feedback from operations personnel regarding the target selection strategy, ensuring that the generation of subsequent candidate selection strategies continuously benefits from historical success. Additionally, during the generation of candidate selection strategies, the strategy generation model needs to define relevant execution processes and typical examples to avoid the strategy generation model experiencing execution illusions and failing to achieve the selection objective. Furthermore, in subsequent selection processes, relevant examples will be automatically updated based on feedback from operations personnel.

[0180] Based on the above, while the general model can understand the needs of operations personnel and refer to the knowledge base to execute the selection task, it lacks an understanding of the platform users' daily selection business. As a result, it can only output a relatively simple selection strategy, which cannot meet the users' selection needs. Therefore, it is necessary to use relevant data within the site to perform a supervised fine-tuning of the general model to obtain a strategy generation model. This model can then be used to output selection strategies based on the sample selection needs information described in natural language.

[0181] In one specific implementation, historical data from real-world selection scenarios within the station are collected to construct an SFT training dataset. It's important to understand that each data point in this SFT training dataset contains a sample selection requirement (i.e., the query entered by the operations staff) and its corresponding manually labeled or historically adopted candidate selection strategies (i.e., Answers). Approximately 20,000 samples are selected from this dataset and divided into training and test sets.

[0182] Supervised fine-tuning (SFT) of the general large language model is performed using the partitioned training set, enabling the model to learn the platform's unique selection business knowledge, tag usage logic, and strategy generation paradigm, thereby gaining the basic ability to generate candidate selection strategies that conform to the platform's specifications based on natural language queries.

[0183] Furthermore, based on the supervised fine-tuning policy generation model, multiple candidate selection policies are generated using the SFT test dataset. Then, a DPO training dataset is constructed based on the preference ranking (Chosen / Rejected) of manual or auxiliary model annotations. The policy generation model is further trained using the Direct Preference Optimization algorithm, internalizing the user's quality preferences for selection policies into the model parameters. This strengthens the tendency to generate good policies with high intent fidelity and low policy drift risk, while suppressing the generation probability of unreasonable policies.

[0184] Based on the above, a policy generation model with semantic understanding capabilities is obtained.

[0185] Furthermore, for the SFT training dataset, keeping the Query unchanged, the actual labels and enumeration values ​​used are extracted from the corresponding Answer (selection strategy) as new Answer targets. At the same time, the annotators label each extracted label and enumeration value with constraint dimensions, and label them as strong constraints (must), soft constraints (should), or forbidden extrapolation according to their semantic role and necessity in user needs.

[0186] The new Answer target and Query obtained above are used as the new SFT training dataset to train the feature extraction model. It is important to understand that during the training of the feature extraction model, the input is the sample selection requirement information (i.e., the Query entered by the operations personnel), and the output is a set of labels and enumerated values ​​marked with strong constraints (must), soft constraints (should), or forbidden extrapolation.

[0187] Optionally, the training process of the feature extraction model can be understood as supervised fine-tuning based on a general large language model to obtain the feature extraction model (or anchor extraction model). The feature extraction model is used to extract target constraints from the input selection requirement information, that is, a structured anchor set containing three types of constraint dimensions: strong constraints, soft constraints, and prohibition of extrapolation. This transforms the fuzzy natural language selection requirements into a structured and legal semantic boundary that can be rigorously verified by the subsequent policy discrimination model.

[0188] Given the strategy generation model and feature extraction model, in order to improve the accuracy of determining the target selection strategy, this specification uses a strategy discrimination model to evaluate the selection strategy and generate corresponding strategy adjustment attribute information.

[0189] The following is in conjunction with the appendix Figure 2 The training process of the policy discrimination model is explained. Figure 2 A flowchart illustrating the training process of a policy discrimination model provided in one embodiment of this specification is shown.

[0190] exist Figure 2 In the reuse strategy generation model SFT stage, the real on-site SFT training dataset used (each containing sample selection requirement information Query and its corresponding selection strategy Answer) is used. The sample selection requirement information in the SFT training dataset is kept unchanged, and the platform labels and enumeration values ​​actually used are extracted from the selection strategy as new output targets.

[0191] Each extracted label / enumeration value is labeled with constraint dimensions by the annotators, and classified into the following three categories according to its semantic role in the sample selection requirement information: Strong constraints (must): clearly defined necessary conditions; their absence will cause the selected objects to deviate completely from the user's intent. Soft constraints (should): Expressed tendencies and preferences, allowing for additions or deletions during policy generation; Forbidden extrapolation: Explicit or implicit range boundaries that must not be exceeded when generating strategies.

[0192] Based on the above, a new dataset is obtained and used as training data for supervised fine-tuning (SFT) on a general large language model. Through this training, the model learns to accurately map fuzzy natural language selection requirements into a structured set of anchor points, resulting in target constraints containing three types of constraint dimensions.

[0193] Furthermore, based on the above, the policy discrimination model is trained.

[0194] In one specific embodiment provided in this specification, input data is acquired.

[0195] Specifically, based on the sample selection requirement information (Query), a strategy generation model is used to generate candidate selection strategies, resulting in the original DPO dataset. Then, a feature extraction model is used to obtain the target constraints (anchor point set).

[0196] The input data is fed into a general large language model for supervised fine-tuning to obtain the shared semantic encoder backbone of the policy discrimination model.

[0197] Backbone can understand the semantics of policy discrimination tasks, encode the input (candidate policies, target constraints) into a hidden layer representation that integrates global semantics, and generate natural language text through the LM Head (language model head), outputting a description of the specific problems of the candidate selection policies and modification suggestions, as well as the judgment conclusion.

[0198] After the final layer output of Backbone, separate MLP task heads are connected for different problem types. Each task head is trained independently, using only manually labeled binary positive and negative samples (0 / 1 labels) as training data, without relying on text-generated data. Specifically, this includes: Specifically, based on the intent drift detection head (Head 1), it calculates whether the selection strategy adds additional conditions outside the forbidden list, or performs semantic extrapolation on the must condition. It outputs the intent drift probability.

[0199] Based on the knowledge gap detection head (Head 2), each tag / enumeration value in the selection strategy is hard-matched against a preset tag knowledge base. The knowledge gap probability is output based on whether the selection strategy references a tag that does not exist in the knowledge base.

[0200] Based on the constraint missing detection header (Head 3), check one by one whether all constraints in the must list appear in the selection strategy; check whether forbidden constraints are violated. Output the constraint missing probability according to the severity of the missing / violation.

[0201] The intention drift probability, knowledge gap probability, and constraint missing probability corresponding to each selection strategy are weighted and fused to obtain the strategy score corresponding to each selection strategy.

[0202] Based on the above, the trained policy discrimination model is obtained.

[0203] Furthermore, for ease of understanding, the information processing method provided in this specification will be used as an example in the application of selecting target groups to further illustrate the information processing method. Among other things, Figure 3 This specification shows a flowchart of an information processing method for selecting a target population, provided in one embodiment.

[0204] like Figure 3 As shown, in a specific implementation provided in this specification, the target selection requirement information (Query) is input into the strategy generation model. During inference, the strategy generation model simultaneously obtains a Reference document (containing a preset tag knowledge base and historical selection examples) as a reference, and generates at least one candidate selection strategy based on the strategy generation model.

[0205] Similarly, the target selection requirement information (Query) is input into the feature extraction model (i.e., the anchor extraction model). Using the feature extraction model, constraint features of three constraint dimensions—must (strong constraint), should (soft constraint), and forbidden (prohibition of extrapolation)—are extracted from the Query to generate structured target constraints, i.e., the anchor set. This anchor set defines the legitimate semantic boundaries of subsequent policy generation and serves as a benchmark for quality evaluation by the policy discrimination model.

[0206] Furthermore, each candidate selection strategy, target constraints, and a pre-defined label knowledge base are input into the strategy discrimination model. In the strategy discrimination model, a text generation module (LM Head) generates natural language text information containing a description of the specific problems with the candidate selection strategy and targeted modification suggestions. The scoring task module (MLP Heads) outputs three independent scoring task heads, each outputting a probability score for each constraint dimension: intent drift probability P(drift) ∈ [0,1], knowledge gap probability P(gaps) ∈ [0,1], and constraint missing probability P(missing) ∈ [0,1]. The weighted average of these three scores yields the comprehensive strategy score_t for the current candidate strategy. The text information and the comprehensive score_t are used as the strategy adjustment attribute information corresponding to the current candidate selection strategy.

[0207] The strategy adjustment attribute information and the current candidate selection strategy are input into the strategy generation model. The strategy generation model treats the strategy adjustment attribute information as a correction instruction, and makes targeted adjustments to the candidate strategies to obtain updated candidate selection strategies.

[0208] Specifically, the updated candidate selection strategy is obtained by inserting missing must constraints, deleting extrapolation conditions that exceed the forbidden boundary, and replacing empty tags that are not in the preset tag knowledge base with valid tags.

[0209] Calculate the difference between the updated candidate selection strategy and the current candidate selection strategy. ).

[0210] In one case, the difference is large and there is still a problem, so the iteration continues.

[0211] In another case, if the difference is large and there are no longer any problems, then stop and output.

[0212] In another case, the difference is small and there are no longer any problems. This can be understood as natural convergence, so stop and output.

[0213] In another scenario, if the difference is small but there are still issues, then a bottleneck has been reached, and the process should be stopped and feedback from operations personnel should be awaited.

[0214] In one specific implementation provided in this specification, the details of the target selection strategy that meets the update stopping conditions and the scores of each dimension (including the probability of intent drift, the probability of knowledge gaps, the probability of constraint missing, and the comprehensive score) are output to the operations personnel to demonstrate the quantitative quality assessment results of the target selection strategy.

[0215] To facilitate understanding, this specification also provides a segmentation architecture diagram for an embodiment of audience segmentation operations on a content application platform. As shown in Figure 4, Figure 4 illustrates a segmentation architecture diagram provided in an embodiment of this specification.

[0216] In one specific embodiment provided in this specification, the selection architecture diagram shown in Figure 4 includes a model pre-training module and an intelligent processing unit self-evolution module.

[0217] The model pre-training module is used in the model pre-training stage. It includes: knowledge base construction, SFT training of policy generation model data, SFT training of anchor extraction model data, generation of DPO data and policy discrimination model SFT data, training of policy discrimination model, and training of policy generation model DPO.

[0218] The intelligent processing unit's self-evolution module is applied to the agent assembly self-evolution stage. It includes: a policy generation model generating candidate selection strategies based on user input queries; an anchor point extraction model extracting constraint anchor points; a policy discrimination model performing multi-dimensional review of candidate selection strategies; iterative correction by the policy generation model based on the problem description; termination judgment; final output; and post-processing operations.

[0219] In one specific embodiment provided in this specification, the accompanying drawings are used in conjunction with the provided drawings. Figure 4A The process of building the knowledge base is explained. Figure 4A This specification illustrates a schematic diagram of a knowledge base construction method provided in one embodiment.

[0220] This manual constructs a knowledge base from two aspects: a pre-defined tag knowledge base and reference documents.

[0221] In one specific implementation, for a pre-defined tag knowledge base (Agent tag knowledge base), all available tags, tag descriptions, and enumeration values ​​are collected from the platform's tag data source and stored in a structured format. For example, the enumeration values ​​corresponding to the tag "user age range" are "18-24", "25-30", "31-35", etc. Furthermore, a scheduled task (executed daily) retrieves the latest tag information from the platform and synchronously updates the knowledge base to ensure timeliness. It should be understood that the tasks mentioned in this specification can be understood as filtering tasks or selection tasks.

[0222] In one specific implementation, execution process specifications and typical selection examples are preset for the reference document (working reference knowledge base). For example, the workflow specifications for the user selection agent (such as standard output format, steps that must be followed for strategy generation) and typical selection examples (including correct queries and corresponding strategies) are pre-written to avoid execution illusions during agent reasoning. Simultaneously, in subsequent interactions, based on scheduled tasks, if operations personnel confirm the effectiveness of a certain strategy, the strategy and its query are added as new examples to the knowledge base for subsequent reasoning reference.

[0223] In one specific embodiment provided in this specification, the accompanying drawings are used in conjunction with the provided drawings. Figure 4B The process of training the policy generation model (SFT) and the anchor extraction model (SFT) is explained. Figure 4B This diagram illustrates the construction of a supervised training dataset according to an embodiment of this specification.

[0224] In one specific implementation, the strategy generation model, SFT, is trained as follows: Approximately 20,000 historical user targeting data entries are selected from existing site data. Each entry includes a user description (the user's input query indicating their user targeting needs) and a corresponding targeting strategy (a candidate targeting strategy answer). This dataset is split into an SFT training dataset (first training dataset) and an SFT test dataset (first test dataset). The general-purpose model is then supervised fine-tuned (SFT) using the SFT training dataset to obtain the initial strategy generation model. This model can generate multiple candidate targeting strategies (e.g., 3-5 strategies) based on the user's input natural language query and knowledge base information.

[0225] In one specific implementation, the SFT training for the anchor extraction model involves the following steps: Secondary annotation is performed on the queries in the aforementioned SFT training dataset: all actual labels and enumerated values ​​involved in the query are extracted as Answers; simultaneously, each constraint is manually categorized and labeled as a strong constraint (must), a soft constraint (should), or a forbidden constraint (forbidden). A new SFT dataset is constructed to obtain a new SFT training dataset, and the anchor extraction model is fine-tuned. This anchor extraction model is used to extract three types of structured anchor sets from any query to obtain the target constraints, providing precise boundaries for subsequent discrimination.

[0226] In one specific embodiment provided in this specification, the accompanying drawings are used in conjunction with the provided drawings. Figure 4C The process of generating DPO data and discriminant model SFT data is explained. Figure 4C This diagram illustrates an embodiment of the present specification of constructing direct preference data and policy discrimination model data.

[0227] exist Figure 4C In this process, using the queries from the aforementioned SFT test dataset, the trained policy generation model generates 3-5 candidate policies for each query, forming the original DPO dataset (the second augmented dataset). Then, using the "audience description, anchor information, and candidate selection policy" corresponding to each candidate policy as input, a larger-scale general model is invoked as a temporary discriminator (anchor extraction model), outputting the problem description and conclusion (accept / rejection) for that policy. The generated results are then manually sampled and anomaly labels are corrected, ultimately resulting in two types of datasets: The DPO dataset consists of a query, a chosen policy, and a rejected policy for each sample. Discriminative model SFT dataset (second training dataset): Each sample contains input (Query, anchor, policy) and output (problem description text, decision conclusion).

[0228] In one specific embodiment provided in this specification, the accompanying drawings are used in conjunction with the provided drawings. Figure 4D The training process of the policy discrimination model is explained. Figure 4D This diagram illustrates the training of a policy discrimination model provided in one embodiment of this specification.

[0229] like Figure 4D As shown, the policy discrimination scoring model adopts a dual-head architecture, and the training is divided into two stages: (1) LM Head training: Using the above-mentioned discriminative model SFT dataset, the small-scale general model is fine-tuned with standard SFT so that it can master the semantics of the discriminative task and generate a description of the problems in the strategy and modification suggestions (text output) through LM Head.

[0230] (2) MLP Head Training: After the last layer (last token hidden state) of the Backbone, three independent MLP task heads are fed in parallel. Each task head is a three-layer MLP (3xMLPHead), corresponding to the three dimensions of "intent drift", "knowledge gaps", and "constraint missing". The training data only uses manually labeled 0 / 1 binary samples (positive examples indicate that there is a problem in this dimension, and negative examples indicate that there is no problem). Each task head is trained independently, and the output is the problem probability of this dimension (a continuous value between 0 and 1). For example, the knowledge gap task head inputs a tag that appears in the policy. If the tag is not in the knowledge base, it is labeled as 1; otherwise, it is labeled as 0.

[0231] During inference, the same input is simultaneously processed by the Backbone, the LM Head outputs the problem description and conclusion, and three MLPHeads output the probabilities of each dimension. The overall score is then calculated by weighted averaging. This design ensures that the score is consistent with the semantics of the text and that the result is deterministic (without sampling randomness).

[0232] In one specific embodiment provided in this specification, the accompanying drawings are used in conjunction with the provided drawings. Figure 4E The training process of Direct Preference Optimization (DPO) for policy generation models is explained.

[0233] like Figure 4E As shown, Figure 4E This diagram illustrates a direct preference optimization of a policy generation model provided in one embodiment of this specification.

[0234] Using the DPO dataset, we trained a policy generation model with SFT (Simplified Policy Theory) using Direct Preference Optimization (DPO). The training objective was to maximize the generation probability of positive samples (chosen) while minimizing the generation probability of negative samples (rejected), so that the model parameters internalized the intent-fidelity objective, resulting in an optimized policy generation model.

[0235] In one specific embodiment provided in this specification, the accompanying drawings are used in conjunction with the provided drawings. Figure 4F The self-evolutionary stage of Agent assembly is explained.

[0236] Figure 4F A flowchart illustrating the self-evolution of an intelligent processing unit provided in one embodiment of this specification is shown.

[0237] exist Figure 4F In the process, the system receives the target selection requirement information, i.e., the selection query.

[0238] The strategy generation model combines a work reference knowledge base (including process specifications and examples) and a tag knowledge base to generate 3 to 5 candidate strategies for identifying people (each strategy is composed of several tag conditions).

[0239] The anchor extraction model parses the same query and outputs a set of anchor points.

[0240] The policy discrimination model takes a "Query, a set of anchor points, and a candidate policy" as input, and performs knowledge verification using a pre-defined label knowledge base. The model outputs in parallel: Description of existing problems with the LM Head output strategy and suggestions for improvement; The three MLP Heads output the probability of intent drift (P1), the probability of knowledge gaps (P2), and the probability of constraint missing (P3), respectively, from which the overall score is obtained. Where α, β, and γ are preset weights.

[0241] The problem description and modification suggestions output by the LM Head are fed back to the policy generation model. Based on this, the policy generation model adjusts the current policy and generates a revised policy. The revised policy is then fed back into the policy discrimination model to obtain a new score, score_{t+1}.

[0242] calculate The update stop condition is determined based on the following rules: If Δscore is greater than a preset threshold (e.g., 0.1) and there are still problems (any probability in each dimension is greater than 0.5), then continue optimization; If Δscore is greater than the preset threshold and there are no more problems (probabilities of all dimensions are ≤0.5), then output; If Δscore is less than or equal to the threshold and there are no more problems, it is considered to have converged naturally and is output. If Δscore is less than or equal to the threshold and problems still exist, it is determined that the Agent has reached the self-evolution bottleneck, will no longer continue iterating, will directly enter Step 6, and will output the strategy to the operations staff to wait for human feedback.

[0243] The final selected selection strategy (or all candidate strategies and their scores) will be used as the target selection strategy and presented to the operations staff in a structured manner, along with scores for each dimension and a description of any issues. After confirmation by the operations staff, the system will execute subsequent link operations such as selection and push notifications (e.g., ad placement and content recommendation).

[0244] After execution, based on feedback from operations personnel (such as whether it was adopted or modified), the reference document (work reference knowledge base) is updated via prompt words, and the successful experience or corrected examples are added to the knowledge base for subsequent reasoning.

[0245] Corresponding to the above method embodiments, this specification also provides embodiments of an information processing apparatus. Figure 5 A schematic diagram of the structure of an information processing apparatus according to an embodiment of this specification is shown. Figure 5 As shown, the device includes: The acquisition unit 502 is configured to acquire target screening requirement information and determine the target screening semantics corresponding to the target screening requirement information; and based on the target screening semantics, determine the target constraints corresponding to the target screening requirement information.

[0246] The generation unit 504 is configured to generate at least one candidate screening strategy based on the target screening requirement information.

[0247] The judgment unit 506 is configured to determine the strategy adjustment attribute information corresponding to each candidate screening strategy based on each candidate screening strategy and the target constraint conditions.

[0248] Processing unit 508 is configured to update each candidate filtering strategy based on the attribute information of each strategy until the update stop condition is met, and to determine at least one target filtering strategy; and to determine the target filtering object based on each target filtering strategy.

[0249] Furthermore, the acquisition unit 502 is further configured as follows: Based on the target filtering semantics, extract constraint features of at least one constraint dimension; Based on the characteristics of each constraint, the target constraint conditions are determined.

[0250] Furthermore, the acquisition unit 502 is also configured as follows: The target screening requirements are input into the feature extraction model; Based on the feature extraction model, constraint features of at least one constraint dimension corresponding to the target screening requirement information are extracted, and target constraint conditions are generated based on each constraint feature.

[0251] Furthermore, the generating unit 504 is further configured as follows: The target screening requirements information is input into the strategy generation model; The target filtering semantics are determined in the strategy generation model, and at least one candidate filtering strategy output by the strategy generation model is generated based on the target filtering semantics.

[0252] Furthermore, the judgment unit 506 is further configured as follows: Determine an initial candidate screening strategy, wherein the initial candidate screening strategy is any one of the candidate screening strategies; Input prompt words are constructed based on the initial candidate selection strategy and the target constraints; The input prompts are fed into the policy discrimination model to obtain the policy adjustment attribute information corresponding to each candidate filtering policy output by the policy discrimination model.

[0253] Furthermore, the policy discrimination model includes a semantic encoder, a text generation module, and at least one scoring task module; The judgment unit 506 is further configured as follows: The input prompt words are input into the semantic encoder for feature extraction to obtain the initial policy features output by the semantic encoder; The initial strategy features are input into the text generation module, and initial text information containing a problem description and modification suggestions is generated based on the text generation module. The initial policy features are input into each scoring task module, and the policy score for each constraint dimension is determined based on each scoring task module. The text information and the strategy score are determined as the strategy adjustment attribute information.

[0254] Furthermore, the scoring task module can be any one of the following: intent drift detection head, knowledge gap detection head, and constraint missing detection head; The judgment unit 506 is further configured as follows: Based on the intent drift detection head, it detects whether the initial candidate screening strategy meets the set semantic boundary and outputs the intent drift probability, wherein the set semantic boundary is determined based on the constraint features corresponding to each constraint dimension in the target constraint condition; Based on the knowledge void detection head, it detects whether the initial candidate screening strategy meets the preset label knowledge base and outputs the knowledge void probability. Based on the constraint missing detection head, detect whether the initial candidate screening strategy satisfies the target constraint condition, and output the constraint missing probability; The strategy score is determined based on the intent drift probability, the knowledge gap probability, and the constraint missing probability.

[0255] Furthermore, the processing unit 508 is further configured as follows: Determine the candidate filtering strategy to be processed and the corresponding adjustment attribute information of the candidate filtering strategy to be processed, wherein the candidate filtering strategy to be processed is any one of the candidate filtering strategies; The candidate filtering strategy to be processed and the adjustment attribute information of the strategy to be processed are input into the strategy generation model to obtain the updated candidate filtering strategy output by the strategy generation model. The updated candidate filtering strategy is input into the strategy discrimination model to obtain the updated strategy adjustment attribute information corresponding to the updated candidate filtering strategy output by the strategy discrimination model; When the updated attribute information meets the update stop condition, the target filtering strategy is obtained.

[0256] Furthermore, the processing unit 508 is also configured as follows: Obtain the update strategy score corresponding to the update strategy adjustment attribute information, and obtain the strategy score corresponding to the strategy adjustment attribute information to be processed; Based on the update strategy score and the difference between the strategy scores, it is determined whether the update strategy adjustment attribute information meets the update stop condition.

[0257] Furthermore, the processing unit 508 is further configured as follows: If at least two candidate screening strategies meet the update stopping condition, determine the strategy score corresponding to each candidate screening strategy. Based on the scores of each strategy, at least one target screening strategy is determined among the candidate screening strategies.

[0258] The above is an illustrative scheme of an information processing device according to this embodiment. It should be noted that the technical solution of this information processing device and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the information processing device, please refer to the description of the technical solution of the information processing method described above.

[0259] See Figure 6 , Figure 6 This specification illustrates an architecture diagram of an information processing system according to one embodiment of the present specification. The information processing system may include a client 100 and a server 200. Client 100 is used to send target filtering requirement information to server 200; Server 200 is used to obtain target filtering requirement information and determine the target filtering semantics corresponding to the target filtering requirement information; based on the target filtering semantics, determine the target constraints corresponding to the target filtering requirement information; based on the target filtering requirement information, generate at least one candidate filtering strategy; based on each candidate filtering strategy and the target constraints, determine the strategy adjustment attribute information corresponding to each candidate filtering strategy; update each candidate filtering strategy based on each strategy adjustment attribute information until the update stop condition is met, and determine at least one target filtering strategy; based on each target filtering strategy, determine the target filtering object.

[0260] Send the target filtering object to client 100; Client 100 is also used to receive target filtering objects sent by server 200.

[0261] The information processing system may include multiple clients 100 and a server 200. Clients 100 can be referred to as end-side devices, and the server 200 can be referred to as cloud-side devices. Multiple clients 100 can establish communication connections through the server 200. In the information processing scenario, the server 200 is used to provide information processing services between the multiple clients 100. Each client 100 can act as a sender or receiver, communicating through the server 200.

[0262] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In information processing scenarios, users can publish data streams to server 200 through client 100, server 200 can generate target selection objects based on the data stream, and push the target selection objects to other clients that have established communication.

[0263] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.

[0264] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0265] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0266] It is worth noting that the information processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the information processing methods provided in the embodiments of this specification. In other embodiments, the information processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0267] Figure 7 A structural block diagram of a computing device according to an embodiment of this specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.

[0268] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0269] In one embodiment of this specification, the above-described components of the computing device 700 and Figure 7 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 7 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0270] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.

[0271] The processor 720 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described information processing method.

[0272] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the information processing method described above.

[0273] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the information processing method described above.

[0274] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the information processing method described above.

[0275] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described information processing method.

[0276] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the information processing method described above.

[0277] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0278] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0279] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.

[0280] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0281] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An information processing method, characterized in that, include: Obtain target filtering requirement information and determine the target filtering semantics corresponding to the target filtering requirement information; Based on the target filtering semantics, determine the target constraints corresponding to the target filtering requirement information; Based on the target screening requirements information, at least one candidate screening strategy is generated; Based on each candidate screening strategy and the target constraints, determine the strategy adjustment attribute information corresponding to each candidate screening strategy; The candidate filtering strategies are updated based on the attribute information of each strategy until the update stop condition is met, and at least one target filtering strategy is determined. Based on the screening strategies for each target, the target screening objects are determined.

2. The method as described in claim 1, characterized in that, Based on the target filtering semantics, the target constraints corresponding to the target filtering requirement information are determined, including: Based on the target filtering semantics, extract constraint features of at least one constraint dimension; Based on the characteristics of each constraint, the target constraint conditions are determined.

3. The method as described in claim 2, characterized in that, The method further includes: The target screening requirements are input into the feature extraction model; Based on the feature extraction model, constraint features of at least one constraint dimension corresponding to the target screening requirement information are extracted, and target constraint conditions are generated based on each constraint feature.

4. The method as described in claim 1, characterized in that, Based on the target screening requirements information, at least one candidate screening strategy is generated, including: The target screening requirements information is input into the strategy generation model; The target filtering semantics are determined in the strategy generation model, and at least one candidate filtering strategy output by the strategy generation model is generated based on the target filtering semantics.

5. The method as described in claim 1, characterized in that, Based on each candidate screening strategy and the target constraints, the strategy adjustment attribute information corresponding to each candidate screening strategy is determined, including: Determine an initial candidate screening strategy, wherein the initial candidate screening strategy is any one of the candidate screening strategies; Input prompt words are constructed based on the initial candidate filtering strategy and the target constraints; The input prompts are fed into the policy discrimination model to obtain the policy adjustment attribute information corresponding to each candidate filtering policy output by the policy discrimination model.

6. The method as described in claim 5, characterized in that, The policy discrimination model includes a semantic encoder, a text generation module, and at least one scoring task module; The input prompts are fed into the policy discrimination model to obtain the policy adjustment attribute information corresponding to each candidate filtering policy output by the policy discrimination model, including: The input prompt words are input into the semantic encoder for feature extraction to obtain the initial policy features output by the semantic encoder; The initial strategy features are input into the text generation module, and initial text information containing a problem description and modification suggestions is generated based on the text generation module. The initial policy features are input into each scoring task module, and the policy score for each constraint dimension is determined based on each scoring task module. The text information and the strategy score are determined as the strategy adjustment attribute information.

7. The method as described in claim 6, characterized in that, The scoring task module can be any one of the intent drift detection head, knowledge gap detection head, and constraint missing detection head. The policy score for each constraint dimension is determined based on each scoring task module, including: Based on the intent drift detection head, it detects whether the initial candidate screening strategy meets the set semantic boundary and outputs the intent drift probability, wherein the set semantic boundary is determined based on the constraint features corresponding to each constraint dimension in the target constraint condition; Based on the knowledge void detection head, it detects whether the initial candidate screening strategy meets the preset label knowledge base and outputs the knowledge void probability. Based on the constraint missing detection head, detect whether the initial candidate screening strategy satisfies the target constraint condition, and output the constraint missing probability; The strategy score is determined based on the intent drift probability, the knowledge gap probability, and the constraint missing probability.

8. The method as described in claim 1, characterized in that, Based on the attribute information adjusted by each strategy, each candidate filtering strategy is updated until the update stop condition is met, and at least one target filtering strategy is determined, including: Determine the candidate filtering strategy to be processed and the corresponding adjustment attribute information of the candidate filtering strategy to be processed, wherein the candidate filtering strategy to be processed is any one of the candidate filtering strategies; The candidate filtering strategy to be processed and the adjustment attribute information of the strategy to be processed are input into the strategy generation model to obtain the updated candidate filtering strategy output by the strategy generation model. The updated candidate filtering strategy is input into the strategy discrimination model to obtain the updated strategy adjustment attribute information corresponding to the updated candidate filtering strategy output by the strategy discrimination model; When the updated attribute information meets the update stop condition, the target filtering strategy is obtained.

9. The method as described in claim 8, characterized in that, The method further includes: Obtain the update strategy score corresponding to the update strategy adjustment attribute information, and obtain the strategy score corresponding to the strategy adjustment attribute information to be processed; Based on the update strategy score and the difference between the strategy scores, it is determined whether the update strategy adjustment attribute information meets the update stop condition.

10. The method as described in claim 1, characterized in that, Determine at least one target screening strategy, including: If at least two candidate screening strategies meet the update stopping condition, determine the strategy score corresponding to each candidate screening strategy. Based on the scores of each strategy, at least one target screening strategy is determined among the candidate screening strategies.

11. An information processing device, characterized in that, include: The acquisition unit is configured to acquire target screening requirement information and determine the target constraints corresponding to the target screening requirement information. The generation unit is configured to generate at least one candidate filtering strategy based on the target filtering requirement information; The judgment unit is configured to input each candidate screening strategy and the target constraint into the strategy discrimination model, and obtain the strategy adjustment attribute information corresponding to each candidate screening strategy output by the strategy discrimination model; The processing unit is configured to update each candidate filtering strategy based on the attribute information of each strategy until the update stop condition is met, and to determine at least one target filtering strategy. Based on the screening strategies for each target, the target screening objects are determined.

12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 10.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 10.