Model training sampling method and device, poster generation method and device, equipment and medium

By introducing a dual-section management mechanism of challenge zones and diversity zones in poster generation, the problem of inefficient training is solved, and the quality and efficiency of poster layout generation is improved, and intelligent poster generation scenarios are suitable for financial technology and medical and health fields.

CN120580321APending Publication Date: 2025-09-02PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510687324.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing technology has problems of inefficient training efficiency and waste of computing resources in poster generation. Especially in multi-objective optimization tasks that require taking into account aesthetic rules, content levels and spatial constraints. Traditional sampling strategies are difficult to adapt to poster layout feature learning needs at different levels of difficulty, resulting in limited improvement in model performance.

Method used

The dual-section management mechanism of challenge zone and diversity zone is adopted, and the sample set is divided into challenge zone and diversity zone by obtaining the difficulty score of the sample, so as to collect training samples. The challenge zone focuses on the layout combination that is difficult to deal with in the current model for intensive training. The diversity zone maintains the sample distribution of different semantic categories to prevent the model from falling into the local optimal solution.

Benefits of technology

The optimal allocation of training resources is achieved, the quality and training efficiency of poster layout generation is improved, and the redundancy of repeated training of simple samples is avoided, ensuring that the model maintains generalization ability while overcoming difficulties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580321A_ABST
    Figure CN120580321A_ABST
Patent Text Reader

Abstract

The invention discloses a model training sampling method and device, a poster generation method and device, equipment and a medium, and relates to the financial science and technology field, the medical health field and the artificial intelligence technology field, and the method comprises the steps: obtaining the difficulty score of each sample in a sample set; dividing the sample set into a challenge area and a diversity area according to the difficulty score of the sample; training samples are collected from the challenge zone and the diversity zone. According to the method, a challenge area focuses on a layout combination (such as multi-element asymmetric arrangement) which is difficult to process by a current model, and convergence is accelerated through intensified training; and the diversity region is used for preventing the model from falling into a local optimal solution by maintaining sample distribution of different semantic categories (such as posters of different styles). The double-section management mechanism realizes optimal allocation of training resources, avoids repeated training redundancy of simple samples, ensures that the model keeps generalization ability while overcoming difficulties, and finally remarkably improves layout generation quality and training efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of financial technology, medical health and artificial intelligence technology, and in particular to a model training sampling method, a poster generation method, an apparatus, equipment and a medium. Background Art

[0002] Posters are widely used in the fields of financial technology and healthcare. How to quickly and accurately generate posters is a difficult problem in these fields.

[0003] In the field of poster element layout generation based on deep learning, traditional methods generally adopt fixed batch sampling strategies for model training.

[0004] This training method has obvious technical limitations: on the one hand, the model often requires a large number of repeated trainings to reach a state of convergence, resulting in low training efficiency. On the other hand, all samples are treated equally during the training process. This indiscriminate treatment not only prolongs the model's learning cycle for key difficult samples, but also results in a significant waste of computing resources. Especially in poster design, a multi-objective optimization task that requires considering aesthetic rules, content hierarchy, and spatial constraints, traditional sampling strategies are difficult to adapt to the learning requirements of poster layout features at different difficulty levels, seriously restricting further improvement of model performance. Summary of the Invention

[0005] The embodiments of the present invention provide a model training sampling method, a poster generation method, an apparatus, a device and a medium, which aim to solve the technical problem of low training efficiency of existing model training methods.

[0006] In a first aspect, an embodiment of the present invention provides a model training sampling method, which includes:

[0007] Get the difficulty score of each sample in the sample set;

[0008] Dividing the sample set into a challenge zone and a diversity zone according to the difficulty scores of the samples, wherein the difficulty scores of the samples in the challenge zone are greater than the difficulty scores of the samples in the diversity zone;

[0009] Training samples are collected from the challenge area and the diversity area.

[0010] In a second aspect, an embodiment of the present invention provides a poster generation method, comprising:

[0011] Receiving poster design elements, extracting features of the poster design elements, and obtaining an input vector;

[0012] Inputting the input vector into a preset poster layout generation model, so that the poster layout generation model outputs poster layout parameters based on the input vector, wherein, during the training process of the poster layout generation model, training samples are collected based on the model training sampling method as described in the first aspect;

[0013] A poster file is generated based on the poster layout parameters.

[0014] In a third aspect, an embodiment of the present invention further provides a model training sampling device, which includes a unit for executing the above method.

[0015] In a fourth aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0016] In a fifth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program can implement the above method when executed by a processor.

[0017] The embodiments of the present invention provide a model training sampling method, a poster generation method, an apparatus, a device and a medium. The method includes: obtaining the difficulty score of each sample in a sample set; dividing the sample set into a challenge area and a diversity area according to the difficulty score of the sample, wherein the difficulty score of the sample in the challenge area is greater than the difficulty score of the sample in the diversity area; and collecting training samples from the challenge area and the diversity area. In the present invention, the challenge area focuses on layout combinations that are difficult for the current model to handle (such as asymmetric arrangement of multiple elements), and accelerates convergence through intensive training; the diversity area prevents the model from falling into a local optimal solution by maintaining the distribution of samples of different semantic categories (such as posters of different styles). This dual-segment management mechanism achieves the optimal allocation of training resources, avoids the redundancy of repeated training of simple samples, and ensures that the model maintains generalization ability while overcoming difficulties, and ultimately significantly improves the quality of poster layout generation and training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A flow chart of a model training sampling method provided by an embodiment of the present invention;

[0020] Figure 2A flowchart of a poster generation method provided by an embodiment of the present invention;

[0021] Figure 3 A schematic block diagram of a model training sampling device provided by an embodiment of the present invention;

[0022] Figure 4 A schematic block diagram of a poster generating device provided by an embodiment of the present invention;

[0023] Figure 5 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0025] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0026] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0027] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0029] The present invention can be widely used in intelligent poster generation scenarios in the fields of financial technology and medical health, and its technical solution can effectively adapt to the complex business needs of the two fields.

[0030] In the field of financial technology, the poster display covers the entire chain of payment transaction scenarios, including payment agreement optimization, cross-border settlement systems, electronic signature and intelligent invoice management, while supporting e-commerce promotion strategy generation, dynamic pricing advertising and customer relationship management in commercial service scenarios, and extending to core financial business scenarios, such as the visualization of credit risk assessment in online banking services, dynamic charting of securities market data, intelligent underwriting process description of insurance products, and personalized generation of tax declaration guidelines.

[0031] In the field of medical health, the poster display focuses on serving the entire process of smart medical care, covering collaborative management from digital registration and appointment, cloud-based medical record analysis to remote medical monitoring, supporting the integrated display of wearable device data and electronic health records in the medical Internet of Things scenario, and adapting to the doctor-patient interaction interface design of the Internet hospital platform, the generation of medical community resource scheduling dashboards, and patient medication reminders and health trend analysis in chronic disease management scenarios. It can also realize the mixed text and image output of AI medical imaging reports and the visualization of full-link tracking of the pharmaceutical supply chain.

[0032] This solution can convert multi-dimensional business data into visual poster content that complies with industry standards in real time, and demonstrates significant technical advantages in scenarios that require integrating heterogeneous data across systems and realizing multi-dimensional information fusion presentation.

[0033] See also Figure 1 , an embodiment of the present invention provides a model training sampling method, such as Figure 1 As shown, the method includes the following steps:

[0034] S1, obtain the difficulty score of each sample in the sample set.

[0035] In a specific implementation, the sample set includes multiple samples, which may be poster design samples, but the present invention does not specifically limit this. In the present invention, the difficulty of the sample is evaluated to obtain a difficulty score of the sample.

[0036] For example, in some preferred embodiments, the above step of "obtaining the difficulty score of each sample in the sample set" specifically includes the following steps:

[0037] S11, determining a first scoring item based on the gradient norm of the loss function of the sample.

[0038] In specific implementations, the first scoring term can be referred to as the gradient dynamic evaluation term. The loss function gradient norm reflects the model's sensitivity to samples under the current parameters. Samples with high gradients indicate that parameter adjustments have a significant effect on error correction, indicating room for optimization in the corresponding layout decision. The loss function gradient norm is positively correlated with the first scoring term.

[0039] For example, in a preferred embodiment, by the formula Calculate the first scoring item, where represents the first score item of sample i at iteration t, α is the smoothing coefficient (default 0.3), is the gradient norm of the loss function for sample i in the current batch. This formula maintains the stability of the difficulty evaluation through exponential weighted averaging, avoiding random fluctuations in a single calculation.

[0040] For example, the loss function gradient norm of a financial product instruction manual containing a complex nested structure (including structured derivative clauses) is relatively high.

[0041] When hospitals generate differentiated recruitment posters for different types of cancer (lung cancer / breast cancer), the gradient norm of complex samples containing a comparison table of pathological section microscopic images and genetic test results is 2 to 3 times higher than that of ordinary text posters.

[0042] S12: Determine a second scoring item based on the attention weight matrix of the sample.

[0043] In specific implementations, the second scoring term can be referred to as the layout complexity term. The attention weight matrix is ​​composed of the attention distribution entropy values ​​between the elements of the sample. The attention weight entropy value measures the complexity of the decision-making process regarding the relationship between elements. Higher entropy values ​​indicate that the spatial relationships between layout elements are more difficult to coordinate.

[0044] For example, in a preferred embodiment, by the formula Calculate the second scoring item.

[0045] Among them, A ij Represents the attention weight matrix between elements of the layout generation model output sample; the entropy function calculates the attention distribution entropy of each element. This indicator reflects the complexity of the layout decision. The higher the value, the more difficult it is to determine the element relationship; N represents the number of elements.

[0046] For example, a poster for a portfolio asset management product needs to simultaneously display multiple elements, including asset allocation ratios, historical drawdown data, and fund manager resumes. Using the above formula, we calculate the attention distribution entropy of each element in a sample containing five elements (product name, yield curve, risk matrix, compliance statement, and QR code).

[0047] When hospitals display combination therapy posters, they need to coordinate multiple pieces of information, including animated frames about drug mechanisms of action, lists of side effects, and enrollment flowcharts. The above formula is used to calculate the attention distribution entropy of each element in a sample containing eight elements (title, molecular diagram, flowchart, list of contraindications, etc.).

[0048] S13: Determine a third scoring item based on the aesthetic feature vector of the sample.

[0049] In practice, the third scoring term is called the aesthetic deviation term. It is determined based on the sample's aesthetic feature vector and evaluates layout quality from the perspective of visual rules, ensuring that the model not only learns spatial arrangements but also conforms to human aesthetic standards.

[0050] For example, in one embodiment, by formula d i =||v i -v ref ||1Calculate the third scoring item.

[0051] Among them, v i is the aesthetic feature vector of the current layout of the sample (including 6-dimensional indicators of hue, chroma, complementary color, purity, coldness and warmth, and negative space ratio), v ref The third scoring item quantifies the degree of deviation of the layout from the ideal aesthetic standard.

[0052] For example, when a bank generates a poster for a private VIP client, the reference standard vector is set to [Hue = Dark Blue (HSB 215°), Chroma = 85%, Complementary Color = Champagne Gold, Negative Space ≥ 30%]. If the generated poster shows excessive chroma (e.g., excessive red content) or insufficient negative space (information overload), the third score item will be significantly increased. If the poster's warm / cold color index deviates from the standard (the proportion of cool colors should be greater than 70%), the system automatically reduces the transparency of the warm-toned decorative elements until the aesthetic deviation is less than the threshold.

[0053] For example, when a hospital generates a poster for a pediatric clinical trial, the reference standard vector is set to [hue = light blue (soothing color), saturation ≤ 60%, purity = medium, negative space ≥ 40%]. If a high-saturation red warning frame (saturation = 90%) is used in the generated poster, the third scoring item will be automatically replaced with a soft orange border until the aesthetic deviation is less than the threshold.

[0054] S14: Determine the difficulty score based on the first scoring item, the second scoring item, and the third scoring item.

[0055] In a specific implementation, the difficulty score is determined based on the first scoring item, the second scoring item, and the third scoring item, and the three scores are summed up to avoid one-sided evaluation and greatly improve the accuracy of the evaluation.

[0056] For example, a weighted sum of the first scoring item, the second scoring item, and the third scoring item may be calculated as the difficulty score by using a weighted summation method.

[0057] Specifically, by the following formula S i =w1s i +w2c i +w3d i Calculate the difficulty score.

[0058] The weight coefficients w1, w2, and w3 are determined by adjusting the validation set (typical values ​​of w1, w2, and w3 are 0.5, 0.3, and 0.2).

[0059] S2. Divide the sample set into a challenge area and a diversity area according to the difficulty scores of the samples, wherein the difficulty scores of the samples in the challenge area are greater than the difficulty scores of the samples in the diversity area.

[0060] In a specific implementation, the sample set is divided into a challenge zone and a diversity zone according to the difficulty scores of the samples. The number of samples in the challenge zone and the diversity zone can be set by those skilled in the art and is not specifically limited in the present invention.

[0061] For example, in some preferred embodiments, the above step of "dividing the sample set into a challenge zone and a diversity zone according to the difficulty scores of the samples" specifically includes the following steps: dividing the samples whose difficulty scores are greater than a preset score threshold into the challenge zone; and dividing the samples whose difficulty scores are not greater than the preset score threshold into the diversity zone.

[0062] In a specific implementation, the score threshold can be set by those skilled in the art, and is not specifically limited in the present invention.

[0063] Furthermore, in some preferred embodiments, the method further comprises the following steps: clustering the samples in the diversity region to obtain a plurality of categories; and dividing the samples in each category into a sample group.

[0064] In a specific implementation, semantic features of samples in the diversity region are calculated. Using a pre-set clustering algorithm, the samples in the diversity region are clustered based on these features to obtain multiple categories. The samples in each category are then divided into sample groups. Clustering sample groups ensures that the samples in a sample group are similar.

[0065] Furthermore, in some preferred embodiments, the method further includes the following steps: for the samples in the challenge area, determining the sampling probability of the samples based on the difficulty score of the samples, wherein the sampling probability of the samples is positively correlated with the difficulty score of the samples; for the sample groups in the diversity area, determining the sampling probability of the sample groups based on the number of samples in the sample groups, wherein the sampling probability of the sample groups is negatively correlated with the number of samples in the sample groups.

[0066] In a specific implementation, for samples in the challenge area, the sampling probability calculation formula is as follows:

[0067]

[0068] Among them, τ h is the high difficulty threshold (the high difficulty threshold can be set by those skilled in the art, and is not specifically limited in the present invention. It is usually taken as the 75th percentile of the difficulty score), γ is the focusing strength (the default is 1.5, and the focusing strength can be set by those skilled in the art, and is not specifically limited in the present invention), ∈ is the threshold value (the threshold value can be set by those skilled in the art, and is not specifically limited in the present invention). The larger ∈ is, the smaller the sampling probability is.

[0069] The sampling probability of the sample is positively correlated with the difficulty score of the sample, so the sample with higher difficulty is selected first during sampling.

[0070] Furthermore, for the sample groups in the diversity zone, the sampling probability calculation formula is as follows:

[0071]

[0072] Among them, n g represents the number of samples in the g-th group. The sampling probability of the sample group is negatively correlated with the number of samples in the sample group, ensuring that the sample group with a small number of samples is not ignored.

[0073] S3: Collect training samples from the challenge area and the diversity area.

[0074] In specific implementations, when training the model, training samples are collected from the challenge zone and the diversity zone. The challenge zone focuses on layout combinations that are difficult for the current model to handle (such as asymmetric arrangements of multiple elements) and accelerates convergence through intensive training; the diversity zone prevents the model from falling into local optimal solutions by maintaining the distribution of samples of different semantic categories (such as posters of different styles). This dual-segment management mechanism achieves the optimal allocation of training resources, avoiding the redundancy of repeated training of simple samples, and ensuring that the model maintains generalization capabilities while overcoming difficulties, ultimately significantly improving layout generation quality and training efficiency.

[0075] For example, in some preferred embodiments, the above step of "collecting training samples from the challenge area and the diversity area" specifically includes the following steps:

[0076] S31, determining a sampling ratio of the challenge area based on the accuracy information of the model.

[0077] In a specific implementation, the sampling ratio of the challenge zone is determined based on the accuracy information of the model. The sampling ratio of the challenge zone is positively correlated with the accuracy of the model. The lower the accuracy of the model, the lower the sampling ratio of the challenge zone; the higher the accuracy of the model, the higher the sampling ratio of the challenge zone.

[0078] For example, in one embodiment, the sampling ratio of the challenge area is outputted through a policy network, and the policy of the policy network is:

[0079] λ=sigmoid(W2ReLU(W1x+b1)+b2)

[0080] in, is a learnable parameter, x is the input state vector, ReLU is the activation function, b1 and b2 are bias terms, and λ is the sampling ratio of the challenge area.

[0081] The input state vector includes: the average precision (mAP) of the current model on the validation set, the mean difficulty of the challenge area samples, the loss reduction rate of the last k batches, and the category distribution entropy of the diversity area samples.

[0082] W1, W2, b1 and b2 are all randomly initialized and then updated through model training.

[0083] Furthermore, the reward function of the policy network is designed as:

[0084]

[0085] Here, η is a tuning parameter used to balance quality and efficiency (default is 0.1). The number of valid samples refers to the number of samples for which the loss drops below a threshold. The goal of the policy network is to maximize the reward function, that is, to increase the proportion of valid samples. The model updates its parameters by maximizing the reward function r.

[0086] S32: Collect training samples from the challenge area and the diversity area based on the sampling ratio.

[0087] In specific implementations, the sampling ratio represents the proportion of samples in the challenge zone. For example, a sampling ratio of 20% means that 20% of the collected training samples are from the challenge zone, and the remaining 80% are from the diversity zone. The training samples are used to train the model.

[0088] An embodiment of the present invention proposes a model training sampling method, comprising: obtaining the difficulty score of each sample in a sample set; dividing the sample set into a challenge area and a diversity area according to the difficulty score of the sample, wherein the difficulty score of the sample in the challenge area is greater than the difficulty score of the sample in the diversity area; and collecting training samples from the challenge area and the diversity area. In the present invention, the challenge area focuses on layout combinations that are difficult for the current model to handle (such as asymmetric arrangement of multiple elements), and accelerates convergence through intensive training; the diversity area prevents the model from falling into a local optimal solution by maintaining the distribution of samples of different semantic categories (such as posters of different styles). This dual-segment management mechanism achieves the optimal allocation of training resources, avoids the redundancy of repeated training of simple samples, and ensures that the model maintains generalization ability while overcoming difficulties, ultimately significantly improving the layout generation quality and training efficiency.

[0089] See also Figure 2 , an embodiment of the present invention proposes a poster generation method, the method comprising the following steps:

[0090] S10, receiving poster design elements, extracting features of the poster design elements, and obtaining an input vector.

[0091] In a specific implementation, poster design elements include images and text, etc., which are not specifically limited in the present invention. Features of poster design elements are extracted using a neural network model. For example, for images, the CLIP model can be used to extract features, while for text, the ResNet-50 model can be used to extract features.

[0092] S20, inputting the input vector into a preset poster layout generation model, so that the poster layout generation model outputs poster layout parameters based on the input vector, wherein, during the training process of the poster layout generation model, training samples are collected based on the model training sampling method provided in any of the above embodiments.

[0093] In a specific implementation, the poster layout generation model can be specifically a Transformer model, which calculates the spatial relationship weights between elements through a cross-attention mechanism to obtain poster layout parameters. The Transformer model can specifically include a 12-layer Transformer encoder.

[0094] S30: Generate a poster file based on the poster layout parameters.

[0095] In specific implementations, a differentiable rendering module generates a poster file based on the poster layout parameters. This module converts the abstract relationships output by the layout generation module into specific coordinate parameters, supporting end-to-end gradient propagation. This module maps layout weights to actual rendering results.

[0096] Furthermore, an evaluation and feedback module is set up to calculate the multi-dimensional indicators of the layout quality of the poster file, including aesthetic score, balance coefficient and visual flow coherence, and provide optimization suggestions.

[0097] For example, in a medical application scenario, a hospital produces a diabetes prevention and treatment science poster, the content of which must include medical illustrations, core knowledge points, data charts, and a call to action.

[0098] The design elements received include:

[0099] Images: Medical illustrations of diabetic complications (diagram of retinopathy), healthy diet photography, hospital logo; Text: Core knowledge points (such as "Blood sugar control target value: fasting <7mmol / L"), calls to action ("Make an appointment for screening immediately"), data charts (screenshots of blood sugar monitoring statistics).

[0100] The CLIP model is used to perform cross-modal feature encoding of medical illustrations and photographs to capture semantic associations (such as "food" and "blood sugar control"). ResNet-50 is used to extract text features (such as highlighting the significance of the value "7mmol / L") and analyze key trends in data charts.

[0101] The Transformer encoder's cross-attention mechanism is used to model spatial relationships, assign visual priorities, and optimize data visualization to output poster layout parameters: the association weights between medical illustrations and corresponding knowledge point texts are calculated (for example, the retinopathy diagram and complication description text are placed adjacent to each other); the call to action text ("make an appointment for screening now") is assigned to the golden area at the top of the poster, and the hospital logo is fixed in the lower right corner to meet brand standards; the blood glucose statistics chart is automatically scaled to a reasonable size to avoid visual conflict with the illustrations.

[0102] Generate poster files through differentiable rendering: convert layout parameters into specific coordinates to generate PDF / PNG files (such as text wrapping around illustrations, and charts and text modules aligning to the grid system).

[0103] Through multi-dimensional evaluation and optimization suggestions:

[0104] Aesthetic scoring: Check whether the color scheme is consistent with the seriousness of the medical scene (such as avoiding high-saturation color blocks).

[0105] Balance factor: Ensure that the weight of the LOGO area is balanced with the core content to avoid top-heavy.

[0106] Visual flow coherence: Verify whether the reader's eyes can flow naturally from the title → medical illustration → data chart → call to action.

[0107] Optimization suggestions: If the assessment finds that the data chart is not readable enough, the system recommends increasing the font size or adding a high-contrast border.

[0108] For example, in a financial application scenario, a bank needs to quickly generate customized financial product promotional posters for different customer groups (such as conservative and aggressive investors). The content must include product return data, risk warnings, asset allocation charts and compliance statements.

[0109] The design elements received include:

[0110] Images: asset allocation ratio pie chart, historical yield curve, bank brand visual elements (logo / primary color); Text: key product information (such as "annualized yield 4.2%"), risk level description ("R2 medium-low risk"), compliance clauses ("past performance does not represent future performance"), etc.

[0111] The CLIP model is used to analyze the trend characteristics of the yield curve (such as the risk implied by volatility) and associate it with the "medium-low risk" description in the text; the ResNet-50 model extracts the numerical significance in the text (highlighting "4.2%") and identifies the legal term weight of the compliance clauses.

[0112] The Transformer encoder's cross-attention mechanism is used to assign information priorities, dynamically adapt to compliance, and integrate data visualization to output poster layout parameters: the "annualized rate of return" value that customers are most concerned about is enlarged and placed in the visual focus area, and the risk warning text is forced to retain the bottom 30% area; according to financial regulatory rules, the risk warning font size is automatically detected to see if it meets the minimum font size requirement (such as ≥120% of the main text); the asset allocation pie chart and the yield curve chart are laid out in contrasting colors to avoid overlapping chart information.

[0113] Render and output poster files: Generate files that meet multi-channel requirements (mobile H5 long image / printed PDF), adapting to different screen resolutions and printing DPI.

[0114] Through multi-dimensional evaluation and optimization suggestions:

[0115] Aesthetic score: Verify that the brand's primary colors (e.g., dark blue + gold) account for more than 60% of the content to maintain a sense of professional trust.

[0116] Balance factor: Check the visual weight ratio of return data and risk warning (regulatory requirement ≥1:1).

[0117] Visual flow consistency: Ensure that the reading path from product name → core income → asset allocation → risk warning complies with the regulatory disclosure order.

[0118] Optimization suggestion: If it is detected that the risk warning area is covered by decorative elements, the system will automatically add a semi-transparent background color to improve readability.

[0119] See also Figure 3 , Figure 3 is a schematic block diagram of a model training sampling device 30 provided in an embodiment of the present invention. Corresponding to the above model training sampling method, the present invention further provides a model training sampling device 30. The model training sampling device 30 includes a unit for executing the above model training sampling method. The model training sampling device 30 can be configured in a terminal or a server. Specifically, the model training sampling device 30 includes:

[0120] An acquisition unit 31 is used to obtain a difficulty score of each sample in the sample set;

[0121] a first dividing unit 32, configured to divide the sample set into a challenge zone and a diversity zone according to the difficulty scores of the samples, wherein the difficulty scores of the samples in the challenge zone are greater than the difficulty scores of the samples in the diversity zone;

[0122] The collection unit 33 is configured to collect training samples from the challenge area and the diversity area.

[0123] In some preferred embodiments, obtaining the difficulty score of each sample in the sample set includes:

[0124] Determining a first scoring item based on a gradient norm of a loss function of the sample;

[0125] Determining a second scoring item based on the attention weight matrix of the sample;

[0126] determining a third scoring item based on the aesthetic feature vector of the sample;

[0127] The difficulty score is determined based on the first scoring item, the second scoring item, and the third scoring item.

[0128] In some preferred embodiments, dividing the sample set into a challenge zone and a diversity zone according to the difficulty scores of the samples includes:

[0129] Classifying the samples whose difficulty scores are greater than a preset score threshold into the challenge area;

[0130] The samples whose difficulty scores are not greater than a preset score threshold are divided into the diversity zone.

[0131] In some preferred embodiments, the model training sampling device 30 further includes:

[0132] A clustering unit, configured to cluster the samples in the diversity region to obtain a plurality of categories;

[0133] The second division unit is used to divide the samples of each category into a sample group.

[0134] In some preferred embodiments, the model training sampling device 30 further includes:

[0135] a first determining unit, configured to determine, for a sample in the challenge area, a sampling probability of the sample based on a difficulty score of the sample, wherein the sampling probability of the sample is positively correlated with the difficulty score of the sample;

[0136] The second determining unit is configured to determine, for the sample group in the diversity region, a sampling probability of the sample group based on the number of samples in the sample group, wherein the sampling probability of the sample group is negatively correlated with the number of samples in the sample group.

[0137] In some preferred embodiments, collecting training samples from the challenge area and the diversity area includes:

[0138] Determining a sampling ratio of the challenge area based on the accuracy information of the model;

[0139] Training samples are collected from the challenge area and the diversity area based on the sampling ratio.

[0140] See also Figure 4 , Figure 4 is a schematic block diagram of a poster generation device 40 provided in an embodiment of the present invention. Corresponding to the above poster generation method, the present invention further provides a poster generation device 40. The poster generation device 40 includes a unit for executing the above poster generation method. The poster generation device 40 can be configured in a terminal or a server. Specifically, the poster generation device 40 includes:

[0141] An extraction unit 41 is configured to receive a poster design element, extract features of the poster design element, and obtain an input vector;

[0142] an output unit 42, configured to input the input vector into a preset poster layout generation model, so that the poster layout generation model outputs poster layout parameters based on the input vector, wherein, during the training process of the poster layout generation model, training samples are collected based on the model training sampling method provided in any embodiment of the present invention;

[0143] The generating unit 43 is configured to generate a poster file based on the poster layout parameters.

[0144] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned model training sampling device 30 and the poster generation device 40 can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of the description, it will not be repeated here.

[0145] The above-mentioned model training sampling device 30 and poster generating device 40 can be implemented in the form of a computer program. The computer program can be used in Figure 5 Runs on the computer device shown.

[0146] See also Figure 5 , Figure 5 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 can be a terminal or a server.

[0147] The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0148] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 may execute a model training sampling method and / or a poster generation method.

[0149] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0150] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a model training sampling method and / or a poster generation method.

[0151] The network interface 505 is used to communicate with other devices over the network. Those skilled in the art will appreciate that the above structure is merely a block diagram of a portion of the structure related to the present invention and does not limit the computer device 500 to which the present invention is applied. A specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0152] The processor 502 is configured to run a computer program 5032 stored in the memory to implement steps of a model training sampling method and / or a poster generation method provided in an embodiment of the present invention.

[0153] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0154] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0155] Therefore, the present invention further provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program. When executed by a processor, the computer program causes the processor to perform the steps of a model training sampling method and / or a poster generation method provided in embodiments of the present invention.

[0156] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, etc. Any physical storage medium capable of storing program code can be non-volatile or volatile.

[0157] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0158] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0159] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0160] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0161] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0162] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, to the extent such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

[0163] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A model training sampling method, characterized in that: include: Get the difficulty score of each sample in the sample set; Dividing the sample set into a challenge zone and a diversity zone according to the difficulty scores of the samples, wherein the difficulty scores of the samples in the challenge zone are greater than the difficulty scores of the samples in the diversity zone; Training samples are collected from the challenge area and the diversity area.

2. The model training sampling method according to claim 1, characterized in that Obtaining the difficulty score of each sample in the sample set includes: Determining a first scoring item based on a gradient norm of a loss function of the sample; Determining a second scoring item based on the attention weight matrix of the sample; determining a third scoring item based on the aesthetic feature vector of the sample; The difficulty score is determined based on the first scoring item, the second scoring item, and the third scoring item.

3. The model training sampling method according to claim 1, characterized in that The step of dividing the sample set into a challenge zone and a diversity zone according to the difficulty scores of the samples includes: Classifying the samples whose difficulty scores are greater than a preset score threshold into the challenge area; The samples whose difficulty scores are not greater than a preset score threshold are divided into the diversity zone.

4. The model training sampling method according to claim 1, characterized in that The method further comprises: Clustering the samples in the diversity region to obtain multiple categories; Divide the samples of each category into a sample group.

5. The model training sampling method according to claim 4, characterized in that: The method further comprises: For samples in the challenge area, determining a sampling probability of the sample based on the difficulty score of the sample, wherein the sampling probability of the sample is positively correlated with the difficulty score of the sample; For the sample groups in the diversity region, a sampling probability of the sample group is determined based on the number of samples in the sample group, wherein the sampling probability of the sample group is negatively correlated with the number of samples in the sample group.

6. The model training sampling method according to claim 1, characterized in that The collecting of training samples from the challenge area and the diversity area includes: Determining a sampling ratio of the challenge area based on the accuracy information of the model; Training samples are collected from the challenge area and the diversity area based on the sampling ratio.

7. A poster generation method, characterized in that: include: Receiving poster design elements, extracting features of the poster design elements, and obtaining an input vector; Inputting the input vector into a preset poster layout generation model, so that the poster layout generation model outputs poster layout parameters based on the input vector, wherein, during the training process of the poster layout generation model, training samples are collected based on the model training sampling method according to any one of claims 1 to 6; A poster file is generated based on the poster layout parameters.

8. A model training sampling device, characterized in that: The method comprises a unit for executing the method according to any one of claims 1 to 7.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the computer program can implement the method according to any one of claims 1 to 7.