An aesthetic feature modeling method of generative artificial intelligence
Patent Information
- Application Number
- CN202610715070.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-25
AI Technical Summary
但现有相关技术存在四大核心缺陷,无法满足产业落地的实际需求:
本发明实现了具备严谨因果支撑的美学特征建模,通过流形约束的反事实因果边界估计,精准定位影响美学感知的真实因果变量,完全剔除虚假相关的混淆变量,从根源上避免了传统技术的因果混淆问题,模型的特征选择与权重分配均具备明确的因果依据,可解释性得到本质提升,能够清晰地揭示生成式人工智能的美学生成底层逻辑。
Smart Images

Figure CN122819341A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for modeling aesthetic features in generative artificial intelligence. Background Technology
[0002] With the rapid deployment and application of generative artificial intelligence (AI) technology, the aesthetic quality of AI-generated content has become a key indicator determining a product's core competitiveness. Modeling the aesthetic features of generative AI has become the core foundation for achieving controllable aesthetic generation, automated quality assessment of generated content, and optimization of model aesthetic capabilities. However, existing technologies suffer from four major shortcomings, failing to meet the actual needs of industrial application: First, existing technologies generally suffer from the core problem of confusing correlation with causation. Most of them rely on statistical correlation analysis to screen aesthetically relevant features, but cannot distinguish between causal variables that have a real impact on aesthetic perception and confounding variables with false correlations. This results in extremely poor model generalization ability, and the model fails after changing the generation task or aesthetic style. At the same time, the model has a serious lack of interpretability and cannot locate the core factors that truly determine the aesthetic quality of the generated content.
[0003] Second, existing technologies generally suffer from the problem of complete loss of temporal evolution features. Most solutions only extract static features for the final generated result, completely ignoring that the aesthetic expression of generative AI is a temporal evolution process that runs through the entire generation trajectory. Static features cannot capture the aesthetic laws brought about by changes in hidden states, attention convergence, and probability distribution adjustment during the generation process. The modeling accuracy is inherently limited and cannot reflect the true aesthetic generation logic of generative AI.
[0004] Third, existing technologies cannot effectively eliminate the heterogeneity of multi-subject aesthetic perception. For the fusion of multi-subject aesthetic perception data, most methods adopt simple average weighting and manual subjective assignment, which cannot capture the perceptual differences of different subjects (model self-consistency, ordinary user groups, professional aesthetic practitioners, machine cognitive systems). The resulting weight allocation has no broad consensus, serious subjective bias, and no rigorous theoretical and causal basis. It is only the result of data fitting and has extremely poor robustness.
[0005] Fourth, the existing modeling process is disconnected and lacks closed-loop optimization capabilities. Each step, from feature extraction and weight allocation to model verification, operates independently and unidirectionally. Errors in the preceding steps accumulate and amplify continuously, and optimization signals from subsequent steps cannot be fed back to the preceding steps to achieve collaborative optimization. This results in extremely low upper limits for model performance and an inability to balance quantization accuracy and interpretability.
[0006] Furthermore, existing technologies generally employ conventional Euclidean space metrics, linear regression fitting, and general time-series coding, which cannot adapt to the structural characteristics of the high-dimensional nonlinear latent space of generative AI. There are very few innovative cross-disciplinary technologies, and they cannot break through the core bottlenecks of existing technologies. Summary of the Invention
[0007] This invention provides a method for modeling aesthetic features in generative artificial intelligence, which addresses the core pain points of existing technologies such as causal confusion, loss of temporal information, inability to resolve perceptual heterogeneity, and process disconnect. It enables interpretable, quantifiable, and generalizable modeling of aesthetic features in generative AI, and provides core technical support for optimizing the aesthetic capabilities of various generative artificial intelligence systems, controlling aesthetic generation, and automatically evaluating the quality of generated content.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: A generative artificial intelligence method for modeling aesthetic features includes the following steps: S1: Collect full-time generation trajectory data and final generation results corresponding to the generation task of the generative artificial intelligence target model, and construct the generation trajectory dataset, the generation result benchmark set and the Riemannian manifold benchmark space; S2: Based on the generated trajectory dataset, the generated result benchmark set, and the Riemann manifold benchmark space, the counterfactual causal boundary estimation algorithm under Riemann manifold constraints is used to complete the causal effect boundary estimation of candidate variables. Based on the boundary estimation results, the aesthetic core causal variable set is obtained and the boundary constraint matrix is generated. S3: Based on the generated trajectory dataset, Riemannian manifold reference space, aesthetic core causal variable set and boundary constraint matrix, the orthogonal projection coding algorithm of density matrix temporal evolution is adopted to complete the temporal evolution modeling and feature encoding of the generated trajectory, and construct a standardized aesthetic feature dynamic primitive library and aesthetic primitive coding matrix; S4: Aesthetic scoring is performed on the generated benchmark set to obtain a multi-subject aesthetic perception dataset. Then, based on the aesthetic primitive encoding matrix and the multi-subject aesthetic perception dataset, a stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm is used to fit the comprehensive weight of aesthetic primitives of multi-subject consensus and construct an aesthetic feature hierarchical structure model. S5: Perform full-dimensional validity verification on the hierarchical structure model of aesthetic features, and correct the applicable boundaries of the model based on the verification results to obtain a qualified generative artificial intelligence aesthetic feature quantification model. S6: Based on the fusion framework of the counterfactual causal boundary estimation algorithm under Riemannian manifold constraints, the orthogonal projection coding algorithm of density matrix temporal evolution, and the stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm, an incremental update and lifelong learning mechanism for the generative artificial intelligence aesthetic feature quantification model is constructed.
[0009] In this specification, the specific generation tasks of the generative artificial intelligence target model adapted in S1 are as follows: for text generation models, it covers all categories of text generation tasks; for image generation models, it covers all categories of image generation tasks; for audio generation models, it covers all categories of audio generation tasks; and for multimodal generation models, it covers all corresponding single-modal and cross-modal generation tasks.
[0010] In this specification, the counterfactual causal boundary estimation algorithm under Riemannian manifold constraints in S2 synthesizes counterfactual samples through exponential mapping on the Riemannian manifold, ensuring semantic consistency between the counterfactual samples and the original samples. Based on the counterfactual samples, the upper and lower boundaries of the average treatment effect of each candidate variable are calculated. Variables whose boundary intervals do not contain 0 are selected to form the aesthetic core causal variable set, and aesthetically irrelevant confounding variables whose boundary intervals contain 0 are removed.
[0011] In this specification, the orthogonal projection coding algorithm of the density matrix temporal evolution in S3 models the temporal evolution process of the generated trajectory as the unitary evolution process of the density matrix. The boundary constraint matrix is used as the projection space constraint to ensure that the coding process retains only the features corresponding to the core aesthetic causal variables. Mutually exclusive and non-redundant aesthetic feature dynamic primitives are extracted through pairwise orthogonal projection operators to complete the feature coding.
[0012] In this specification, the S3 standardized aesthetic feature dynamic primitive library includes temporal smoothness primitive, attention convergence primitive, manifold curvature consistency primitive, probability distribution entropy change primitive, cross-layer feature alignment primitive, and cross-modal stability primitive. The projection operators corresponding to all primitives are pairwise orthogonal.
[0013] In this specification, the stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm in S4 uses the causal effect boundary estimation result as the prior distribution constraint of the hierarchical Bayesian model. It captures the aesthetic perception heterogeneity of different rating subjects through stochastic coefficient fitting, and then obtains the comprehensive weight of aesthetic primitives of multi-subject consensus through Nash equilibrium fitting, ensuring that the weight allocation matches the causal effect strength of the variables.
[0014] In this specification, the S4 multi-subject aesthetic perception dataset is obtained by standardizing the aesthetic scores of each generated result in the benchmark set of generated results through four complementary perceptual subjects. The four perceptual subjects are the self-consistency subject of the generative artificial intelligence target model, the human perception group subject, the machine aesthetic cognition system subject, and the professional aesthetic practitioner subject. All scores are strongly bound to the unique identifier of the corresponding generation trajectory.
[0015] In this manual, the full-dimensional validity verification in S5 includes four dimensions: causal validity verification, generalization validity verification, discriminant validity verification, and stability verification. Each dimension has a corresponding pass threshold. If any dimension fails the verification, the corresponding preceding steps are returned to complete the model optimization until all dimensions pass the verification.
[0016] In this specification, the incremental update and lifelong learning mechanism in S6 is implemented by constructing a dynamically updated incremental learning sample pool. The data in the sample pool comes from user-authorized and privacy-de-identified model interaction log data. Implicit aesthetic feedback annotation of samples is completed based on user interaction behavior, and an elastic weight consolidation algorithm is used to complete the incremental update of the model. While retaining the original core capabilities, it adapts to the newly added generation tasks and aesthetic styles.
[0017] In summary, the present invention has at least the following beneficial effects: This invention achieves aesthetic feature modeling with rigorous causal support. By estimating the counterfactual causal boundary with manifold constraints, it accurately locates the real causal variables that affect aesthetic perception and completely eliminates falsely related confounding variables. This fundamentally avoids the causal confusion problem of traditional techniques. The feature selection and weight allocation of the model have clear causal basis, and the interpretability is fundamentally improved. It can clearly reveal the underlying logic of aesthetic generation in generative artificial intelligence.
[0018] This invention fully preserves the temporal aesthetic information of the generation process, breaks through the essential limitations of traditional static feature modeling, and accurately models the temporal evolution process of the generation trajectory through unitary evolution of the density matrix. It fully captures the aesthetic laws brought about by dynamic processes such as latent state changes and attention convergence during the generation process, and greatly improves the ability to characterize and model the aesthetic features of generative AI.
[0019] This invention effectively eliminates the heterogeneity of multi-subject aesthetic perception. It accurately captures the differences in aesthetic perception among different subjects through a hierarchical Bayesian model and obtains the weight allocation of multi-subject consensus through Nash equilibrium fitting. It completely avoids the subjective bias problem of traditional technology. The constructed aesthetic model has broad group consensus and can adapt to the aesthetic perception needs of different scenarios and groups.
[0020] This invention achieves closed-loop collaborative optimization throughout the entire process. The three core algorithms interact bidirectionally in pairs, and the optimization signals of subsequent steps can be fed back to the preceding steps, forming a complete closed-loop optimization framework. This completely solves the problems of disconnected steps and error accumulation in traditional technologies. The overall performance and generalization ability of the model are significantly improved, and it can adapt to the generative AI aesthetic modeling needs of different modalities, tasks, and styles.
[0021] This invention possesses long-term stable scene adaptation capabilities and constructs an incremental update and lifelong learning mechanism based on implicit user feedback. It can adapt to new generation tasks, new aesthetic styles, and dynamic changes in user preferences while fully preserving the original core capabilities of the model. This avoids the performance degradation problem that traditional models experience over time and has extremely strong industrial application and long-term application value. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the aesthetic feature modeling method for generative artificial intelligence involved in this invention.
[0024] Figure 2 This is a flowchart illustrating the counterfactual causal boundary estimation algorithm under Riemannian manifold constraints involved in this invention.
[0025] Figure 3 This is a flowchart illustrating the orthogonal projection coding algorithm for the temporal evolution of the density matrix involved in this invention.
[0026] Figure 4 This is a schematic diagram illustrating the process of the stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm and aesthetic feature hierarchical modeling involved in this invention. Detailed Implementation
[0027] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0028] The following disclosure provides many different implementations or examples for carrying out different structures of the embodiments of the present invention. To simplify the disclosure of the embodiments of the present invention, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the embodiments of the present invention. Furthermore, reference numerals and / or reference letters may be repeated in different examples of the embodiments of the present invention; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various implementations and / or arrangements discussed.
[0029] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0030] like Figure 1 As shown, this embodiment provides a generative artificial intelligence aesthetic feature modeling method, including: S1: Full-scale trajectory acquisition and Riemannian manifold reference space construction This step completes three core tasks: full-scene generation task coverage, full-time trajectory acquisition, and construction of a benchmark metric space adapted to the latent space structure of the generation model. This provides a unified data source and metric benchmark for all causal analysis, feature encoding, and weight fitting.
[0031] First, for the generative AI target model to be modeled, generation tasks are matched. For text generation models, the tasks include narrative text generation, poetry text generation, advertising copy generation, academic text generation, dialogue text generation, script generation, official document text generation, and creative copywriting generation. For image generation diffusion models, the tasks include realistic image generation, impressionistic image generation, abstract image generation, industrial design image generation, scene design image generation, portrait generation, landscape design generation, and illustration design generation. For audio generation models, the tasks include pure music generation, song melody generation, speech synthesis generation, environmental sound effect generation, film and television score generation, and radio drama sound effect generation. For multimodal generation models, the tasks include image-text matching generation, audio-visual synchronization generation, video clip generation, and digital human motion-driven generation, simultaneously covering the above-mentioned text, image, and audio generation tasks. The sample size for each type of generation task is no less than 100,000 to ensure sufficient statistical power for subsequent analysis.
[0032] For each generation task execution process of the generative AI target model, the native single-step decoding or denoising step size of the model is used as the smallest data acquisition unit. All temporal data for each step is fully collected, and all collected data is strongly bound to the unique identifier of the generation task and the generation step size number. The aesthetic characteristics of generative AI are not only reflected in the final generation result, but also permeate the temporal evolution trajectory throughout the entire generation process. Changes in the hidden state, convergence of attention, and adjustments to the probability distribution at each step directly affect the aesthetic presentation of the final result. Collecting only the final result would lose more than 90% of the aesthetically relevant temporal information. Therefore, the generation step size of a single task covers all native decoding or denoising steps of the model without truncation. For different modalities, the collected content is matched one-to-one with the modality. For text generation models, the token probability distribution tensor, the hidden state tensors of each layer of the decoder, the multi-head attention weight matrix, and the token sampling mask are collected at each decoding step. For image generation models, the noise prediction tensor, the hidden feature tensors of each layer of UNet, the cross-attention weight matrix, and the noise scheduling coefficients of the sampling step are collected at each denoising step. For audio generation models, the Mel spectrum tensor generated at each step, the hidden state tensors of each layer of the decoder, the attention weight matrix, and the sampling probability distribution tensor are collected. For multimodal generation models, all the above-mentioned text, image, and audio content are collected, and an additional cross-modal alignment feature mapping tensor is collected. Finally, the generated trajectory time series dataset is obtained.
[0033] While collecting the generated trajectories, the final generated result corresponding to each generated trajectory is collected simultaneously to construct a benchmark set of generated results. Semantic annotation, modal attribute annotation, and generation parameter annotation are completed for each result in the benchmark set. All annotation results are strongly bound to the unique identifier of the corresponding generated trajectory.
[0034] To construct a Riemannian manifold baseline space, traditional aesthetic modeling commonly uses Euclidean space as the metric space for hidden states. However, the hidden states of generative artificial intelligence models are essentially high-dimensional nonlinear manifold structures, and calculating straight-line distances in Euclidean space introduces significant semantic biases, failing to accurately reflect the semantic relationships between two hidden states. Therefore, a Riemannian manifold learning algorithm is employed to map the collected full set of hidden state tensors to a unified, complete Riemannian manifold. Let be the complete Riemannian manifold corresponding to the hidden state space of the generative artificial intelligence target model. For Riemannian manifold The geodesic distance between the two points is used to measure the true semantic distance between the two hidden states in the manifold space. The origin of the Riemannian manifold reference space is the total hidden state tensor in the manifold. The Fraser mean on the coordinate axis is a manifold. The cumulative variance contribution rate of the principal geodesic direction is not less than 95%, ensuring that all generated trajectories have uniform comparability in the same space, and providing a unified measurement benchmark for subsequent calculations.
[0035] S2: Counterfactual Causal Boundary Estimation and Selection of Core Aesthetic Variables Based on the Riemannian manifold reference space, the full-volume generated trajectory time-series dataset, and the generated result reference set, a counterfactual causal boundary estimation algorithm under Riemannian manifold constraints (RMCCBE algorithm) is introduced to locate the core causal variables of aesthetics and eliminate irrelevant confounding variables. Traditional methods often include noisy variables with spurious correlations to aesthetic labels in the modeling, resulting in extremely poor generalization ability and lack of interpretability in the final model. This step, through counterfactual causal analysis, retains only variables with real causal effects on aesthetic perception, ensuring the effectiveness of the modeling from the root, while providing strict causal constraints for feature encoding. Furthermore, through a two-way feedback mechanism, it receives optimization signals to improve its own accuracy. The counterfactual causal boundary estimation process under Riemannian manifold constraints is as follows: Figure 2 As shown.
[0036] The core construction logic of the RMCCBE algorithm is to locate variables that have a real causal effect on the aesthetic perception of the generated results through nonparametric counterfactual boundary estimation under the constraint of the Riemannian manifold reference space, and completely eliminate confounding variables without causal effect, thus solving the problem that traditional linear causal models cannot adapt to the nonlinear manifold structure of the latent space of generative AI. To determine the total number of generated trajectories in the generated trajectory time series dataset, the first... This represents the total number of generation steps contained in a single generated trajectory. For the generative artificial intelligence target model in the first The first generated trajectory The hidden state tensor corresponding to each generation step size, where , ,all All are Riemannian manifolds The point on, For the first The aesthetic perception label for the final generated result corresponding to each generated trajectory ranges from [0,1]. A higher value indicates a better aesthetic perception evaluation. For the first The first generated trajectory The generation step size of the first The intervention status of each candidate variable takes a value of 0 or 1. A value of 1 indicates that a counterfactual intervention is performed on that variable, while a value of 0 indicates that no intervention is performed and the original state is maintained. This is the boundary constraint matrix for the aesthetic core causal variable set output by the RMCCBE algorithm. Each element in the matrix corresponds to a causal effect boundary estimate of a candidate variable. is the Riemannian manifold regularization coefficient of the RMCCBE algorithm, used to control the strength of manifold constraints, and its value ranges from (0,1]. This is the loss function of the RMCCBE algorithm, used to optimize the accuracy of causal boundary estimation.
[0037] The RMCCBE algorithm first synthesizes counterfactual samples under manifold constraints. While ensuring semantic consistency between the counterfactual samples and the original samples, it synthesizes counterfactual samples that conform to the intervention state, avoiding the semantic shift problem caused by traditional Euclidean space synthesis. The geodesic distance between the counterfactual samples and the original samples must not exceed a preset threshold, and the core semantic similarity between the counterfactual samples and the original samples must not be lower than a preset threshold. This ensures that the intervention only changes the target variable and does not alter the core semantics of the generated content, thus eliminating the interference of semantically confusing variables on the estimation of causal effects. The calculation formula for counterfactual sample synthesis is as follows: ; in the formula For the first The first generated trajectory The generation step size of the first The counterfactual hidden state tensors of the candidate variables after intervention are Riemannian manifolds. The point on, For Riemannian manifold above The exponential mapping with base points is used to map vectors in the tangent space of a manifold back to points on the manifold, ensuring that counterfactual samples are always on the manifold. Within this space, it will not deviate from the effective latent space of the generative model, thus avoiding the synthesis of meaningless outliers. The step size for intervention is the standard deviation of the corresponding candidate variable across the entire dataset. For Riemannian manifold above The unit vector in the tangent space of the base point, with direction i. The gradient direction of each candidate variable.
[0038] After completing the counterfactual sample synthesis, nonparametric causal boundary estimation is performed. Traditional parametric causal estimation requires pre-setting the model form, which easily introduces model specification bias. Nonparametric boundary estimation, however, can obtain rigorous upper and lower bounds of causal effects without pre-setting variable distributions, ensuring the reliability of variable selection. The formula for calculating the average treatment effect boundary estimate is as follows: ; ; in the formula For the first The lower bound estimate of the average treatment effect of the candidate variables. For the first The upper bound estimate of the average treatment effect of the candidate variables. For the first The first generated trajectory After intervention on each candidate variable, the predicted aesthetic perception label value of the corresponding generated result. For the first The first generated trajectory The predicted aesthetic perception label values of the corresponding generated results when no intervention is performed on the candidate variables. For the first The boundary estimation error of each candidate variable is taken as the standard error at a 95% confidence level.
[0039] After estimating the causal effect boundary, confounding variable removal and boundary constraint matrix construction are performed. The selection criteria for variables are: a lower bound estimate of the average treatment effect greater than 0, or an upper bound estimate less than 0, and the boundary interval does not contain 0, indicating that the causal effect of that variable is statistically significant. Boundary constraint matrix. The construction calculation formula is as follows: ; in the formula Boundary constraint matrix The There are 10 elements. A value of 0 indicates that the variable is an aesthetically irrelevant confounding variable and will be completely removed later. A value greater than 0 indicates that the variable is a core causal variable in aesthetics. The value represents the degree of uncertainty of the causal effect of the variable. The smaller the value, the more accurate the causal effect estimate.
[0040] The RMCCBE algorithm employs an iterative optimization approach during training, with the core optimization objective being to minimize the loss function. This simultaneously satisfies the Riemannian manifold constraint and incorporates bidirectional feedback signals from subsequent steps, achieving co-optimization with the two following algorithms. Loss function The calculation formula is as follows: ; in the formula The total number of candidate variables. The weighting coefficients for the separability index in subsequent feature encoding backfeedback are in the range of (0,1]. The primitive encoding separability index output by the orthogonal projection encoding algorithm for subsequent density matrix temporal evolution is used to back-optimize the variable selection accuracy of the RMCCBE algorithm, realizing bidirectional interaction with subsequent feature encoding. These are the weighting coefficients for the causal consistency index of the subsequent multi-agent weighted fitting back feedback, with values ranging from (0,1]. The multi-agent causal consistency index output by the subsequent stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm is used to back-optimize the manifold regularization strength of the RMCCBE algorithm, realizing bidirectional interaction with the subsequent weight fitting.
[0041] The iterative steps of the training process begin with initializing the Riemannian regularization coefficients. Reverse feedback weighting coefficient and The initial iteration count is set to 0, and the maximum iteration count is set to 100. The second step is based on the current boundary constraint matrix. The third step involves synthesizing counterfactual samples and calculating the upper and lower bounds of the average treatment effect for each candidate variable; the third step is to calculate the loss function. The values are then used to update the regularization coefficients using the Riemannian gradient descent algorithm. With boundary constraint matrix The numerical value; the fourth step is to receive the latest primitive encoding separability index from the subsequent feature encoding output. The latest multi-agent causal consistency index output by subsequent weight fitting The fifth step is to determine whether the iteration has converged, with the convergence condition being the loss function. If the change is less than 1e-6, or the maximum number of iterations is reached, and convergence is not achieved, return to the second step to continue iterating; if convergence is achieved, the training process ends, and the final boundary constraint matrix is output. The core causal variable set for aesthetics, the set of confounding variables irrelevant to aesthetics, and the estimated boundary of the average treatment effect.
[0042] The trained RMCCBE algorithm has three core contributions. First, it accurately identifies the core variables that have a real causal effect on aesthetic perception, completely eliminating spurious confounding variables and ensuring the effectiveness of subsequent modeling from the root. Second, the constructed boundary constraint matrix provides strict causal constraints for subsequent feature encoding, ensuring that the encoding process only focuses on variables with causal effects and avoids interference from noisy features. Third, through a two-way feedback mechanism, it forms a collaborative optimization with the two subsequent core algorithms, allowing the selection of causal variables to simultaneously consider causal effects, encoding separability, and consistency of multi-subject perception, thus solving the bias problem that exists in single-step optimization.
[0043] In some embodiments, the Riemannian regularization coefficients in the RMCCBE algorithm are... Its preferred value range is 0.3 to 0.7; further, in the embodiment of aesthetic feature modeling for text generation models, The optimal value is 0.45; in the embodiment of aesthetic feature modeling for image generation diffusion model, The optimal value is 0.6; in the embodiment of aesthetic feature modeling for audio generation models, The optimal value is 0.35; in the embodiment of aesthetic feature modeling for multimodal generative models, The optimal value is 0.55. This embodiment matches different modal models with differentiated values. The optimal value maximizes the accuracy of the causal effect boundary estimation while ensuring semantic consistency between the counterfactual sample and the original sample.
[0044] In some embodiments, the weighting coefficients for the primitive encoding separability index in the RMCCBE algorithm. Its preferred value range is 0.2 to 0.5; further, in the embodiment of the model pre-training stage, The optimal value is 0.25 to prioritize the rigor of causal variable selection and anchor the selection boundary of the core causal variables in aesthetics; in the embodiment of the model fine-tuning and optimization stage, The optimal value is 0.45, which strengthens the reverse optimization effect of feature coding separability on causal variable screening, improves the efficiency of whole-process collaborative optimization, and realizes bidirectional adaptation between causal variable screening and feature coding.
[0045] S3: Temporal Evolutionary Orthogonal Projection Encoding and Aesthetic Feature Primitive Construction The aesthetic core causal variable set and boundary constraint matrix obtained based on S2 screening This paper utilizes the S1 Riemannian manifold reference space and the full generated trajectory data, and introduces the orthogonal projection coding algorithm (DMTE-OPC algorithm) based on the temporal evolution of the density matrix to extract the temporal evolution features of the generated trajectory and construct standardized aesthetic feature dynamic primitives. Traditional methods often simplify the generated trajectory to a static feature vector, losing the core temporal evolution information. This step, however, models the temporal process of trajectory generation through the unitary evolution of the quantum density matrix, extracts mutually exclusive and complete aesthetic feature dynamic primitives through orthogonal projection coding, provides standardized feature inputs for multi-agent weight fitting, and optimizes the accuracy of prior causal variable selection and subsequent weight fitting through a bidirectional feedback mechanism. The orthogonal projection coding process based on the temporal evolution of the density matrix is as follows: Figure 3 As shown.
[0046] The core construction logic of the DMTE-OPC algorithm is to map each hidden state of the generated trajectory to a quantum density matrix, model the temporal evolution of the generated trajectory as a unitary evolution of the density matrix, and extract the dynamic primitives with aesthetic features that have temporal evolution characteristics through orthogonal projection encoding. The temporal evolution of the generated trajectory has a natural isomorphism with the unitary evolution of the quantum system. The density matrix can simultaneously represent the amplitude information and correlation information of the hidden state. The unitary evolution can accurately model the temporal evolution law of the generation process without losing temporal information, which is something that traditional temporal coding algorithms cannot achieve.
[0047] For the first The first generated trajectory The hidden state density matrix corresponding to each generation step size is a positive semi-definite Hermitian matrix with a trace of 1, and its dimension is consistent with the dimension of the core causal variable set of aesthetics. For the first The unitary evolution operator corresponding to each generation step size satisfies ,in It is the identity matrix. The sign for conjugate transpose. For the first The orthogonal projection operator corresponding to each aesthetic feature dynamic primitive satisfies , Furthermore, the projection operators corresponding to different primitives are pairwise orthogonal, i.e. when hour, This is the aesthetic primitive encoding matrix output by the DMTE-OPC algorithm. Each row in the matrix corresponds to the full aesthetic primitive encoding result of a generated trajectory. These are the causal constraint regularization coefficients for the DMTE-OPC algorithm, used to control the boundary constraint matrix. The constraint strength for the projection operator takes values in the range (0,1]. This is the loss function of the DMTE-OPC algorithm, used to optimize the accuracy of orthogonal projection coding. The total number of dynamic primitives representing aesthetic features.
[0048] The DMTE-OPC algorithm first performs the hidden state density matrix mapping. The core of this step is to map the hidden state tensors on the Riemannian manifold into a density matrix that conforms to quantum mechanics rules. Simultaneously, it completely eliminates aesthetically irrelevant confounding variables, retaining only the dimensions corresponding to the core aesthetic causal variables, ensuring that subsequent encoding processes focus only on features with causal effects. The formula for calculating the density matrix mapping is as follows: ; in the formula For the first The first generated trajectory The hidden state tensor of each generation step size, after removing confounding variables, becomes a core variable vector. Only the dimensions corresponding to the core causal variable set of aesthetics are retained, while the dimensions corresponding to all aesthetically irrelevant confounding variables are removed. The trace operation is used to calculate the sum of the elements on the main diagonal of a matrix.
[0049] After completing the density matrix mapping, the unitary evolution modeling of the density matrix is performed. The core of this step is to model the temporal evolution process of the generated trajectory as a unitary evolution process of the density matrix, capturing the dynamic changes of aesthetic features during the generation process, while introducing boundary constraint matrices. As a constraint, the evolutionary process is ensured to focus only on variables with causal effects. The formula for calculating unitary evolution is as follows: ; in the formula For the first The density matrix corresponding to the initial hidden state of each generated trajectory, i.e., the density matrix corresponding to the initial input of the generation process. Unitary evolution operator. The constraint condition is that its action space contains only the boundary constraint matrix. The dimensions corresponding to variables with values greater than 0 are used to ensure that the evolution process completely eliminates the interference of irrelevant and confusing variables.
[0050] After completing the temporal evolution modeling, orthogonal projection primitive encoding is performed. This step uses pairwise orthogonal projection operators to project and encode the density matrix of the temporal evolution, extracting mutually exclusive and complete aesthetic feature dynamic primitives. The core advantage of orthogonal projection is that it ensures no information overlap between different primitives, with each primitive corresponding to an independent aesthetic dimension. The aesthetic feature dynamic primitives are: temporal smoothness primitive, attention convergence primitive, manifold curvature consistency primitive, probability distribution entropy change primitive, cross-layer feature alignment primitive, and cross-modal stability primitive. The formula for calculating the encoded value of each aesthetic feature dynamic primitive is as follows: ; in the formula For the first The first generated trajectory corresponds to the The encoded values of each aesthetic feature dynamic primitive are in the range [0,1]. Aesthetic primitive encoding matrix. The construction method is as follows: the first... Line number The elements of the column are ,Right now .
[0051] The DMTE-OPC algorithm employs an iterative optimization approach during training, with the core optimization objective being to minimize the loss function. It simultaneously satisfies orthogonal projection constraints and causal boundary constraints, and incorporates bidirectional feedback signals from preceding and subsequent steps, achieving collaborative optimization with the other two core algorithms. Loss function The calculation formula is as follows: ; in the formula For the first The mean value of the encoded values of each aesthetic feature dynamic primitive on the full generation trajectory. Operations to convert a vector into a diagonal matrix, The Frobenius norm of a matrix is used to measure the degree of dissimilarity of the matrix. The constraint coefficients for the equilibrium weights in the subsequent multi-agent weight fitting back feedback are in the range of (0,1]. This represents the total number of subjects in the subsequent multi-subject perception system. The output of the subsequent stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm is the first... The equilibrium weight coefficients of each subject For the first The first subject corresponds to the first The optimal projection operator for each primitive is obtained by fitting subsequent multi-subject perception data and is used to back-optimize the projection operator of the DMTE-OPC algorithm, realizing bidirectional interaction with subsequent weight fitting.
[0052] During training, the separability index of primitive codes is calculated synchronously. This is used in the RMCCBE algorithm for backfeedback to S2, enabling bidirectional interaction with the selection of preceding causal variables. Primitive encoding separability index. The calculation formula is as follows: ; in the formula The inter-class scatter matrix encodes aesthetic primitives, with the categories divided based on the aesthetic perception labels of the generated results. Grouping by high and low, The intra-class scatter matrix encoding the aesthetic primitives. The larger the value, the stronger the ability of the aesthetic primitive encoding to distinguish between high and low aesthetic perception results, and the better the encoding effect.
[0053] The iterative steps of the training process strictly follow closed-loop optimization logic. The first step is to initialize the causal constraint regularization coefficients. Reverse feedback constraint coefficient The initial iteration count is set to 0, the maximum iteration count is set to 100, and the orthogonal projection operator is initialized. The first step is to use a random orthogonal matrix; the second step is based on the current projection operator. Calculate the aesthetic primitive encoding values for all generated trajectories and construct the aesthetic primitive encoding matrix. The third step is to calculate the loss function. The values are then used to update the orthogonal projection operator using the gradient descent algorithm. With unitary evolution operator The numerical value is determined while ensuring orthogonality and unitary constraints; the fourth step is to calculate the separability index of the primitive code. The output is then fed back to the RMCCBE algorithm in S2 to update the loss function value of the RMCCBE algorithm; the fifth step receives the latest equilibrium weight coefficients from the subsequent weight fitting output. With the optimal projection operator The sixth step is to update the calculated value of the loss function; the convergence condition is the loss function. If the change is less than 1e-6, or the maximum number of iterations is reached, and convergence is not achieved, return to the second step to continue iterating; if convergence is achieved, the training process ends, and the final aesthetic primitive encoding matrix is output. Standardized aesthetic feature dynamic primitive library, orthogonal projection operator Primitive code separability index .
[0054] The trained DMTE-OPC algorithm has three core contributions. First, it accurately captures the temporal evolution characteristics of the generated trajectory through unitary evolution of the density matrix, solving the core problem of losing temporal information in traditional methods. The extracted aesthetic features fully reflect the aesthetic evolution law of the generation process. Second, it constructs a mutually exclusive and complete dynamic primitive library of aesthetic features through orthogonal projection encoding, eliminating feature redundancy and ensuring the interpretability of subsequent modeling. Third, through a bidirectional feedback mechanism, it forms a synergistic optimization with the selection of preceding causal variables and subsequent weight fitting, allowing feature encoding to simultaneously take into account causal constraints, consistency of multi-subject perception, and distinguishability, solving the problem of disconnect between traditional encoding methods and preceding and subsequent processes.
[0055] In some embodiments, the causality constraint regularization coefficients in the DMTE-OPC algorithm are... Its preferred value range is 0.4 to 0.8; further, in the embodiment for professional content generation scenarios with high aesthetic accuracy requirements, The optimal value is 0.75 to strengthen causal boundary constraints, ensuring that the encoding process retains only the features corresponding to the core causal variables and eliminates irrelevant noise interference to the greatest extent. In the implementation example for general content generation scenarios where generalization is a priority, The optimal value is 0.45, which retains a moderate amount of feature redundancy while ensuring the core causal constraints, thereby improving the model's adaptability to different aesthetic styles and different generation tasks.
[0056] S4: Multi-Subject Perceptual Fusion and Construction of Aesthetic Hierarchical Structure Model Based on the average treatment effect boundary estimation results and boundary constraint matrix Aesthetic primitive encoding matrix This paper introduces a standardized dynamic primitive library of aesthetic features and a hierarchical Bayesian equilibrium fitting algorithm with random coefficients (RCH-BEF algorithm) to complete the heterogeneous fitting of multi-subject aesthetic perception data and solve the consensus weights, thus constructing a three-level progressive aesthetic feature hierarchical structure model. Traditional methods often use simple average weighting or subjective assignment, which cannot eliminate the differences in aesthetic perception among different subjects, and the resulting weights lack consensus and causal support. This step captures the perceptual heterogeneity of different subjects through a hierarchical Bayesian model, obtains the weights of multi-subject consensus through Nash equilibrium fitting, and introduces the results of prior causal effects as a prior constraint to ensure that the weight allocation is consistent with the strength of the causal effect. Furthermore, through a two-way feedback mechanism, the accuracy of the two core algorithms in the preceding steps is optimized, completing the core construction of the entire aesthetic feature model. The hierarchical Bayesian equilibrium fitting and aesthetic feature hierarchical modeling process is as follows: Figure 4 As shown.
[0057] The core construction logic of the RCH-BEF algorithm is to fit the aesthetic perception data of multiple subjects through a hierarchical Bayesian model, introduce random coefficients to capture the perceptual heterogeneity of different subjects, obtain the aesthetic feature weights of multi-subject consensus through Nash equilibrium fitting, and construct a three-level progressive hierarchical structure model from the bottom quantifiable features to the top interpretable aesthetic concepts, thus solving the problem that traditional aesthetic models cannot balance quantitative accuracy and interpretability.
[0058] For the first The subject to the first The aesthetic perception score of the generated result corresponding to each generated trajectory ranges from [0,1]. The higher the value, the better the aesthetic evaluation of the subject. For the first Each subject corresponds to a random coefficient vector, where each element in the vector corresponds to the perceptual weight of a dynamic primitive of aesthetic features. The dimension is consistent with the number of dynamic primitives of aesthetic features. This is the comprehensive weight matrix of aesthetic primitives for multi-agent consensus output by the RCH-BEF algorithm. Each element in the matrix corresponds to the final weight of a dynamic primitive of aesthetic features. For the first The weights of each subject in the equilibrium fitting are such that the sum of all subject weights is 1. , This is the loss function of the RCH-BEF algorithm, used to optimize the fitting accuracy and balanced fitting effect of the hierarchical Bayesian model. is the causal prior regularization coefficient of the RCH-BEF algorithm, used to control the constraint strength of the causal effect boundary estimation result of S2 output on the prior distribution, and its value ranges from (0,1).
[0059] The RCH-BEF algorithm first completes multi-subject perceptual data collection, constructing four complementary perceptual subjects: the generative AI target model self-consistency subject, the human perceptual group subject, the machine aesthetic cognitive system subject, and the professional aesthetic practitioner subject. The scoring data collection rules for these four subjects are complete and clearly defined. The score for the generative AI target model self-consistency subject is calculated based on the model's own reconstruction error of the generated results. For text-based models, the normalized value of perplexity is used; for image-based models, the normalized value of FID is used; for audio-based models, the normalized value of Mel-spectrum reconstruction error is used; and for multimodal models, the weighted average of the reconstruction errors of each modality is used. The score for the human perceptual group subject comes from healthy adult subjects with basic aesthetic cognitive abilities who have passed ethical review, with a sample size of no less than 300 people, covering different ages, genders, educational backgrounds, and aesthetic preferences. For the preference group, a pairwise comparison experimental paradigm was used, and normalized scores were calculated using the Bradley-Terry model. For the machine aesthetic cognitive system, a pre-trained cross-modal aesthetic cognitive model was used to quantify the generated results across six dimensions: formal harmony, emotional delivery, creative novelty, narrative completeness, structural rationality, and expressiveness of artistic conception. A normalized comprehensive score was obtained through principal component analysis. For professional aesthetic practitioners, scores were derived from aesthetic professionals with over 5 years of experience. A 5-point Likert scale was used to rate the generated results, normalized to the [0,1] interval. All collected scoring data were strongly bound to the unique identifier of the generation trajectory and the aesthetic primitive encoding matrix. The rows correspond one-to-one.
[0060] After completing the multi-subject perception data collection, a hierarchical Bayesian fitting with random coefficients is performed. The core of this step is to fit the random coefficient vector of each subject using a hierarchical Bayesian model, while simultaneously introducing the average treatment effect boundary estimate as a prior distribution to constrain the range of random coefficient values and ensure that the weights are consistent with the strength of the causal effect. The hierarchical Bayesian model has a two-layer structure. The first layer is the likelihood layer, which describes the relationship between subject ratings and aesthetic primitive encodings. The calculation formula is as follows: ; in the formula It follows a normal distribution, with the first parameter being the mean and the second parameter being the variance. Encoding matrix for aesthetic primitives The line, corresponding to the first The full primitive encoding values of the generated trajectory. For the first The variance of the score fitting error for each subject.
[0061] The second layer is the prior layer, which describes the random coefficient vector. The prior distribution is used, and the average treatment effect boundary estimation result is introduced as a constraint to ensure that the weight allocation has a causal basis. The calculation formula is as follows: ; in the formula Let be the population mean vector of random coefficient vectors, where each element corresponds to the midpoint value of the average treatment effect of a core aesthetic causal variable, i.e. , Let be the population covariance matrix of the random coefficient vector, and be a diagonal matrix. Each element on the diagonal corresponds to the interval length of the average treatment effect of a core aesthetic causal variable, i.e. .
[0062] After completing the hierarchical Bayesian fitting, multi-agent equilibrium weight fitting is performed. The core of this step is to obtain the subject weights and comprehensive weight matrix of the multi-agent consensus through Nash equilibrium fitting based on the fitting results of the random coefficients of each subject. This addresses the problem of perceptual heterogeneity among different subjects, yields unbiased consensus weights, and avoids the bias of a single subject dominating the consensus. (Comprehensive weight matrix of multi-agent consensus) The calculation formula is as follows: Main weight The equilibrium fitting condition is that the weight of each subject is positively correlated with the fitting accuracy of that subject and negatively correlated with the scoring deviation of that subject from other subjects. The formula for calculating the equilibrium solution is as follows: ; in the formula For the first Mean squared fitting error of hierarchical Bayesian models for each subject The temperature coefficient, which is used to balance the fit, has a value range of (0,1) and is used to control the distribution concentration of the main weights.
[0063] Based on the comprehensive weight matrix A three-level progressive aesthetic feature hierarchical structure model is constructed, which realizes a complete mapping from quantifiable features to interpretable aesthetic concepts from the bottom to the top. The bottom layer is the primitive layer, and the input is all the primitive encoding values of the standardized aesthetic feature dynamic primitive library. Each primitive corresponds to an independent input node, and the node weight is determined by a comprehensive weight matrix. The corresponding elements are assigned values in the first layer; the middle layer is the dimension layer, which contains six mutually exclusive and complete aesthetic dimensions. Each dimension is mapped by the combination of corresponding primitives, without adding any additional features. The six aesthetic dimensions are formal aesthetic dimension, emotional aesthetic dimension, creative aesthetic dimension, narrative aesthetic dimension, structural aesthetic dimension, and artistic conception aesthetic dimension. The mapping relationship for each dimension is that the dimension score is equal to the sum of the product of the encoded value of the corresponding primitive and the weight; the upper layer is the style layer, which contains mainstream aesthetic style classifications. Each style is mapped by the weight combination of the six aesthetic dimensions in the dimension layer. They are minimalism, romanticism, realism, surrealism, classicism, modernism, avant-garde, naturalism, cyberpunk, and traditional Chinese style aesthetics. The mapping relationship for each style is that the style matching degree is equal to the sum of the product of the scores of the six aesthetic dimensions and the corresponding style weights.
[0064] The RCH-BEF algorithm employs Markov chain Monte Carlo sampling during training, with the core optimization objective being to minimize the loss function. This simultaneously satisfies the prior constraints and equilibrium fitting conditions of the hierarchical Bayesian model, and incorporates bidirectional feedback signals from the previous two steps, achieving synergistic optimization with the other two core algorithms. Loss function The calculation formula is as follows: ; in the formula This is the vector of midpoint values of the average treatment effect. is the constraint coefficient for the causal consistency index, and its value ranges from (0,1]. This is a multi-agent causal consistency index, calculated as the inverse of the Pearson correlation coefficient between the multi-agent consensus weight and the average treatment effect. It is used to back-output the results to the RMCCBE algorithm in S2, optimizing the accuracy of causal boundary estimation and enabling bidirectional interaction with the screening of preceding causal variables. The smaller the value, the higher the consistency between the weight and the causal effect, and the better the fit.
[0065] During training, the optimal projection operator for each subject is calculated synchronously. This is used for backfeedback to the DMTE-OPC algorithm in S3, enabling bidirectional interaction with the preceding feature encoding. Optimal projection operator. The fitting formula is as follows: The constraints are: , That is, orthogonal projection constraint.
[0066] The iterative steps of the training process strictly follow closed-loop optimization logic. The first step is to initialize the causal prior regularization coefficients. Constraint coefficient Equilibrium fitting temperature coefficient The initial iteration count is set to 0, the maximum sampling count is set to 10,000, and the burn-in period is set to 2,000. The second step uses the Gibbs sampling method to process the random coefficient vector of the hierarchical Bayesian model. ,variance population mean vector Overall covariance matrix The third step involves sampling and updating parameter values; based on the random coefficient vector obtained from the sampling... Calculate the mean squared fitting error for each subject. The main body weights are calculated using the equilibrium fitting formula. The comprehensive weight matrix is obtained. The fourth step is to calculate the loss function. The value is used to calculate the multi-agent causal consistency index. The output is then fed back to the RMCCBE algorithm in S2 to update the loss function value of the RMCCBE algorithm; the fifth step is to fit the optimal projection operator for each subject. Combine it with the main weight The reverse output is fed back to the DMTE-OPC algorithm in S3 to update the loss function value of the DMTE-OPC algorithm; the sixth step determines whether the sampling has converged. The convergence condition is that the Gelman-Rubin diagnostic index is less than 1.1, or the maximum number of samplings has been reached. If it has not converged, return to the second step to continue sampling. If it has converged, the training process ends, and the final comprehensive weight matrix is output. Aesthetic feature hierarchical structure model, multi-subject causal consistency index Main weight Optimal projection operator .
[0067] The trained RCH-BEF algorithm has three core contributions. First, it effectively captures the heterogeneity of aesthetic perception among different subjects through a hierarchical Bayesian model and obtains weights based on multi-subject consensus through Nash equilibrium fitting, solving the problems of serious subjective bias and lack of consensus in traditional methods. Second, it introduces causal effects as prior constraints, ensuring that the weight allocation has a rigorous causal basis rather than spurious correlation fitting results, thus improving the model's generalization ability. Third, it constructs a three-level progressive aesthetic feature hierarchical structure model, achieving a balance between quantification accuracy and interpretability. At the same time, through a two-way feedback mechanism, it forms a complete closed-loop optimization framework with the two preceding core algorithms, allowing the three algorithms to work together in pairs, achieving a 1+1+1>3 effect and solving the core problem of disconnect in the traditional modeling process.
[0068] In some embodiments, the equilibrium fitting temperature coefficient in the RCH-BEF algorithm is considered. Its preferred value range is 0.1 to 0.6; further, in embodiments targeting mass consumer-level production scenarios, The optimal value is 0.2 to increase the concentration of the main subject weight distribution, prioritize amplifying the weight proportion of the human perception group, and match the mainstream aesthetic preferences of the general public; in the embodiment targeting professional-level content creation scenarios, The optimal value is 0.5 to balance the weight distribution of each perceptual subject, taking into account the perceptual consensus of professional aesthetic practitioners, machine aesthetic cognitive systems, and the human population, and adapting to the diverse aesthetic needs of professional creation; in an embodiment oriented towards model self-optimization scenarios, The optimal value is 0.15, in order to improve the weight priority of the model's self-consistency subject and adapt to the iterative optimization needs of the generative AI model's own aesthetic capabilities.
[0069] S5: Model Validation and Application Boundary Correction Based on the closed-loop fusion framework constructed by the three core algorithms in the preceding steps and the preliminary aesthetic feature hierarchical structure model, this step completes the full-dimensional validity verification and application boundary correction of the model, ensuring the causal validity, generalization ability, discrimination ability and stability of the model, solving the problems of traditional aesthetic models having a single verification dimension, no clear application boundary and uncontrollable actual deployment effect, and providing a rigorous validity guarantee for the actual deployment and application of the model.
[0070] First, a full multi-dimensional validity validation is performed. Each validation dimension has a clearly defined pass / fail threshold and a failure feedback path. These feedback paths directly connect to their preceding steps, forming a complete optimization loop. The first validation dimension is causal validity validation. Its core objective is to verify the true causal effect of the core aesthetic causal variables included in the model on aesthetic scores, avoiding spurious correlations and noise features in the model. The validation method involves repeating the counterfactual intervention experiment in S2, performing targeted interventions for each core aesthetic causal variable, recording the changes in the corresponding aesthetic dimension scores after the intervention, and calculating the causal effect size. The pass / fail threshold is an absolute value of the causal effect size greater than 0.3 and a p-value less than 0.05, indicating that the causal effect of the variable is statistically significant. If the validation fails, the process returns to S2 to re-execute the RMCCBE algorithm training process, optimizing the causal variable selection results until the validation is passed.
[0071] The second validation dimension is generalization validity validation. Its core objective is to verify the model's generalization ability on tasks and styles not previously trained on, preventing overfitting of the training data and ensuring that the model learns general aesthetic principles rather than task-specific features. The validation method involves extracting 20% of the generation tasks from the entire set of generation tasks in S1 that were not modeled, along with three unmodeled aesthetic styles. The model is then used to score the generation results, and the Pearson correlation coefficient between the model score and the multi-agent consensus score is calculated. A passing threshold is a correlation coefficient greater than or equal to 0.85, indicating that the model still possesses stable evaluation capabilities on unseen tasks and styles. If the validation fails, the process returns to S3 to re-execute the DMTE-OPC algorithm training process, optimizing the primitive encoding results until the validation is passed.
[0072] The third validation dimension is discriminant validity validation. The core objective is to verify the model's ability to distinguish between high and low aesthetic quality generated results, ensuring that the model can effectively identify differences in aesthetic quality among the generated results. The validation method involves constructing 1000 pairs of samples. Each pair contains two generated results that satisfy semantic consistency constraints but have a multi-subject consensus aesthetic score difference greater than or equal to 2 points. The model scores each pair of samples, and the ROC-AUC value for correctly distinguishing high aesthetic score samples is calculated. The passing threshold is an ROC-AUC value greater than or equal to 0.9, indicating that the model possesses extremely strong aesthetic quality discrimination ability. If the validation fails, the process returns to S4 to re-execute the RCH-BEF algorithm training process, optimizing the comprehensive weight matrix until the validation is passed.
[0073] The fourth verification dimension is stability verification. Its core objective is to verify the stability of the model's output, preventing the model from giving significantly different evaluation results for multiple generation attempts of the same input, thus ensuring the model's reliability. The verification method involves performing 10 independent generation processes for the same generation task and the same initial parameters. The model scores the aesthetic features of the 10 generated trajectories and calculates the coefficient of variation (COP). The passing threshold is a COP of less than or equal to 5%, indicating that the model output possesses extremely high stability. If the verification fails, the process returns to S1 to re-optimize the Riemannian manifold reference space, improving the stability of the trajectory encoding until the verification passes.
[0074] After completing full-dimensional validity verification, the model boundary specifications are constructed. For models that pass verification, the applicable task scope, modal boundaries, and style boundaries of the model are clarified, and the model application boundary specifications are constructed to define the applicable and inapplicable scenarios of the model, avoiding model misuse. Simultaneously, the docking standards between the model and the generative AI target model, the parameter specifications for real-time intervention, and the aesthetic intervention interface specifications for the generative process are defined, providing standardized interface support for the actual deployment and application of the model. The final output includes a verified generative AI aesthetic feature quantification model, model application boundary specifications, and generative process aesthetic intervention interface specifications.
[0075] S6: Incremental Model Update and Construction of Lifelong Learning Mechanism Based on the fusion framework of a validated aesthetic feature quantification model and three preceding core algorithms, an incremental update and lifelong learning mechanism for the model is constructed. This addresses the pain point that traditional aesthetic models cannot adapt to new generation tasks, new aesthetic styles, and changes in user preferences, ensuring the performance stability and adaptability of the model during long-term deployment and preventing performance degradation over time.
[0076] First, an incremental learning sample pool was constructed, which is dynamically updated. The sample data comes from user interaction logs after the generative AI target model is deployed. The data has undergone user authorization and privacy anonymization, fully complying with relevant laws and regulations on personal information protection. The samples in the sample pool include users' full generation trajectory data and users' implicit aesthetic feedback tags, namely user adoption, user likes, user positive modifications, user abandonment, and user negative modifications. No manual annotation is required; the data is obtained entirely based on users' natural interaction behavior, significantly reducing the cost of incremental updates.
[0077] Secondly, an implicit aesthetic anchoring mechanism was constructed, using the constructed Riemannian manifold reference space as the permanent anchoring space and the primitive encoding mapping function as the fixed encoding rule. This ensures that the encoding space of the original primitives does not shift during incremental learning, avoiding catastrophic forgetting and guaranteeing that the model's original capabilities are not lost with incremental updates. Simultaneously, based on the user's implicit feedback labels, the aesthetic scores of incremental samples are automatically labeled: positive feedback samples are labeled with high aesthetic scores, and negative feedback samples are labeled with low aesthetic scores, providing standardized labeled data for incremental learning.
[0078] Then, incremental updates and lifelong learning are executed. An elastic weight consolidation algorithm is used to incrementally update the validated generative AI aesthetic feature quantification model. The core of the algorithm is to protect the weights of the original core aesthetic features from being overwritten, while simultaneously learning and updating the primitives and weights corresponding to new aesthetic styles and new task scenarios. This retains existing capabilities while learning new content, perfectly solving the catastrophic forgetting problem. The incremental update is triggered when the cumulative sample size in the incremental sample pool reaches 100,000 samples, or when the correlation coefficient between the model's aesthetic score and implicit user feedback in actual applications is less than 0.75. This ensures that the model is not unstable due to frequent updates, nor does it suffer from decreased adaptability due to prolonged periods without updates. During the incremental update process, bidirectional interactive optimization of the three core algorithms (S2, S3, and S4) is executed simultaneously, ensuring that the updated model meets all constraints and does not disrupt the original closed-loop optimization framework.
[0079] Finally, incremental model validation and specification updates are completed. After each incremental update, the full-dimensional validity validation in S5 is repeated. If the validation is successful, the model parameters, model application boundary specifications, and generative process aesthetic intervention interface specifications are updated to ensure that the incrementally updated model still has sufficient validity and stability. If the validation fails, the model is rolled back to the previous version, and incremental sample selection and learning are performed again to avoid performance degradation caused by ineffective updates. The final output is a generative AI aesthetic feature final state model with lifelong learning capabilities, incremental update process specifications, and implicit feedback anchoring rule set, completing the entire process of generative AI aesthetic feature modeling.
[0080] In some embodiments, the weighting coefficients for the multi-agent causal consistency index in the RMCCBE algorithm are... Its preferred value is 0.2 to 0.4, and its optimal value is 0.3; this is the constraint coefficient for the balance weight in the DMTE-OPC algorithm. Its preferred value is 0.3 to 0.6, and its optimal value is 0.4; regarding the causal prior regularization coefficient in the RCH-BEF algorithm The preferred value is 0.5 to 0.8, and the optimal value is 0.65. These preferred values can be matched with the aforementioned core coefficients to ensure the bidirectional interaction and closed-loop optimization effect of the three core algorithms, and further reduce the implementation threshold for technical personnel in this field.
[0081] In some embodiments, the orthogonal projection operator corresponding to the temporal smoothness primitive Constructed within the smoothness subspace of the temporal evolution of the generated trajectory, its projection space constraint is the temporal correlation dimension of the density matrix corresponding to adjacent generation step sizes within a single generated trajectory, satisfying the idempotency and Hermitian constraints of the orthogonal projection operator, i.e.: , Specifically, The construction uses the trace distance between adjacent generated step size density matrices as the core metric, and the kernel function of the projection operator is: In the formula, For the first The generated trajectory number The hidden state density matrix corresponding to each generation step size. For the trace operation of a matrix, This involves the absolute value operation of the matrix elements. The operator projection aims to extract the smoothness features of the temporal evolution of the hidden states during the generation process, filtering out non-smooth noise interference caused by temporal jumps; simultaneously, The projection subspace of the given element is linearly independent of the subspaces of the projection operators corresponding to the other five primitives, ensuring pairwise orthogonality constraints, i.e., for any ,satisfy .
[0082] In some embodiments, the orthogonal projection operator corresponding to the attention convergence primitive Constructed on the temporal convergence subspace of the attention weights in the generation process, its projection space constraint is the temporal evolution dimension of the multi-head attention weight matrix corresponding to each generation step size within a single generation trajectory, satisfying the idempotency and Hermitian constraints of the orthogonal projection operator, i.e.: , Specifically, The construction of the projection operator is based on the Gini coefficient and temporal change rate of the attention weight distribution at each generation step. The kernel function of the projection operator is: In the formula, For the first The generated trajectory number The multi-head attention weight matrix corresponding to each generation step size. This is the function for calculating the Gini coefficient. The operator projection objective is to extract the convergence characteristics of the attention weights from dispersion to focus during the generation process, filtering out noise interference caused by disordered fluctuations in attention; simultaneously, The projection subspace of the given element is linearly independent of the subspaces of the projection operators corresponding to the other five primitives, ensuring pairwise orthogonality constraints, i.e., for any ,satisfy .
[0083] In some embodiments, the orthogonal projection operator corresponding to the manifold curvature uniformity primitive The curvature stability subspace is constructed on the reference space of the Riemannian manifold. Its projection space constraint is the geodesic curvature evolution dimension of each generation step hidden state within a single generation trajectory on the Riemannian manifold, satisfying the idempotency and Hermitian constraints of the orthogonal projection operator, i.e.: , Specifically, The construction of hidden states with generation step size in the Riemannian manifold The deviation between the local curvature and the global mean curvature of the manifold is the core metric, and the kernel function of the projection operator is: In the formula, For the first The generated trajectory number A hidden state with a generation step size in a Riemannian manifold Local curvature on, Let be the global mean curvature of the Riemannian manifold reference space. This is an absolute value operation. The operator projection objective is to extract the curvature consistency features of the latent states evolving in the manifold space during the generation process, filtering out semantic offset noise caused by anomalous jumps in manifold curvature; simultaneously... The projection subspace of the given element is linearly independent of the subspaces of the projection operators corresponding to the other five primitives, ensuring pairwise orthogonality constraints, i.e., for any ,satisfy .
[0084] In some embodiments, the orthogonal projection operator corresponding to the probability distribution entropy-variant primitive is... The entropy evolution subspace of the sampling probability distribution in the generation process is constructed. Its projection space constraint is the temporal entropy change dimension of the token sampling probability distribution and the noise prediction probability distribution corresponding to each generation step size within a single generation trajectory, satisfying the idempotency and Hermitian constraints of the orthogonal projection operator, i.e.: , Specifically, The construction of the projection operator uses the temporal change of Shannon entropy for each generation step probability distribution as the core metric, and the kernel function of the projection operator is: In the formula, For the first The generated trajectory number The sampling probability distribution tensor corresponding to each generation step size. The Shannon entropy calculation function has the following standard formula: ;in It is a discrete probability distribution vector. The dimension of the probability distribution is defined as follows. The operator projection aims to extract the entropy evolution characteristics of the probability distribution during the generation process, filtering out generation stability noise caused by disordered entropy changes in the probability distribution; simultaneously, The projection subspace of the given element is linearly independent of the subspaces of the projection operators corresponding to the other five primitives, ensuring pairwise orthogonality constraints, i.e., for any ,satisfy .
[0085] In some embodiments, the orthogonal projection operator corresponding to the cross-layer feature alignment primitive The feature alignment subspace built between the layers of the generative model network is constrained by the inter-layer alignment dimension of the hidden state tensors of each encoder / decoder / UNet layer corresponding to each generation step size within a single generation trajectory. This satisfies the idempotency and Hermitian constraints of the orthogonal projection operator, i.e.: , Specifically, The construction of the projection operator uses the cosine similarity of the density matrices after mapping the hidden states of adjacent network layers as the core metric. The kernel function of the projection operator is: In the formula, For the first The generated trajectory number The generation step size is... The hidden state density matrix corresponding to the layer network, The cosine similarity calculation function has the following standard formula: ; The operator projection aims to extract feature alignment consistency features between different network layers during the generation process, filtering out semantic breakdown noise caused by feature misalignment between layers; simultaneously... The projection subspace of the given element is linearly independent of the subspaces of the projection operators corresponding to the other five primitives, ensuring pairwise orthogonality constraints, i.e., for any ,satisfy .
[0086] In some embodiments, the orthogonal projection operator corresponding to the cross-modal stability primitive The cross-modal feature stability subspace is constructed in a multimodal generation scenario. Its projection space constraint is the cross-modal alignment dimension of the hidden state tensors of different modalities corresponding to each generation step size within a single generation trajectory, satisfying the idempotency and Hermitian constraints of the orthogonal projection operator, i.e.: , Specifically, The construction of the projector uses the temporal variation of the mutual information of the hidden state density matrices of different modes as the core metric, and the kernel function of the projector is: In the formula, , The first The generated trajectory number The generation step size is... , No. The hidden state density matrix corresponding to each mode. The standard formula for calculating mutual information is: ;in , For edge entropy, The joint entropy is used. The operator projection objective is to extract the temporal stability features of feature mappings between different modes during multimodal generation and filter out modal misalignment noise caused by cross-modal feature shifts; for single-modal generation scenarios, The projection output uses a fixed reference value, which does not affect the consistency of the primitive encoding; at the same time, The projection subspace of the given element is linearly independent of the subspaces of the projection operators corresponding to the other five primitives, ensuring pairwise orthogonality constraints, i.e., for any ,satisfy .
[0087] The embodiments described above are for illustrative purposes only and are not intended to limit the invention. Therefore, any changes in numerical values or substitutions of equivalent elements should still fall within the scope of this invention.
[0088] The above detailed description will enable those skilled in the art to understand that the present invention can indeed achieve the aforementioned objectives and has complied with the provisions of the Patent Law.
[0089] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention. The above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the invention should be included within the scope of protection of the invention.
[0090] It should be noted that the above description of the process is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to the process under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.
[0091] The basic concepts have been described above. Obviously, for those skilled in the art who have read this application, the above disclosure is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore, such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this application.
[0092] Furthermore, this application uses specific terms to describe its embodiments. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different positions in this specification do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application can be appropriately combined.
Claims
1. A generative artificial intelligence method for modeling aesthetic features, characterized in that, Includes the following steps: S1: Collect full-time generation trajectory data and final generation results corresponding to the generation task of the generative artificial intelligence target model, and construct the generation trajectory dataset, the generation result benchmark set and the Riemannian manifold benchmark space; S2: Based on the generated trajectory dataset, the generated result benchmark set, and the Riemann manifold benchmark space, the counterfactual causal boundary estimation algorithm under Riemann manifold constraints is used to complete the causal effect boundary estimation of candidate variables. Based on the boundary estimation results, the aesthetic core causal variable set is obtained and the boundary constraint matrix is generated. S3: Based on the generated trajectory dataset, Riemannian manifold reference space, aesthetic core causal variable set and boundary constraint matrix, the orthogonal projection coding algorithm of density matrix temporal evolution is adopted to complete the temporal evolution modeling and feature encoding of the generated trajectory, and construct a standardized aesthetic feature dynamic primitive library and aesthetic primitive coding matrix; S4: Aesthetic scoring is performed on the generated benchmark set to obtain a multi-subject aesthetic perception dataset. Then, based on the standardized aesthetic feature dynamic primitive library, aesthetic primitive encoding matrix and multi-subject aesthetic perception dataset, a random coefficient hierarchical Bayesian equilibrium fitting algorithm is used to fit the comprehensive weight of aesthetic primitives of multi-subject consensus and construct an aesthetic feature hierarchical structure model. S5: Perform full-dimensional validity verification on the hierarchical structure model of aesthetic features, and correct the applicable boundaries of the model based on the verification results to obtain a qualified generative artificial intelligence aesthetic feature quantification model.
2. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, It also includes S6: a fusion framework based on the counterfactual causal boundary estimation algorithm under Riemannian manifold constraints, the orthogonal projection coding algorithm of density matrix temporal evolution, and the stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm, to construct an incremental update and lifelong learning mechanism for the generative artificial intelligence aesthetic feature quantification model.
3. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, The specific generation tasks of the generative AI target models adapted in S1 are as follows: for text generation models, they cover all types of text generation tasks; for image generation models, they cover all types of image generation tasks; for audio generation models, they cover all types of audio generation tasks; and for multimodal generation models, they cover all corresponding unimodal and crossmodal generation tasks.
4. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, The counterfactual causal boundary estimation algorithm under Riemannian manifold constraints in S2 synthesizes counterfactual samples through exponential mapping on the Riemannian manifold, ensuring semantic consistency between the counterfactual samples and the original samples. Based on the counterfactual samples, it calculates the upper and lower boundaries of the average treatment effect of each candidate variable, selects variables whose boundary intervals do not contain 0 to form the aesthetic core causal variable set, and removes aesthetically irrelevant confounding variables whose boundary intervals contain 0.
5. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, The orthogonal projection coding algorithm of density matrix temporal evolution in S3 models the temporal evolution process of generated trajectory as the unitary evolution process of density matrix. The boundary constraint matrix is used as the projection space constraint to ensure that the coding process retains only the features corresponding to the core aesthetic causal variables. Mutually exclusive and non-redundant aesthetic feature dynamic primitives are extracted through pairwise orthogonal projection operators to complete feature coding.
6. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, The S3 standardized aesthetic feature dynamic primitive library includes temporal smoothness primitives, attention convergence primitives, manifold curvature consistency primitives, probability distribution entropy change primitives, cross-layer feature alignment primitives, and cross-modal stability primitives. The projection operators corresponding to all primitives are pairwise orthogonal.
7. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, The stochastic coefficient hierarchical Bayesian equilibrium fitting algorithm in S4 uses the causal effect boundary estimation results as the prior distribution constraint of the hierarchical Bayesian model. It captures the aesthetic perception heterogeneity of different rating subjects through stochastic coefficient fitting, and then obtains the comprehensive weight of aesthetic primitives of multi-subject consensus through Nash equilibrium fitting, ensuring that the weight allocation matches the causal effect strength of the variables.
8. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, The S4 multi-subject aesthetic perception dataset is obtained by standardizing the aesthetic scores of each generated result in the benchmark set of generated results through four complementary perceptual subjects. The four perceptual subjects are the self-consistency subject of the generative artificial intelligence target model, the human perceptual group subject, the machine aesthetic cognition system subject, and the professional aesthetic practitioner subject. All scores are strongly bound to the unique identifier of the corresponding generation trajectory.
9. The method for modeling aesthetic features in generative artificial intelligence according to claim 1, characterized in that, The full-dimensional validity verification in S5 includes four dimensions: causal validity verification, generalization validity verification, discriminant validity verification, and stability verification. Each dimension has a corresponding pass threshold. If any dimension fails the verification, the corresponding preceding steps are returned to complete the model optimization until all dimensions pass the verification.
10. The method for modeling aesthetic features in generative artificial intelligence according to claim 2, characterized in that, The incremental update and lifelong learning mechanism in S6 is implemented by constructing a dynamically updated incremental learning sample pool. The data in the sample pool comes from user-authorized and privacy-de-sensitized model interaction log data. Implicit aesthetic feedback annotation of samples is completed based on user interaction behavior. The elastic weight consolidation algorithm is used to complete the incremental update of the model, adapting to the newly added generation tasks and aesthetic styles while retaining the original core capabilities.