Attractive evaluation processing system based on multi-source data fusion
By integrating multi-source data and using modular processing, the problems of single data source and subjective evaluation in existing aesthetic education evaluation systems have been solved, achieving full automation and objectivity in aesthetic education evaluation and improving the accuracy and applicability of evaluation results.
Patent Information
- Application Number
- CN202511383462.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-01-20
AI Technical Summary
Existing aesthetic education evaluation systems rely on a single data source or subjective manual assessment, resulting in incomplete evaluation dimensions, insufficient objectivity, and a lack of dynamic adaptability, making them unable to meet the needs of diverse aesthetic education scenarios.
An aesthetic education evaluation and processing system based on multi-source data fusion is adopted, including data acquisition, preprocessing, fusion, feature extraction and evaluation engine modules. It integrates multimodal data such as digital images, videos, audio, and text, and generates objective evaluation results through adaptive weight allocation and machine learning models.
It has achieved full automation and objectivity in aesthetic education evaluation, improved the accuracy and reliability of evaluation results, provided personalized feedback, adapted to different forms of aesthetic education and data environments, and reduced labor costs.
Smart Images

Figure CN121365899A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of data processing, in particular to an aesthetic education evaluation processing system based on multi-source data fusion. BACKGROUND
[0002] The existing aesthetic education evaluation system relies on a single data source or subjective manual evaluation, which has the problems of incomplete evaluation dimension, insufficient objectivity and poor scalability. The traditional method is difficult to effectively integrate multi-modal data such as images, audio, and text, resulting in one-sided evaluation results and lack of depth analysis, which cannot meet the needs of diversified aesthetic education scenarios. In addition, the existing system usually lacks dynamic adaptive ability and cannot flexibly adjust the evaluation strategy according to the data characteristics, limiting its wide application in actual educational environments.
[0003] Therefore, a better solution is needed. SUMMARY
[0004] Therefore, the aesthetic education evaluation processing system based on multi-source data fusion is provided to solve the technical defects in the prior art.
[0005] According to a first aspect of the embodiments of the present specification, an aesthetic education evaluation processing system based on multi-source data fusion is provided, which includes a data acquisition module, a data preprocessing module, a data fusion module, a feature extraction module, an evaluation engine module and a result output module. The data acquisition module acquires aesthetic education related data from multiple heterogeneous data sources to determine the original data. The data preprocessing module cleans and standardizes the original data to determine the multi-source data. The data fusion module integrates the preprocessed multi-source data into a unified data representation to determine the fusion data. The feature extraction module extracts key features from the fusion data. The evaluation engine module generates an aesthetic education evaluation result based on the key features, and the result output module outputs the aesthetic education evaluation result.
[0006] In one possible implementation, the original data includes digital image data, video data, audio data, text data and auxiliary data.
[0007] In one possible implementation, the data preprocessing module performs denoising, brightness adjustment, size normalization and format conversion on the image data. Frame extraction, resolution unification and key frame selection are performed on the video data. Noise reduction, sampling rate standardization and silent segment removal are performed on the audio data. Text data is segmented, stop words are removed, spelling is corrected, and encoding is unified.
[0008] In a possible implementation, the data fusion module adopts a strategy of combining feature-level fusion and decision-level fusion, and adjusts the fusion proportion through an adaptive weight allocation method, where the fusion proportion is based on the reliability and correlation of the data sources.
[0009] In a possible implementation, the feature extraction module extracts visual features, auditory features, text features, and comprehensive features. Among them, the visual features include color distribution, texture pattern, and composition balance; the auditory features include pitch contour, rhythm stability, and timbre diversity; the text features include sentiment tendency, creative vocabulary density, and syntax complexity; and the comprehensive features include multi-modal interaction mode.
[0010] In a possible implementation, the evaluation engine module applies a pre-trained evaluation model based on a machine learning algorithm, and outputs a score or classification result, and further includes adaptive learning to update the evaluation model.
[0011] In a possible implementation, when calculating the aesthetic evaluation score, the evaluation engine module calculates a standardized feature value for each feature:
[0012] Among them, indicates the feature index, from 1 to N, N is the total number of features, indicates the ith feature value, obtained from the feature extraction module, indicates the mean value of the ith feature, obtained from the pre-trained model parameters, indicates the standard deviation of the ith feature, obtained from the pre-trained model parameters; Calculate the aesthetic evaluation score:
[0013] Among them, indicates the weight of the ith feature, obtained from the pre-trained model parameters, and N indicates the total number of features.
[0014] In a possible implementation, when calculating the comprehensive aesthetic evaluation score, the evaluation engine module uses the formula:
[0015] Among them, indicates the comprehensive aesthetic evaluation score, indicates the creativity score, obtained from the evaluation engine module, indicates the technical skill score, obtained from the evaluation engine module, indicates the aesthetic perception score, obtained from the evaluation engine module, indicates the creativity weight coefficient, obtained from the pre-trained model parameters, denotes a technical skill weight coefficient, obtained from pre-training model parameters, denotes an aesthetic perception weight coefficient, obtained from pre-training model parameters.
[0016] In a possible implementation, the data preprocessing module further includes a data augmentation function for increasing image sample diversity through rotation or cropping.
[0017] In a possible implementation, the evaluation engine module further includes an adaptive learning function for updating model parameters based on new data.
[0018] The embodiment of the present specification provides an aesthetic education evaluation processing system based on multi-source data fusion. The system includes a data acquisition module, a data preprocessing module, a data fusion module, a feature extraction module, an evaluation engine module, and a result output module. The data acquisition module acquires aesthetic education related data from multiple heterogeneous data sources to determine the original data. The data preprocessing module cleans and standardizes the original data to determine the multi-source data. The data fusion module integrates the preprocessed multi-source data into a unified data representation to determine the fusion data. The feature extraction module extracts key features from the fusion data. The evaluation engine module generates an aesthetic education evaluation result based on the key features, and the result output module outputs the aesthetic education evaluation result. Through multi-source data fusion and modular processing, the system realizes comprehensive automation and objectivity of aesthetic education evaluation, significantly improves the accuracy and reliability of the evaluation result. The system can comprehensively utilize multi-modal data, deeply mine aesthetic education features, and provide personalized feedback, effectively supporting education decision-making. At the same time, the adaptive design and expandable architecture of the system enable it to flexibly cope with different aesthetic education forms and data environments, enhancing the practicality and scope of application, and providing efficient technical support for aesthetic education. BRIEF DESCRIPTION OF DRAWINGS Figure 1 is a system schematic diagram of an aesthetic education evaluation processing system based on multi-source data fusion provided by an embodiment of the present specification. DETAILED DESCRIPTION
[0019] In the following description, many specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced in many different ways beyond the specific embodiments described herein, and it is understood that one of ordinary skill in the art can make and use the present specification without departing from the scope of the present specification.
[0020] The terminology used in this disclosure of one or more embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0021] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal sequence, but are used only to distinguish one piece of information from another. For example, without departing from the scope of one or more embodiments of the disclosure, first can be termed second; likewise, second can be termed first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining."
[0022] In the present specification, a beauty education evaluation processing system based on multi-source data fusion is provided, which is described in detail one by one in the following embodiments.
[0023] Referring to Figure 1 , Figure 1 A system schematic diagram of a beauty education evaluation processing system based on multi-source data fusion according to an embodiment of the present disclosure is shown, which specifically includes a data acquisition module, a data preprocessing module, a data fusion module, a feature extraction module, an evaluation engine module and a result output module; the data acquisition module acquires beauty education related data from multiple heterogeneous data sources to determine the original data; the data preprocessing module performs cleaning and standardization processing on the original data to determine the multi-source data; the data fusion module integrates the preprocessed multi-source data into a unified data representation to determine the fusion data; the feature extraction module extracts key features from the fusion data; the evaluation engine module generates a beauty education evaluation result based on the key features, and the result output module outputs the beauty education evaluation result.
[0024] The data collection module can refer to a hardware or software unit responsible for collecting aesthetic education-related data from multiple sources to ensure the diversity and integrity of the raw data. The multiple heterogeneous data sources can refer to data sources of different types and formats, including digital images, videos, audio, text, and auxiliary data, which can provide comprehensive information on aesthetic education activities. Aesthetic education-related data can refer to digitized content related to aesthetic education, such as student artwork, performance recordings, or written records, to support subsequent analysis and evaluation. Raw data can refer to initial data obtained directly from data sources without processing, serving as the starting point for the entire system processing. The data preprocessing module can refer to a component that cleans and standardizes raw data to eliminate noise and inconsistencies and improve data quality. Cleaning can refer to the process of removing errors, duplicates, or invalid parts of data to ensure data accuracy and usability. Standardization can refer to the operation of converting data to a uniform format or range, which can make data from different sources comparable. Multi-source data can refer to clean and consistent multi-category data obtained after preprocessing, which can be input into the subsequent fusion module. The data fusion module can refer to a unit that integrates multi-source data into a unified representation to generate a coordinated and consistent data view. The unified data representation can refer to a data structure or format compatible with different data types, which can facilitate the processing of the feature extraction module. The fusion data can refer to the integrated data output by the fusion module, which can be input into the feature extraction module. The feature extraction module can refer to a component that identifies and extracts key features from fusion data to capture the core information required for aesthetic education evaluation. Key features can refer to quantitative indicators that represent the essential attributes of aesthetic education activities, such as color distribution or rhythm patterns, which can be input into the evaluation engine. The evaluation engine module can refer to an algorithm or model unit that generates aesthetic education evaluation results based on key features to produce objective and comprehensive evaluation conclusions. The aesthetic education evaluation results can refer to the analysis output of aesthetic education performance, including scores, ratings, or feedback suggestions, to help educators and students understand progress. The result output module can refer to a component responsible for presenting and delivering aesthetic education evaluation results, which can display the results in a visual or readable form.
[0025] As a specific example: One embodiment of the system involves the evaluation of school art courses. The data collection module collects 50 digital images of student paintings from tablets, 10 video clips of the painting process from cameras, 10 audio clips of student creation descriptions from recording devices, and 20 student creation diaries from text input boxes as raw data. The data preprocessing module denoises and normalizes the images to 1024x768 pixels, extracts key frames from the videos and unifies them to 30 frames per second, denoises and normalizes the audio to 44.1kHz sampling rate, and performs word segmentation and spelling correction on the text, outputting multi-source data. The data fusion module aligns the video key frames, corresponding audio clips, and text diary entries according to timestamps, adopts a weighted fusion strategy to assign image data a weight of 0.5, video a weight of 0.3, audio a weight of 0.1, and text a weight of 0.1, and generates fusion data stored as a multi-dimensional matrix. The feature extraction module extracts visual features such as dominant color HSV values and composition symmetry ratios, auditory features such as voice frequency stability, and text features such as sentiment positive word frequency from the fusion data, and outputs a key feature vector. The evaluation engine module loads a pre-trained deep learning model, inputs the feature vector to calculate a creativity score of 85 points, a technical skill score of 78 points, and an aesthetic perception score of 82 points, and generates a comprehensive aesthetic education evaluation score of 83 points by weighting. The result output module generates a PDF report showing a score curve graph and text improvement suggestions, and pushes it to a teacher management platform for viewing through an API.
[0026] The present application realizes the comprehensive automation of aesthetic education evaluation through multi-module cooperation, significantly improves the objectivity and accuracy of the evaluation, and overcomes the limitations of traditional methods relying on subjective judgment. The system can efficiently integrate multi-source heterogeneous data, deeply mine aesthetic education features, provide multi-dimensional in-depth analysis feedback, and support personalized teaching decisions. The extensible architecture design allows flexible adaptation to different art forms and data types, enhancing the practicality and applicability of the system, providing efficient and reliable technical support for aesthetic education, and reducing labor costs and time expenditure through automated processing.
[0027] In one possible implementation, the raw data includes digital image data, video data, audio data, text data, and auxiliary data.
[0028] Among them, digital image data can refer to visual information stored in the form of a pixel matrix obtained through a digital device, used to record static aesthetic education works such as paintings, photographs, etc. Video data can refer to a dynamic visual sequence composed of consecutive image frames, capable of capturing time dimension information such as performances, creative processes, etc. Audio data can refer to an electronic signal that records sound waveforms to store music, recitations, or other auditory art forms. Text data can refer to language information existing in the form of character encoding, used to save written materials such as creation instructions, comments, or diaries. Auxiliary data can refer to supplementary information other than core artistic data, such as user operation logs or environmental parameters, capable of providing context background support analysis.
[0029] As a specific example: in a music teaching evaluation, the system collects digital image data such as 20 photos of students' instrument playing posture, video data such as 5 video recordings of the playing process, each 3 minutes, audio data such as 5 audio recordings of the playing catalog, text data such as 15 student practice notes, and auxiliary data such as practice duration records and classroom temperature and humidity sensor readings. The preprocessing module uniformly adjusts the images to 1920x1080 resolution and corrects the color, converts the video to MP4 format and extracts the audio track, standardizes the audio to -6dB volume and eliminates background noise, encodes and keyword labels the text, and cleans and time-aligns the auxiliary data. The fusion module fuses with audio data as the core weight 0.6, video data weight 0.25, image data weight 0.1, text and auxiliary data each weight 0.025, generates a spatiotemporally synchronized multi-modal data package. The feature extraction module extracts audio features including pitch deviation rate, rhythm stability index, video features including hand movement accuracy, posture coordination, image features including instrument holding angle, text features including practice method description quality, and auxiliary features including practice environment stability index. The evaluation engine calculates the skill accuracy score as 82 points, the artistic expressiveness score as 85 points, and the practice effectiveness score as 80 points, and finally generates a comprehensive score of 83 points. The result output module generates a visual radar chart to display the scores in each dimension, and generates a personalized improvement suggestion list and pushes it to the student terminal.
[0030] The present application significantly improves the dimension and depth of aesthetic education evaluation by comprehensively collecting multiple types of raw data, making the evaluation results more comprehensive and objective. The system can effectively capture the details of artistic performance that are easily overlooked, provide more accurate formative evaluation, and help learners improve specifically. The multi-modal data fusion mechanism enhances the system's ability to analyze complex artistic performances, adapting to the special needs of different artistic forms. The automated processing flow greatly improves the evaluation efficiency, reduces the teachers' work burden, and at the same time ensures the consistency of the evaluation standard. The flexible architecture design makes the system widely applicable to various aesthetic education scenes, promoting the standardization and digitization development of aesthetic education.
[0031] In a possible implementation, the data preprocessing module performs denoising, brightness adjustment, size normalization and format conversion on image data; performs frame extraction, resolution unification and key frame selection on video data; performs noise reduction, sampling rate standardization and silent segment removal on audio data; and performs word segmentation, stop word removal, spelling correction and encoding unification on text data.
[0032] Denoising can refer to a process of removing random interference pixels in digital images through algorithms to improve the purity and clarity of image signals. Brightness adjustment can refer to a processing operation that changes the overall brightness of an image, enabling images captured under different lighting conditions to be comparable. Size normalization can refer to an operation of adjusting images to a standard pixel size to eliminate size inconsistency caused by differences in acquisition devices. Format conversion can refer to a processing process that changes the encoding format of image files to unify data specifications for system processing. Frame extraction can refer to an operation of extracting specific image frames from a video stream, which can reduce the amount of data and retain key information. Resolution unification can refer to a process of adjusting video frames to a standard pixel size to ensure consistency in subsequent processing. Key frame selection can refer to a process of identifying and retaining representative frames in a video to capture the core change nodes of dynamic content. Noise reduction can refer to a processing technique that eliminates background noise in audio signals, which can improve the clarity of voice or music signals. Sampling rate standardization can refer to a process of unifying audio signals to a specific sampling frequency, which can enable audio from different sources to have the same time resolution. Silent segment removal can refer to an operation of deleting silent or low-volume segments in audio to improve the effective data density and processing efficiency. Word segmentation can refer to a processing process of dividing continuous text into independent lexical units as the basis for text analysis. Removing stop words can refer to an operation of filtering out high-frequency but low-information content words in text to improve the representativeness of text features. Spelling correction can refer to a process of detecting and correcting spelling errors in text to ensure the accuracy of text data. Encoding unification can refer to a process of converting text data to a standard character encoding format to eliminate compatibility problems caused by encoding differences.
[0033] As a specific example: in the dance art evaluation scene, the system receives the original dance video data (1080p, 30fps, MP4 format) and the accompanying audio data (44.1kHz, MP3 format). The preprocessing module first performs frame extraction on the video data, extracting 2 key frames per second, a total of 180 frames; the resolution is uniformly adjusted to 1280x720 pixels; 35 key frames containing typical dance movements are selected. The audio data is denoised, with a signal-to-noise ratio of 30dB; the sampling rate is standardized to 48kHz; the first and last silent segments and the middle pause segments are removed, reducing the total length from 3 minutes to 2 minutes and 45 seconds. At the same time, the accompanying text training notes are processed, and 325 word units are obtained after word segmentation; 28 stop words such as "of" and "of" are removed; 3 spelling errors are corrected; the encoding is uniformly converted to UTF-8 format. The processed data is output as a standardized data package for subsequent modules.
[0034] The present application significantly improves the quality and consistency of multi-source data through a systematic preprocessing process, providing a reliable foundation for subsequent analysis. The comprehensive application of various data processing techniques effectively eliminates various noise and bias introduced in the data collection process, ensuring the comparability of data from different sources. Standardization processing enables heterogeneous data to be processed and analyzed under a unified framework, greatly improving the compatibility and processing efficiency of the system. Special processing methods are adopted for different data types, which not only retain the key information of the original data, but also optimize the data structure and storage space. These preprocessing operations create good conditions for subsequent feature extraction and fusion analysis, ultimately improving the accuracy and reliability of the entire aesthetic education evaluation system.
[0035] In one possible implementation, the data fusion module adopts a combination of feature-level fusion and decision-level fusion strategies, and adjusts the fusion ratio through an adaptive weight allocation method, where the fusion ratio is based on the reliability and correlation of the data sources.
[0036] Among them, feature-level fusion can refer to the processing strategy of integrating feature vectors of different data sources in the feature extraction stage, which is used to realize multi-source information integration while maintaining the original characteristics of the data. Decision-level fusion can refer to the strategy of making a comprehensive judgment after each data source independently generates a preliminary evaluation result, which can improve the reliability of the final result through multi-angle decision-making. The adaptive weight allocation method can refer to an algorithm that dynamically adjusts the importance of each data source according to real-time data characteristics, which can automatically optimize the fusion effect according to the change of data quality. The fusion ratio can refer to the proportion of different data sources in the final fusion result, which is used to control the contribution degree of each type of data to the overall result. The reliability of the data source can refer to the degree of confidence that a specific data source provides accurate and consistent data, which can be an important basis for weight allocation. The relevance of the data source can refer to the correlation strength between the data source and the current evaluation task, which can ensure that the fusion process focuses on the most relevant information content.
[0037] As a specific example: in drama performance evaluation, the system processes video data (actor performance), audio data (dialogue expression) and text data (script content) simultaneously. The data fusion module first performs feature-level fusion, combining the body movement features extracted from the video, the speech emotion features extracted from the audio, and the semantic features extracted from the text into a multi-modal feature vector. Then, decision-level fusion is performed, with the video analysis module outputting a performance appeal score of 86 points, the audio analysis module outputting a dialogue expressiveness score of 82 points, and the text analysis module outputting a content fit degree score of 90 points. The adaptive weight allocation method assigns a weight of 0.5 to the video data based on the data quality evaluation result: high reliability (good clarity), a weight of 0.3 to the audio data: medium reliability (with environmental noise), and a weight of 0.2 to the text data: high reliability but low relevance (script is fixed content). Finally, the comprehensive performance score is calculated according to the weighted proportion to obtain a comprehensive performance score of 85.2 points, and the fusion evaluation result is generated.
[0038] The present application significantly improves the comprehensiveness and accuracy of aesthetic education evaluation through a dual fusion strategy, which can not only preserve the original data characteristics but also integrate multiple decision-making perspectives. The adaptive weight mechanism enables the system to intelligently respond to different data quality conditions, automatically highlighting the contribution of high-quality data while suppressing the influence of low-quality data. The proportion allocation based on reliability and relevance ensures that the fusion result is more objective and reasonable, effectively avoiding subjective bias. This fusion method enhances the system's comprehensive analysis capability for complex artistic performances, making the evaluation result more reflective of the true artistic level. The modular fusion architecture has good scalability, allowing new data sources and fusion algorithms to be easily integrated. The entire fusion process realizes the organic integration of multi-source data rather than simple superposition, producing a synergistic effect, and providing a more reliable technical foundation for aesthetic education evaluation.
[0039] In a possible implementation, the feature extraction module extracts visual features, auditory features, textual features, and comprehensive features; wherein the visual features include color distribution, texture pattern, and composition balance; the auditory features include pitch contour, rhythm stability, and timbre diversity; the textual features include sentiment tendency, creative vocabulary density, and grammatical complexity; and the comprehensive features include multi-modal interaction mode.
[0040] The color distribution can refer to the statistical characteristics of the proportion of different color values in the image, which is used to quantify the color application characteristics and style tendency of the work. The texture pattern can refer to the visual structure law of repeated appearance on the surface of the image, which can reflect the material performance and brushstroke technique characteristics of the work. The composition balance can refer to the balance degree of the distribution of visual elements in the picture, to evaluate the rationality of the spatial arrangement of the work. The pitch contour can refer to the trajectory curve of the fundamental frequency in the audio signal changing with time, which is used to analyze the pitch accuracy and melody line. The rhythm stability can refer to the uniformity of the interval of audio beats, which can measure the rhythm control ability of performance or singing. The timbre diversity can refer to the richness of the sound spectrum characteristics, which can evaluate the timbre change and expressiveness in performance. The sentiment tendency can refer to the emotional polarity direction expressed by the text content, which is used to analyze the emotional expression characteristics in creation. The creative vocabulary density can refer to the proportion of unconventional words or metaphorical expressions in the text, which can measure the innovation degree of language creation. The grammatical complexity can refer to the complexity and change degree of sentence structure, to evaluate the proficiency of language use. The multi-modal interaction mode can refer to the association and synchronization characteristics between different media data, which is used to analyze the coordination of cross-media artistic performance.
[0041] As a specific example: in evaluating a group of abstract painting works, the feature extraction module analyzes digital images to extract visual features, calculates the color distribution to get the proportion of main color system (blue accounts for 40%, red accounts for 35%), analyzes the texture pattern to identify 3 main brushstroke types, and evaluates the composition balance to calculate the symmetry index 0.82. At the same time, the creation description text is analyzed, the sentiment tendency score is extracted to be 0.75 (positive emotion), the creative vocabulary density is calculated to be 28%, and the grammatical complexity is measured to be an average sentence length of 22 words. For the accompanying creation process audio commentary, the pitch contour fluctuation range is extracted to be ±1.5 semitones, the rhythm stability deviation rate is 12%, and the timbre diversity index is 3.8. Finally, the multi-modal interaction mode is analyzed, and it is found that the color distribution is highly correlated with the sentiment tendency (correlation coefficient 0.85), and the texture pattern is synchronized with the rhythm of the audio commentary. All the features are encoded into a feature vector to input the evaluation engine.
[0042] The present application realizes the deep analysis of aesthetic education works through multi-level feature extraction, which can capture subtle and key feature information in artistic performance. Multi-dimensional feature combination provides a comprehensive and objective analysis basis, avoiding the one-sidedness of single feature evaluation. The professional feature design closely combines the characteristics of aesthetic education, making the evaluation results more professional and persuasive. The introduction of cross-modal features reveals the internal relationship between different artistic forms and enhances the system's understanding of comprehensive artistic performance. The extracted features contain both quantitative indicators and artistic characteristics, providing rich data support for accurate evaluation. This feature extraction method establishes a scientific analysis framework for aesthetic education evaluation and promotes the transition from subjective perception to objective analysis.
[0043] In one possible implementation, the evaluation engine module applies a pre-trained evaluation model based on a machine learning algorithm and outputs a score or classification result, and also includes adaptive learning to update the evaluation model.
[0044] The pre-trained evaluation model can refer to a machine learning model trained based on historical aesthetic education data to provide accurate and consistent aesthetic education evaluation benchmarks. The machine learning algorithm can refer to a mathematical method that allows a computer to automatically learn rules and patterns from data, enabling the establishment of complex mapping relationships between features and evaluation results. The score or classification result can refer to the quantitative score or category label output by the evaluation model to intuitively reflect the level of aesthetic performance. Adaptive learning can refer to a learning mechanism that allows the model to automatically adjust parameters based on new data to continuously optimize evaluation accuracy.
[0045] As a specific example: in music creation evaluation, the evaluation engine loads a pre-trained deep neural network model trained based on 5000 annotated works. The model receives a 12-dimensional feature vector provided by the feature extraction module, including pitch contour stability index 0.85, rhythm complexity coefficient 1.2, timbre diversity score 88 points, etc. Through forward propagation calculation, it outputs the creation skill score 92 points (percentage), innovation level A (total five levels ABCDE), and emotional expression intensity classification "intense". At the same time, the system adds this evaluation data to the training set and starts the adaptive learning process, fine-tunes the model weight parameters through the backpropagation algorithm, and makes the model better adapt to modern music creation characteristics. The updated model performs better in subsequent evaluations, especially in the evaluation accuracy of electronic music works.
[0046] The pre-trained model ensures the professionalism and consistency of the evaluation criteria, avoiding subjective bias in manual evaluation. The application of machine learning algorithms enables the system to handle complex aesthetic characteristics relationships and discover artistic rules that are difficult for the human eye to detect. The adaptive learning mechanism ensures that the system evolves over time and with data accumulation, continuously improving evaluation accuracy and applicability. The output form of scoring combined with classification provides both precise quantitative results and qualitative judgments required for artistic evaluation. The entire evaluation process realizes automated intelligent analysis, greatly improving the efficiency and scalability of aesthetic education evaluation, and providing reliable technical support for large-scale aesthetic education assessment.
[0047] In one possible implementation, when the evaluation engine module calculates the aesthetic evaluation score, the standardized feature value is calculated for each feature:
[0048] wherein, i represents the feature index, from 1 to N, N represents the total number of features, xi represents the i-th feature value, obtained from the feature extraction module, μi represents the mean value of the i-th feature, obtained from the pre-trained model parameters, σi represents the standard deviation of the i-th feature, obtained from the pre-trained model parameters; The aesthetic evaluation score is calculated as:
[0049] wherein, wi represents the weight of the i-th feature, obtained from the pre-trained model parameters, N represents the total number of features.
[0050] wherein, the feature index can refer to the sequential number used to uniquely identify each feature, used to distinguish and locate different features in the calculation process. The total number of features can refer to the number of all feature items participating in the evaluation, which can determine the complexity and calculation range of the evaluation system. The feature value can refer to the quantitative value extracted from the specific aesthetic data, which is used to reflect the actual performance level of the feature in the specific work. The mean value can refer to the average value of a certain feature in the training data set, which can provide a reference benchmark for the central tendency of the feature value. The standard deviation can refer to the dispersion degree index of a certain feature in the training data set, which can be used to measure the fluctuation range of the feature value. The standardized feature value can refer to the feature value after normalization by mean and standard deviation, which is used to eliminate the dimensional differences between different features. The feature weight can refer to the pre-set coefficient representing the importance of each feature, which can adjust the contribution proportion of different features to the final score. The aesthetic evaluation score can refer to the comprehensive score result calculated by weighted summation, which is used to quantify the overall level of aesthetic works.
[0051] As a specific example: in the evaluation of an oil painting work, the feature extraction module outputs 6 feature values: color contrast F1=85, composition balance F2=78, brushstroke complexity F3=92, color saturation F4=80, theme distinctness F5=88, and innovation F6=75. The evaluation engine obtains the corresponding mean values (μ1=70, μ2=75, μ3=80, μ4=78, μ5=82, μ6=72) and standard deviations (σ1=12, σ2=8, σ3=10, σ4=9, σ5=11, σ6=7) from the pre-trained model parameters, as well as the feature weights (W1=0.2, W2=0.15, W3=0.18, W4=0.12, W5=0.2, W6=0.15).
[0052] First, calculate the standardized value of each feature: Z1=(85-70) / 12=1.25, Z2=(78-75) / 8=0.375, Z3=(92-80) / 10=1.2, Z4=(80-78) / 9=0.222, Z5=(88-82) / 11=0.545, Z6=(75-72) / 7=0.429. Then calculate the aesthetic evaluation score: S1=0.2×1.25+0.15×0.375+0.18×1.2+0.12×0.222+0.2×0.545+0.15×0.429=0.25+0.056+0.216+0.027+0.109+0.064=0.722. Finally, map the score to the percentage system to get 88 points.
[0053] The present application eliminates the dimensional difference and distribution difference between different features through standardization processing, so that the feature values have comparability and additivity. The weighted summation calculation method can reasonably reflect the importance difference of different features, so that the evaluation result is more in line with the professional evaluation standard. The standardization benchmark based on the pre-trained model parameters ensures the consistency and objectivity of the evaluation result, avoiding subjective randomness. The whole calculation process has clear mathematical basis and interpretability, which is easy to understand and use. This calculation method can effectively integrate multiple feature dimensions to generate a comprehensive and accurate aesthetic evaluation result, providing a scientific and reliable quantitative tool for aesthetic evaluation.
[0054] In one possible implementation, the evaluation engine module uses the formula:
[0055] wherein, S represents the comprehensive aesthetic evaluation score, F represents the creativity score, which is calculated from the evaluation engine module, a technical skill score, calculated from the evaluation engine module, an aesthetic perception score, calculated from the evaluation engine module, a creativity weight coefficient, obtained from the pre-trained model parameters, a technical skill weight coefficient, obtained from the pre-trained model parameters, an aesthetic perception weight coefficient, obtained from the pre-trained model parameters.
[0056] The comprehensive aesthetic education evaluation score can refer to the final evaluation result calculated by multi-dimensional score weighting, which can reflect the overall quality level of the aesthetic education work. The creativity score can refer to the special score calculated by the evaluation engine according to the innovation and originality characteristics, which can measure the creative thinking performance of the work. The technical skill score can refer to the special score calculated by the evaluation engine according to the skill application and execution accuracy characteristics, which can evaluate the technical proficiency of the author. The aesthetic perception score can refer to the special score calculated by the evaluation engine according to the aesthetic law grasping and emotional expression characteristics, which can measure the aesthetic value of the work. The creativity weight coefficient can refer to the importance parameter of the creativity dimension in the pre-trained model, which can adjust the influence of creativity in the overall evaluation. The technical skill weight coefficient can refer to the importance parameter of the technical skill dimension in the pre-trained model, which can control the proportion of technical factors in the final result. The aesthetic perception weight coefficient can refer to the importance parameter of the aesthetic perception dimension in the pre-trained model, which can adjust the contribution of aesthetic factors to the overall score.
[0057] As a specific example: when evaluating a digital media art project, the evaluation engine first calculates three dimension scores: the creativity score C=85 points is obtained by analyzing the novelty and unique idea of the work, the technical skill score T=80 points is obtained by evaluating the proficiency of software operation and the accuracy of technical implementation, and the aesthetic perception score A=82 points is obtained by analyzing the color matching aesthetics and visual rhythm. Then the weight coefficients are obtained from the pre-trained model parameters: a=0.4 (emphasizing innovation), b=0.3 (technical implementation), c=0.3 (aesthetic value). Finally, the comprehensive aesthetic education evaluation score S2=0.4×85+0.3×80+0.3×82=34+24+24.6=82.6 points. The system automatically adjusts the weight coefficients according to the characteristics of the art form, for example, it increases the technical skill weight to 0.5 and reduces the creativity weight to 0.3 for traditional craft works.
[0058] The present application realizes the comprehensiveness and balance of aesthetic education evaluation through the multi-dimensional weighted fusion calculation method, focusing on both technical innovation and artistic value. The adjustable weight coefficient enables the system to adapt to the characteristics of different art forms, providing more accurate and personalized evaluation results. The hierarchical evaluation mechanism provides both specialized ability analysis and comprehensive conclusions, helping users to identify strengths and improvement directions. The weight setting based on the pre-trained model ensures the professionalism and consistency of the evaluation standard, avoiding subjective randomness. This comprehensive evaluation method can truly reflect the multi-faceted value of aesthetic works, providing a more scientific and reasonable quantitative system for aesthetic education evaluation.
[0059] In one possible implementation, the data preprocessing module further includes a data augmentation function for increasing image sample diversity through rotation or cropping.
[0060] The data augmentation function can refer to a processing component that generates training data variants through algorithms, which is used to expand the size of the data set and improve the generalization ability of the model. Rotation can refer to the operation of rotating the image around the center point by a certain angle, which can simulate the image changes caused by different shooting angles. Cropping can refer to a processing method that extracts part of the area from the original image, which can generate image samples with different composition characteristics. Image sample diversity can refer to the richness of images in the data set in terms of perspective, composition, content, etc., which is used to improve the adaptability of the evaluation model to different changes.
[0061] As a specific example: in the calligraphy work evaluation system, the data preprocessing module performs data augmentation on the collected 500 original calligraphy images. First, rotate each image by 5 degrees clockwise, 3 degrees counterclockwise, and 8 degrees, generating 1500 rotated samples. Then, perform cropping operations to randomly crop subgraphs from different regions of each image, including left upper corner, right lower corner, and center region cropping, generating a total of 4500 cropped samples. After the enhancement processing, the original 500 images are expanded to 6500 training samples, significantly improving the sample diversity. These enhanced images are input into the subsequent feature extraction module together with the original images, used to train a more robust evaluation model, enabling the model to accurately identify calligraphy works of different angles and close-up shots.
[0062] The present application effectively solves the problem of insufficient training data in aesthetic education evaluation through data augmentation, significantly improving the generalization ability and robustness of the model. Rotation and cropping operations simulate the possible changes in viewing angle and close-up situations in real-world scenarios, enabling the system to better handle data under various actual collection conditions. The increased sample diversity reduces the risk of model overfitting and improves the consistency and reliability of evaluation results. This data augmentation method provides more abundant training resources for aesthetic education evaluation systems, helping to establish more accurate and stable evaluation models. At the same time, the automated enhancement process greatly reduces the workload of manual data collection and labeling, improves the efficiency of system construction, and provides strong support for the promotion and application of aesthetic education evaluation.
[0063] In one possible implementation, the evaluation engine module further includes an adaptive learning function to update model parameters based on new data.
[0064] The adaptive learning function can refer to the system's ability to automatically adjust and optimize model parameters based on new input data, allowing the evaluation model to continuously improve and adapt to new forms of aesthetic expression. New data can refer to newly collected or newly labeled aesthetic evaluation data during system operation, providing the necessary training samples for model updates. Model parameters can refer to internal variables in a machine learning model that need to be determined through learning, which can affect the specific behavior and prediction accuracy of the model.
[0065] As a specific example: In a ceramic art evaluation system, the evaluation engine initially uses a model trained based on traditional pottery data. When the system begins to handle new 3D printed ceramic works, the adaptive learning function is activated. The system collects evaluation data for 100 new works, including expert scores and user feedback as new data. Through incremental learning algorithms, model parameters are gradually adjusted: first, update the weights of the feature extraction layer to identify the texture features of new materials, then adjust the parameters of the fully connected layer to adapt to new aesthetic standards. After 5 iteration cycles, the model's evaluation accuracy for 3D printed ceramic works improves from 65% initially to 89%, while maintaining stability in evaluating traditional pottery works. The updated model parameters are stored and used for all subsequent evaluation tasks.
[0066] The adaptive learning function enables the aesthetic education evaluation system to continuously evolve, constantly adapting to the development and changes of artistic forms. The parameter updating mechanism based on new data ensures that the evaluation standards keep pace with the times, avoiding the problem of outdated or ineffective models. This self-optimization capability reduces the need for human intervention, lowers system maintenance costs, and improves the accuracy and applicability of evaluation results. Adaptive learning enables the system to individually adapt to the aesthetic preferences of user groups, providing more accurate evaluation services. This function also enhances the system's ability to adapt to different regional cultural characteristics, providing technical support for cross-cultural applications of aesthetic education evaluation. Through continuous learning and improvement, the system can establish a more comprehensive and comprehensive aesthetic education evaluation system, providing long-term and reliable technical support for the development of aesthetic education.
[0067] It should be noted that, for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described actions, because according to the embodiments of the present specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of the present specification.
[0068] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0069] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details and do not limit the invention to only the described specific embodiments. Obviously, according to the content of the embodiments of the present specification, many modifications and changes can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is limited only by the claims and their entire scope and equivalents.
Claims
1. An aesthetic education evaluation processing system based on multi-source data fusion, characterized by, The system comprises a data acquisition module, a data preprocessing module, a data fusion module, a feature extraction module, an evaluation engine module, and a result output module. The data acquisition module acquires aesthetic education related data from multiple heterogeneous data sources and determines the original data. The data preprocessing module cleans and standardizes the original data and determines the multi-source data. The data fusion module integrates the preprocessed multi-source data into a unified data representation and determines the fusion data. The feature extraction module extracts key features from the fusion data. The evaluation engine module generates aesthetic education evaluation results based on the key features, and the result output module outputs the aesthetic education evaluation results.
2. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 1, characterized in that, The original data includes digital image data, video data, audio data, text data, and auxiliary data.
3. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 2, characterized in that, The data preprocessing module performs denoising, brightness adjustment, size normalization, and format conversion on the image data. Frame extraction, resolution unification, and key frame selection are performed on the video data. Noise reduction, sampling rate standardization, and silent segment removal are performed on the audio data. Tokenization, stop word removal, spelling correction, and encoding unification are performed on the text data.
4. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 1, characterized in that, The data fusion module adopts a combination of feature-level fusion and decision-level fusion strategies and adjusts the fusion proportion through an adaptive weight allocation method, wherein the fusion proportion is based on the reliability and relevance of the data sources.
5. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 1, characterized in that, The feature extraction module extracts visual features, auditory features, text features, and comprehensive features. The visual features include color distribution, texture pattern, and composition balance, the auditory features include pitch contour, rhythm stability, and timbre diversity, the text features include sentiment orientation, creative vocabulary density, and syntax complexity, and the comprehensive features include multi-modal interaction patterns.
6. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 1, characterized in that, The evaluation engine module applies a pre-trained evaluation model based on machine learning algorithms and outputs a score or classification result, and also includes adaptive learning to update the evaluation model.
7. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 6, characterized in that, When calculating the aesthetic education evaluation score, the evaluation engine module calculates the standardized feature value for each feature: wherein, denotes a feature index, from 1 to N, N being the total number of features, denotes the i-th feature value, obtained from the feature extraction module, denotes the mean of the i-th feature, obtained from the pre-trained model parameters, denotes the standard deviation of the i-th feature, obtained from the pre-trained model parameters; The evaluation engine module calculates the comprehensive aesthetic education evaluation score using the formula: wherein, represents the weight of the i-th feature, obtained from the pre-trained model parameters, and N represents the total number of features.
8. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 6, characterized in that, The data preprocessing module also includes a data augmentation function to increase image sample diversity through rotation or cropping. wherein, represents the comprehensive aesthetic evaluation score, represents the creativity score, calculated from the evaluation engine module, represents the technical skill score, calculated from the evaluation engine module, represents the aesthetic perception score, calculated from the evaluation engine module, represents the creativity weight coefficient, obtained from the pre-trained model parameters, represents the technical skill weight coefficient, obtained from the pre-trained model parameters, represents the aesthetic perception weight coefficient, obtained from the pre-trained model parameters.
9. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 1, characterized in that, The evaluation engine module also includes an adaptive learning function to update model parameters based on new data.
10. The aesthetic education evaluation processing system based on multi-source data fusion according to claim 6, characterized in that,
Citation Information
Patent Citations
Education evaluation method and system driven by multi-modal data
CN120047286A
Building design project intelligent evaluation method and system based on multi-dimensional indexes
CN120410330A
Intelligent assessment feedback method and system for artistic creation
CN120411708A
Modern service industry development level evaluation analysis method and system
CN120525181A
KR20220084751A