A method and system for generating a traditional Chinese visual style based on music features

Through multi-scale time-frequency analysis and genetic coding, the problem of mapping musical characteristics with traditional Chinese visual style has been solved, realizing the stable and controllable generation of traditional Chinese visual style, which is suitable for digital exhibition and cultural dissemination of traditional Chinese art.

CN122135676APending Publication Date: 2026-06-02LANZHOU YINQIAO CULTURAL COMMUNICATION CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LANZHOU YINQIAO CULTURAL COMMUNICATION CO LTD
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to structurally and parametrically map musical features to traditional Chinese visual styles, resulting in poor stability and insufficient stylistic consistency in the generated results. In particular, the lack of controllable parametric adjustment methods in traditional Chinese art application scenarios leads to low cultural recognition and difficulty in automatically generating visual expressions with traditional Chinese aesthetic characteristics.

Method used

Melody contours, rhythmic patterns, phonochromatic patterns, and emotional polarity fingerprints are extracted through multi-scale time-frequency analysis. These fingerprints are then genetically decomposed and parametrically encoded with the compositional paradigms, brushwork types, pattern motifs, and traditional chromatograms of traditional Chinese visual art works to construct a traditional visual style gene library. Cross-modal similarity and feature saliency are calculated, and the expression weights of style genes are dynamically calculated to drive the traditional visual style rendering model to generate images of traditional Chinese visual styles.

Benefits of technology

It achieves a precise mapping between musical features and visual style, improves the stability and stylistic consistency of the generated results, meets the aesthetic requirements of traditional Chinese art, provides controllable parametric adjustment methods, and is suitable for digital display and cultural dissemination of opera, folk instrumental music and multi-voice music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135676A_ABST
    Figure CN122135676A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of image generation technology and discloses a method and system for generating traditional Chinese visual styles based on musical features. The method includes: acquiring a music signal, extracting melody contour fingerprints, rhythmic pattern fingerprints, phonogram fingerprints, and emotional polarity fingerprints, and encoding them into corresponding fingerprint vectors; collecting traditional Chinese visual art works and performing gene-based decomposition and parameterized encoding to obtain multiple visual style genes; calculating the cross-modal similarity between each fingerprint vector and each visual style gene to determine the optimal combination of visual style genes corresponding to each fingerprint vector; dynamically calculating the expression weight and expression degree of each visual style gene to generate a style gene expression parameter set; and generating a traditional Chinese visual style image based on the style gene expression parameter set. This invention can form traditional Chinese visual style images with visual density and layer distribution, realizing intelligent generation of the entire process from music signals to traditional Chinese visual style images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image generation technology, and more specifically, to a method and system for generating traditional Chinese visual styles based on musical features. Background Technology

[0002] With the development of digital media technology, music visualization and visual generation technologies have been widely applied in stage performances, multimedia displays, digital art, and teaching aids. Existing technologies typically drive visual changes through basic acoustic parameters such as volume, spectrum, and rhythm to generate dynamic light and shadow, geometric shapes, or abstract animations. In recent years, deep learning-based image generation technology has made some progress, and some research has begun to explore cross-modal generation between audio and images. However, these studies mainly focus on general style transfer or realistic image generation, and the generated results often present abstract dynamic graphics or Western modern art styles, lacking the ability to systematically express traditional Chinese visual styles. In the context of traditional Chinese art, there is often an inherent aesthetic connection between music and visuals, such as the deep correspondence between rhythm and brushstrokes, melody and compositional trends, and timbre and color emotions. However, existing technologies have not yet formed a complete technical solution that structures, parameterizes, and systematically maps musical features to elements of traditional Chinese visual styles.

[0003] Furthermore, in existing music-driven visual generation technologies, the mapping relationship between musical features and visual elements largely relies on subjective experience or end-to-end black-box models. The mapping process lacks interpretability and reproducibility, resulting in poor stability and insufficient stylistic consistency in the generated results. Especially in the application scenarios of traditional Chinese arts, such as digital exhibitions and stage image generation of opera, folk instrumental music, and polyphonic music, existing technologies struggle to automatically generate visual expressions with traditional Chinese aesthetic characteristics based on the inherent features of the music. The generated content has low cultural recognition and lacks controllable parametric adjustment methods, making it difficult to reliably reuse it in scenarios such as teaching, cultural dissemination, and digital display. Therefore, there is an urgent need for a method that can establish a systematic mapping relationship between musical features and elements of traditional Chinese visual style, and achieve music-driven, style-controllable automatic generation of traditional Chinese visual styles.

[0004] In view of this, the present invention proposes a method and system for generating traditional Chinese visual styles based on musical features to solve the above problems. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art and achieve the above objectives, the present invention provides the following technical solution: a method for generating traditional Chinese visual styles based on musical features, comprising: Step S1: Acquire music signal, perform multi-scale time-frequency analysis on music signal, extract melody contour fingerprint, rhythm pattern fingerprint, phonogram fingerprint and emotion polarity fingerprint respectively, encode each type of fingerprint into corresponding fingerprint vector and calculate the feature salience in music signal; Step S2: Collect traditional Chinese visual art works, and perform genetic decomposition and parameter encoding of the composition paradigms, brushwork types, pattern motifs and traditional color spectrum in traditional Chinese visual art works to construct a traditional visual style gene library containing multiple visual style genes; Step S3: Calculate the cross-modal similarity between each fingerprint vector and each visual style gene in the traditional visual style gene library, and determine the optimal visual style gene combination corresponding to each fingerprint vector according to the preset pairing rules; Step S4: Based on the cross-modal similarity between each fingerprint vector and each visual style gene in the corresponding optimal visual style gene combination, and combined with the feature saliency corresponding to each fingerprint, dynamically calculate the expression weight and expression degree of each visual style gene to generate a style gene expression parameter set. Step S5: Based on the style gene expression parameter set, drive the preset traditional visual style rendering model, and perform parameterized superposition rendering according to the expression weight and degree of each visual style gene to generate a Chinese traditional visual style image.

[0006] Furthermore, the methods for extracting melody contour fingerprints, rhythmic pattern fingerprints, phonogram fingerprints, and emotional polarity fingerprints include: A pre-defined multi-scale analysis window parameter set is used, including window lengths and window jump steps corresponding to three analysis scales: short, medium, and long time windows. Based on the window length and window jump steps for each analysis scale, the music signal is processed sequentially to obtain the time-spectrum matrix for each analysis scale. The time-spectrum matrix corresponding to the medium time window is selected as the melody analysis time-spectrum, and the melody analysis time-spectrum is analyzed to obtain pitch sequences and rhythm sequences, which are then integrated to form a melody contour fingerprint. The time-spectrum matrix corresponding to the short time window is selected as the rhythm analysis time-spectrum, and the rhythm analysis time-spectrum is analyzed to obtain... The average beat interval, beat interval variability, dominant rhythm cycle, and starting interval sequence are collected and integrated to form a rhythmic pattern fingerprint. The time spectrum matrix corresponding to the medium time window is selected as the time spectrum for timbre analysis. The time spectrum for timbre analysis is analyzed to obtain the Mel cepstrum mean vector, Mel cepstrum variability vector, average spectral centroid, and average spectral flux, and is integrated to form a chromatic fingerprint. The time spectrum matrix corresponding to the long time window is selected as the time spectrum for emotion analysis. The time spectrum for emotion analysis is analyzed to obtain the tonality polarity sequence, dynamic intensity sequence, overall emotion polarity value, and emotion variability, and is integrated to form an emotion polarity fingerprint.

[0007] Furthermore, methods for constructing a traditional visual style gene pool include: We acquired data records of multiple traditional Chinese visual art works. Each record contains a unique work code, art category, and digital image of the work. For each digital image, we sequentially extracted features of composition paradigm, brushwork type, pattern motif, and traditional color spectrum, resulting in composition paradigm genes, brushwork type genes, pattern motif genes, and traditional color spectrum genes. The composition paradigm gene includes a set of compositional feature parameters and a composition gene vector; the brushwork type gene includes a set of brushwork feature parameters and a brushwork gene vector; the pattern motif gene includes a set of pattern feature parameters and a pattern gene vector; and the traditional color spectrum gene includes a set of color spectrum feature parameters and a color spectrum gene vector. Compositional paradigm genes, brushwork type genes, pattern motif genes, and traditional color spectrum genes are collectively referred to as visual style genes. All visual style genes corresponding to the digital image of each work are summarized and associated with the corresponding unique work code and art category to form the visual style genome of each Chinese traditional visual art work. The visual style genomes of all Chinese traditional visual art works are summarized to construct a traditional visual style gene library.

[0008] Furthermore, methods for calculating cross-modal similarity between each fingerprint vector and each visual style gene include: All visual style genes are obtained from the traditional visual style gene library, and the gene category and gene vector corresponding to each visual style gene are obtained; a unified embedding space mapping is performed on each fingerprint vector and each gene vector to obtain the standard fingerprint embedding vector and the standard gene embedding vector. A pre-defined set of aesthetic semantic axes is established, containing multiple aesthetic semantic axis vectors. Each standard fingerprint embedding vector is paired with each standard gene embedding vector to form multiple pairing combinations. For each aesthetic semantic axis, the corresponding audio axis activation intensity, visual axis activation intensity, and cross-modal resonance coefficient are calculated for each pairing combination. The audio axis activation intensity and visual axis activation intensity of all aesthetic semantic axes in the same pairing combination are arranged according to the order of aesthetic semantic axes in the set of aesthetic semantic axes to form audio activation mode vectors and visual activation mode vectors, respectively. The Pearson correlation coefficient between the audio activation mode vectors and the visual activation mode vectors is calculated to obtain the cross-axis coordination factor. Each aesthetic semantic axis is assigned a corresponding axis importance weight. Based on the axis importance weight, the cross-modal resonance coefficients of all aesthetic semantic axes in the same pairing are weighted and summed to obtain the basic resonance similarity. According to the basic resonance similarity and the cross-axis coordination factor, the original cross-modal similarity of each pairing is calculated, and the original cross-modal similarity is normalized to obtain the cross-modal similarity.

[0009] Furthermore, methods for determining the optimal visual style gene combination corresponding to each fingerprint vector include: A pre-defined cross-modal pairing rule table records the pairing relationships between each fingerprint type and each gene category; each pairing relationship includes fingerprint type, gene category, applicable pairing marker, and preferred pairing quantity; For each fingerprint vector, based on the corresponding fingerprint type, all gene categories marked as applicable for pairing are obtained from the cross-modal pairing rule table to form a candidate gene category set; all visual style genes whose gene categories belong to the candidate gene category set are screened from the traditional visual style gene library to form a candidate gene set; for each fingerprint vector, the cross-modal similarity between it and each visual style gene in the corresponding candidate gene set is obtained in turn, and visual style genes whose cross-modal similarity is less than a preset cross-modal threshold are removed from the corresponding candidate gene set to obtain an effective candidate gene set; The effective candidate gene set is grouped according to gene category to obtain the category gene group corresponding to each gene category in the effective candidate gene set; for each category gene group, all visual style genes are sorted from largest to smallest according to cross-modal similarity, and the visual style genes with the highest ranking and number not exceeding the corresponding preferred pairing are selected as the preferred genes under the corresponding gene category; the preferred genes under each gene category in the effective candidate gene set corresponding to each fingerprint vector are summarized to form the optimal visual style gene combination corresponding to each fingerprint vector.

[0010] Furthermore, methods for generating style gene expression parameter sets include: For each visual style gene in each optimal visual style gene combination, the corresponding expression excitation intensity is calculated sequentially; the expression excitation intensity of each visual style gene is subjected to nonlinear activation transformation to obtain the corresponding activation expression level; an aesthetic regulation matrix between genes is constructed, and the activation expression level of each visual style gene is iteratively regulated and updated to obtain the steady-state expression level, and the steady-state expression level of each visual style gene is taken as the corresponding expression degree. The optimal visual style gene combination corresponding to all fingerprint vectors is summarized, and all unique visual style genes are extracted to form a global candidate gene set. The global candidate gene set is then grouped according to gene category, resulting in a global gene group for each gene category. For each global gene group, all fingerprint types marked as applicable for the corresponding gene category are obtained from the cross-modal pairing rule table, and the feature saliency for each fingerprint type is acquired. Based on the feature saliency of each global gene group, the standard importance of each gene category is calculated. Based on the steady-state expression level of all visual style genes corresponding to each global gene group, the total expression of each category is calculated, and the relative intensity within each category is determined based on the total expression. The product of the standard importance and the relative intensity within each category is calculated to obtain the expression weight of each visual style gene. For each visual style gene in the global candidate gene set, the corresponding gene category, feature parameter set, expression weight and expression degree are integrated to form a gene expression parameter record; the gene expression parameter records of all visual style genes are summarized to generate a style gene expression parameter set.

[0011] Furthermore, methods for obtaining steady-state expression levels include: The system predefines homology and competition coefficients. For any two different visual style genes in the global candidate gene set, if the unique work codes of the two visual style genes are the same but the gene categories are different, the aesthetic regulation coefficient is set to the homology coefficient. If the gene categories of the two visual style genes are the same but the unique work codes are different, the aesthetic regulation coefficient is set to the negative of the competition coefficient. If the unique work codes and gene categories of the two visual style genes are both different, the aesthetic regulation coefficient is set to zero. The aesthetic regulation coefficients corresponding to the pairwise paired visual style genes in the global candidate gene set are integrated to form an aesthetic regulation matrix. The activation expression levels of all visual style genes are combined to form an expression level vector. A preset control step size, convergence threshold, and maximum number of iterations are established. Iterative control updates are then performed: matrix multiplication is performed on the aesthetic control matrix and the expression level vector to obtain a control influence vector; based on the control step size and the control influence vector, a control increment vector is calculated; the expression level vector and the control increment vector are added element-wise, and the expression level vector is updated based on the calculation results, with the iterative change calculated; if the iterative change is less than the convergence threshold or the maximum number of iterations is reached, the iteration stops, and the values ​​of each element in the expression level vector are taken as the steady-state expression level of the corresponding visual style gene; otherwise, the next iterative update continues.

[0012] Furthermore, methods for generating images in the traditional Chinese visual style include: From the style gene expression parameter set, obtain the gene category, feature parameter group, expression weight, and expression degree corresponding to each gene expression parameter record; for each gene expression parameter record, multiply each parameter value in the corresponding feature parameter group by the corresponding expression degree to obtain the expression scaling parameter group; preset the rendering canvas size and background color value; initialize the rendering canvas according to the rendering canvas size, and set the pixel value of each pixel in the rendering canvas to the background color value; initialize a load margin map with the same size as the rendering canvas, and set the load margin value of each pixel in the load margin map to one; The rendering hierarchy is preset, which specifies the execution order of the rendering layers corresponding to each gene category. The traditional visual style rendering model contains category rendering sub-models corresponding to each gene category. According to the rendering hierarchy, each gene category is used as the current gene category, and layered rendering and load margin adjustment are superimposed. After all gene categories have completed layered rendering and load margin adjustment, the pixel values ​​of all pixels in the rendering canvas are combined to form the final image data, which is then output as a traditional Chinese visual style image.

[0013] Furthermore, the methods for performing layered rendering and load capacity adjustment include: For each visual style gene under the current gene category, the corresponding expression scaling parameter set and the rendering canvas size are input into the corresponding category rendering sub-model to generate a single-gene rendering layer; the sum of the expression weights of all visual style genes under the current gene category is calculated to obtain the category superposition intensity; based on the expression weights and category superposition intensity, the category standard weight of each visual style gene is calculated; for each pixel in the rendering canvas, the pixel values ​​of all single-gene rendering layers under the current gene category at the corresponding pixel are weighted and summed based on the category standard weight to obtain the corresponding category fusion pixel value; Obtain the pixel value of each pixel in the rendering canvas and the corresponding pixel's carrying capacity value in the carrying capacity map; calculate the product of the carrying capacity value corresponding to the same pixel and the category stacking intensity to obtain the effective stacking coefficient of each pixel; calculate the stacked pixel value of each pixel based on the effective stacking coefficient and the category fusion pixel value, and update the pixel value of each pixel in the rendering canvas to the corresponding stacked pixel value; calculate the update carrying capacity value of each pixel based on the effective stacking coefficient and the preset level attenuation coefficient; update the carrying capacity value of the corresponding pixel in the carrying capacity map to the update carrying capacity value.

[0014] A system for generating traditional Chinese visual styles based on musical features, comprising the following steps: The audioprint encoding module is used to acquire music signals, perform multi-scale time-frequency analysis on the music signals, extract melody contour fingerprints, rhythm pattern fingerprints, audio chromatogram fingerprints and emotion polarity fingerprints respectively, encode each type of fingerprint into corresponding fingerprint vectors and calculate the feature salience in the music signal. The visual base construction module is used to collect traditional Chinese visual art works, and to perform genetic decomposition and parameter encoding of compositional paradigms, brushwork types, pattern motifs and traditional color spectrum in traditional Chinese visual art works, and to construct a traditional visual style gene library containing multiple visual style genes. The modality matching module is used to calculate the cross-modal similarity between each fingerprint vector and each visual style gene in the traditional visual style gene library, and to determine the optimal visual style gene combination corresponding to each fingerprint vector according to the preset matching rules. The gene regulation module is used to dynamically calculate the expression weight and expression level of each visual style gene based on the cross-modal similarity between each fingerprint vector and each visual style gene in the corresponding optimal visual style gene combination, and combined with the feature saliency corresponding to each fingerprint, to generate a set of style gene expression parameters. The style rendering module is used to drive a preset traditional visual style rendering model based on the style gene expression parameter set. It performs parameterized superposition rendering according to the expression weight and degree of each visual style gene to generate Chinese traditional visual style images.

[0015] The technical effects and advantages of the method and system for generating traditional Chinese visual styles based on musical features in this invention are as follows: By performing multi-scale time-frequency analysis on music signals and extracting four types of fingerprints—melody contour, rhythmic pattern, phonochromatic pattern, and emotional polarity—a comprehensive feature characterization of music signals in multiple dimensions, such as melody direction, rhythmic structure, timbre texture, and emotional tendency, was achieved. This overcomes the shortcomings of relying solely on a single musical feature dimension for visual generation, which results in a weak connection between the generated result and the musical connotation, thus improving the completeness of the musical feature representation. By genetically deconstructing and parametrically encoding the compositional paradigms, brushwork types, pattern motifs, and traditional color spectrums in traditional Chinese visual art, the complex stylistic elements of traditional visual art are deconstructed into quantifiable and combinable minimum parametric genetic units, providing a structured data foundation for the precise mapping between musical characteristics and visual style. By constructing an aesthetic semantic axis set and introducing a nonlinear resonance amplification mechanism based on minimum activation amplitude and cross-axis coordination factor correction, the music fingerprint vector and visual style gene vector are subjected to axial projection and axial resonance measurement in multiple abstract aesthetic dimensions. This breaks through the limitation of relying solely on the overall similarity of vectors for cross-modal matching, which cannot capture multi-dimensional aesthetic semantic associations. It realizes multi-dimensional aesthetic semantic matching between music fingerprints and visual style genes, and improves the interpretability and matching accuracy of cross-modal mapping. By introducing a cross-source synergistic amplification mechanism, visual style genes that are simultaneously activated by multiple musical features gain a superlinear expression enhancement advantage. Furthermore, the Hill equation from the field of biochemistry is interdisciplinaryly introduced into the expression control of visual style genes to achieve a nonlinear threshold response with adjustable steepness. At the same time, by constructing an aesthetic regulation matrix based on the dual rules of homologous synergy and homologous competition and performing iterative dynamic solutions, each visual style gene converges to a steady-state expression level that conforms to the harmony law of traditional Chinese art styles under the aesthetic mutual regulation, thereby avoiding the visual conflict problem caused by the simultaneous superposition of multiple visual styles. By introducing a load margin map to simulate the limited adsorption and load characteristics of traditional Chinese painting media on visual content, the superposition behavior of each pixel is simultaneously constrained by both the local load state and the global category importance. Furthermore, through a content adaptive attenuation mechanism associated with the effective superposition coefficient, the fully rendered areas automatically tend towards visual stability. Ultimately, a visual density distribution that conforms to the aesthetic principle of the interplay between reality and illusion is naturally formed in the generated traditional Chinese visual style image. The system achieves intelligent generation of the entire process from music signals to images in the traditional Chinese visual style, providing stable and reusable technical support for applications such as digital exhibition of opera, folk instrumental music, and multi-voice music, stage image generation, and cultural dissemination. Attached Figure Description

[0016] Figure 1 This is a flowchart of a method for generating traditional Chinese visual styles based on musical features, according to Embodiment 1 of the present invention. Figure 2 This is a schematic diagram of a traditional Chinese visual style generation system based on musical features, according to Embodiment 2 of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1:

[0018] Please see Figure 1 As shown in this embodiment, a method for generating traditional Chinese visual styles based on musical features includes: Step S1: Acquire music signal, perform multi-scale time-frequency analysis on music signal, extract melody contour fingerprint, rhythm pattern fingerprint, phonogram fingerprint and emotion polarity fingerprint respectively, encode each type of fingerprint into corresponding fingerprint vector and calculate the feature saliency in music signal.

[0019] Methods for acquiring music signals include: The system acquires raw audio data to be processed from the audio input interface. This raw audio data includes the sampling rate, number of channels, and audio sampling sequence. The sampling rate is the number of audio samples acquired per second. The number of channels is the number of independent audio channels in the audio signal. The audio sampling sequence consists of discrete audio amplitudes arranged in chronological order. The raw audio data is preprocessed to obtain a music signal. Specifically, if the number of channels is greater than one, the average of the audio sampling sequences for each channel is calculated to obtain the average sampling sequence. If the number of channels is equal to one, the audio sampling sequence is directly used as the average sampling sequence. The average sampling sequence is then normalized to obtain a standard audio sequence. The amplitude range of the discrete audio amplitudes in the standard audio sequence is [insert range here]. A pre-emphasis filter is applied to the standard audio sequence to compensate for the attenuation of high-frequency energy, thereby obtaining a music signal. The pre-emphasis filter is a first-order high-pass filter, and its specific implementation is a conventional technique in this field, which will not be elaborated on here.

[0020] Methods for multi-scale time-frequency analysis of music signals include: A pre-defined multi-scale analysis window parameter set is used, including window lengths and window jumps corresponding to three analysis scales: short window, medium window, and long window. The short window captures transient rhythmic initiation features in the music signal, the medium window captures melodic progression and timbre changes, and the long window captures overall tonality and emotional evolution. The window length is the number of sampling points in each analysis frame. The window jump is the sampling point offset between the corresponding starting positions of adjacent analysis frames. Each window length and window jump is pre-set by those skilled in the art based on the sampling rate and music analysis accuracy requirements. For each analysis scale, the music signal is framed using the corresponding window length, and frame shifting is performed according to the corresponding window jump, resulting in... Multiple analysis frames are used. A Hanning window function is applied to each analysis frame to suppress spectral leakage, followed by a short-time Fourier transform to calculate the spectral amplitude of each frequency component corresponding to that frame. The Hanning window function is a symmetric bell-shaped window function used to smooth the edges of analysis frames and reduce truncation effects. The spectral amplitude of each frequency component in each analysis frame is squared to obtain the spectral energy value corresponding to each frequency component. The spectral energy values ​​of each analysis frame at the same analysis scale are sorted according to time series to form a time-frequency spectrum matrix for each analysis scale. The row dimension of the time-frequency spectrum matrix corresponds to the frequency axis, the column dimension corresponds to the time axis, and the matrix elements are the spectral energy values ​​at the corresponding time-frequency positions. The time-frequency spectrum matrices of the three analysis scales are integrated to form multi-scale time-frequency representation data.

[0021] Methods for extracting melody contour fingerprints include: From the multi-scale time-frequency characterization data, the time-spectrum matrix corresponding to the mid-time window is selected as the time-spectrum for melody analysis. For each analysis frame of the melody analysis time-spectrum, the frequency component with the maximum spectral energy value is detected within a preset fundamental frequency range, and the frequency value of the corresponding frequency component is used as the fundamental frequency estimate of the corresponding analysis frame. The fundamental frequency range is preset by those skilled in the art based on the fundamental frequency distribution characteristics of human voice and musical instruments. The fundamental frequency estimates of all analysis frames are arranged in chronological order to obtain a fundamental frequency trajectory sequence. A logarithmic frequency transformation is performed on each fundamental frequency estimate in the fundamental frequency trajectory sequence to convert the fundamental frequency estimate into pitch values ​​in semitones, resulting in a pitch sequence. The expression for the logarithmic frequency transformation is: In the formula, The pitch value. This is the estimated fundamental frequency. The preset reference frequency (usually 440 Hz) is used; the difference between adjacent pitch values ​​in the pitch sequence is calculated to obtain the pitch sequence; the pitch sequence is used to describe the size and direction of the interval between adjacent notes in the melody; the pitch sequence and the pitch sequence are integrated to form the melody profile fingerprint; the melody profile fingerprint is used to characterize the melody direction and pitch motion characteristics of the music signal.

[0022] Methods for extracting rhythm pattern fingerprints include: From multi-scale time-frequency characterization data, the time-spectrum matrix corresponding to a short time window is selected as the time-spectrum for rhythm analysis. The energy of the rhythm analysis time-spectrum is summed along the frequency axis to obtain the instantaneous energy value of each analysis frame, forming an energy envelope sequence. The instantaneous energy value is the sum of the spectral energy values ​​of all frequency components within a single analysis frame. The first-order difference is calculated on the energy envelope sequence, and positive values ​​are retained to obtain the initial detection function. Each element in the initial detection function has the same frame index as the corresponding analysis frame. In the initial detection function, the frame indices corresponding to local peaks greater than a preset initial threshold are detected, resulting in a set of peak frame indices. Local peaks are elements in the initial detection function whose values ​​are greater than the values ​​of their adjacent left and right sides. The initial threshold is preset by those skilled in the art based on the dynamic range characteristics of the music type. For each frame index in the peak frame index set, the product of the frame index and the window step corresponding to the short time window is calculated, and then divided by the sampling rate. The rhythmic pattern is obtained by calculating the rate and corresponding beat start times; arranging all beat start times in chronological order to obtain a beat time sequence; calculating the time interval between adjacent beat start times in the beat time sequence to obtain a start interval sequence; performing statistical analysis on the start interval sequence to calculate the mean and standard deviation, obtaining the average beat interval and beat interval volatility, respectively; where the average beat interval reflects the overall tempo characteristics of the music, and the beat interval volatility reflects the regularity of the rhythm; performing autocorrelation calculation on the start interval sequence, extracting the interval values ​​corresponding to the periodic peaks in the autocorrelation function to obtain the dominant rhythm period; integrating the average beat interval, beat interval volatility, dominant rhythm period, and start interval sequence to form a rhythm pattern fingerprint; where the rhythm pattern fingerprint is used to characterize the rhythm tempo, beat regularity, and temporal structure characteristics of the music signal; autocorrelation calculation is a well-known technique in this field, and the specific process will not be elaborated here.

[0023] Methods for extracting chromatographic fingerprints include: From the multi-scale time-frequency characterization data, the time-spectrum matrix corresponding to the mid-time window is selected as the time-spectrum for timbre analysis. Mel frequency cepstral coefficient extraction is performed on each analysis frame in the timbre analysis time-spectrum to obtain the Mel cepstral coefficient set for each analysis frame. The extraction process of the Mel frequency cepstral coefficients is a well-known technique in the field, and the specific process will not be elaborated here. The order of the Mel frequency cepstral coefficients is preset by those skilled in the art according to the accuracy requirements of timbre characterization. The mean and standard deviation of the Mel cepstral coefficient sets for all analysis frames are calculated along the time axis to obtain the Mel cepstral mean vector and the Mel cepstral fluctuation vector. For each analysis frame of the timbre analysis time-spectrum, a weighted average is calculated using the spectral energy value as the weight to obtain the spectral centroid. The spectral centroid values ​​of all analysis frames are arranged in chronological order to form a spectral centroid sequence. The Euclidean distance between adjacent analysis frames for all corresponding spectral amplitudes is calculated to obtain the spectral flux value between every two adjacent analysis frames. All spectral flux values ​​are then arranged in chronological order to form a spectral flux sequence. The spectral centroid sequence reflects the change in timbre brightness over time, while the spectral flux sequence reflects the intensity of timbre changes. The mean values ​​of the spectral centroid sequence and the spectral flux sequence are calculated to obtain the average spectral centroid and average spectral flux, respectively. The Mel cepstral mean vector, Mel cepstral fluctuation vector, average spectral centroid, and average spectral flux are integrated to form an chromatic fingerprint. The chromatic fingerprint is used to characterize the timbre texture and spectral distribution characteristics of the music signal.

[0024] Methods for extracting emotional polarity fingerprints include: From multi-scale time-frequency representation data, the time-spectrum matrix corresponding to a long time window is selected as the time-spectrum for emotion analysis. The time-spectrum for emotion analysis is divided into multiple emotion analysis segments along the time axis, and the duration of each emotion analysis segment is preset by those skilled in the art based on the typical duration of the emotional state. For each emotion analysis segment, the projection intensity of the spectral energy value onto the preset major and minor harmonic frequency templates is calculated, and the major and minor harmonic intensities are obtained respectively. The major and minor harmonic frequency templates are both preset by those skilled in the art based on music harmony theory. The projection intensity is calculated as follows: The spectral energy values ​​of all analysis frames within the emotion analysis segment are summed along the time axis to obtain the cumulative spectral energy value corresponding to each frequency component. The cumulative spectral energy values ​​of all frequency components are then combined to form a spectral energy distribution vector. The inner product between the spectral energy distribution vector and the corresponding harmonic frequency template is calculated to obtain the projection intensity. The difference between the major and minor harmonic intensities is calculated and then divided by the sum of the two to obtain the tonality polarity value. A positive tonality polarity value indicates a bright and positive emotional tone in the corresponding segment, while a negative value indicates a dark and somber tone. The emotional color of stillness; the energy of each emotional analysis segment in the frequency spectrum is summed along the frequency axis to obtain the emotional energy value of each analysis frame within the emotional analysis segment, forming a segment energy envelope; the mean of the energy envelope of each segment is calculated sequentially to obtain the average energy of each emotional analysis segment; the ratio of the average energy of each segment to the mean of the average energy of all emotional analysis segments is calculated to obtain the dynamic intensity ratio of each emotional analysis segment; the dynamic intensity ratio is used to reflect the energy strength of the corresponding segment relative to the whole piece; the tonality values ​​and dynamic intensity ratios of all emotional analysis segments are then compared over time. The sequences are arranged sequentially to form a tonal polarity sequence and a dynamic intensity sequence, respectively. The mean of the tonal polarity sequence is calculated to obtain the overall emotional polarity value. The standard deviation of the dynamic intensity sequence is calculated to obtain the emotional fluctuation degree. The overall emotional polarity value is used to reflect the overall emotional tendency and intensity of the entire piece. The emotional fluctuation degree is used to reflect the amplitude of the emotional energy fluctuation of the entire piece. The tonal polarity sequence, dynamic intensity sequence, overall emotional polarity value and emotional fluctuation degree are integrated to form an emotional polarity fingerprint. The emotional polarity fingerprint is used to characterize the emotional tendency, emotional evolution trajectory and emotional fluctuation characteristics of the music signal.

[0025] Methods for encoding various fingerprints into corresponding fingerprint vectors include: For the melody contour fingerprint, the pitch sequence is resampled at equal intervals to a preset melody encoding dimension to obtain a standardized pitch vector; the interval sequence is also resampled at equal intervals to the melody encoding dimension to obtain a standardized interval vector; where equal-interval resampling is an interpolation operation that resamples the original sequence at equal intervals to adjust the sequence length to the target dimension; the standardized pitch vector and the standardized interval vector are concatenated to form the melody contour fingerprint vector; for the rhythm pattern fingerprint, the starting interval sequence is resampled at equal intervals to a preset rhythm encoding dimension to obtain a standardized rhythm sequence vector; the average beat interval, beat interval variability, and dominant rhythm period are concatenated after the standardized rhythm sequence vector to form the rhythm pattern fingerprint vector; For chromatographic fingerprints, the Mel-Cepstral mean vector, Mel-Cepstral fluctuation vector, average spectral centroid, and average spectral flux are sequentially concatenated to form a chromatographic fingerprint vector. For emotional polarity fingerprints, the tonality polarity sequence and dynamic intensity sequence are resampled at equal intervals to a preset emotional encoding dimension to obtain a standardized tonality polarity vector and a standardized dynamic intensity vector. The standardized tonality polarity vector, standardized dynamic intensity vector, overall emotional polarity value, and emotional fluctuation are sequentially concatenated to form an emotional polarity fingerprint vector. Each encoding dimension is pre-set by those skilled in the art according to the dimensional matching requirements for subsequent cross-modal similarity calculations. L2 norm normalization is performed on each fingerprint vector to make the magnitude of each fingerprint vector equal to one.

[0026] Methods for calculating the saliency of various fingerprint features in music signals include: For the melody profile fingerprint, the ratio of the standard deviation of the pitch sequence to the range of pitch sequence values ​​is calculated to obtain the melody activity level; where the range of pitch sequence values ​​is the difference between the maximum and minimum values ​​in the pitch sequence; melody activity level reflects the degree of activity of melody pitch changes; for each analysis frame of the spectrum during melody analysis, the spectral energy value of the frequency component corresponding to the fundamental frequency estimate is obtained and used as the fundamental frequency energy value; the sum of the spectral energy values ​​of all frequency components in each analysis frame is calculated to obtain the full-band energy value; the fundamental frequency energy values ​​of all analysis frames are summed to obtain the total fundamental frequency energy; the full-band energy values ​​of all analysis frames are summed to obtain the total full-band energy; the ratio of the total fundamental frequency energy to the total full-band energy is calculated to obtain the melody energy proportion; where the melody energy proportion reflects the prominence of the melody component in the full-band energy; corresponding melody significance weights are set for melody activity level and melody energy proportion respectively, and a weighted summation is performed based on the melody significance weights to obtain the feature significance of the melody profile fingerprint; For rhythm pattern fingerprints, the rhythm sharpness is obtained by calculating the ratio between the mean of all local peaks and the mean of all elements in the initial detection function. Rhythm sharpness reflects the prominence of the beat initiation event relative to the overall signal. The rhythm regularity is obtained by calculating the reciprocal of the beat interval fluctuation. Rhythm regularity reflects the uniformity of the beat interval. Corresponding rhythm saliency weights are assigned to rhythm sharpness and rhythm regularity, and a weighted sum is calculated based on these rhythm saliency weights to obtain the feature saliency of the rhythm pattern fingerprint. For chromatic fingerprints, the mean of the absolute values ​​of each element in the Mel-Cepstral fluctuation vector is calculated to obtain the timbre variability; the timbre variability reflects the average fluctuation amplitude of timbre characteristics over time. The ratio of the standard deviation of the spectral centroid sequence to the average spectral centroid is calculated to obtain the timbre dynamics; the timbre dynamics reflects the time-varying degree of timbre brightness and darkness. Corresponding timbre significance weights are assigned to the timbre variability and timbre dynamics, and a weighted sum is calculated based on the timbre significance weights to obtain the feature significance of the chromatic fingerprint. For the emotional polarity fingerprint, the absolute value of the overall emotional polarity value is calculated to obtain the emotional tendency intensity; whereby the emotional tendency intensity is used to reflect the degree to which the overall emotion is biased towards a certain polarity direction; the emotional fluctuation degree is used as the emotional evolution intensity, which is used to reflect the intensity of the overall emotional change; corresponding emotional significance weights are set for the emotional tendency intensity and the emotional evolution intensity, and a weighted sum is calculated based on the emotional significance weights to obtain the feature significance of the emotional polarity fingerprint; It should be noted that each saliency weight is preset by those skilled in the art based on the sensitivity requirements of music feature-driven visual generation; the feature saliency of all fingerprints is normalized so that the sum of all feature saliencies is one.

[0027] Step S2: Collect traditional Chinese visual art works, and perform genetic decomposition and parametric coding of the composition paradigms, brushwork types, pattern motifs and traditional color spectrum in traditional Chinese visual art works to construct a traditional visual style gene library containing multiple visual style genes.

[0028] Methods for collecting traditional Chinese visual art include: Data records of multiple traditional Chinese visual art works were obtained from the Traditional Art Digital Resource Database. The Traditional Art Digital Resource Database is a pre-built data platform used to store and manage digital information of traditional Chinese visual art works. Each work data record includes a unique work code, art category, and digital image of the work. The unique work code is a predefined unique identifier for the work. The art category is used to identify the traditional art category to which the traditional Chinese visual art work belongs, including but not limited to traditional Chinese painting, porcelain patterns, brocade patterns, architectural painting, lacquerware decoration, etc. The digital image of the work is the image data obtained after the traditional Chinese visual art work has been digitally collected.

[0029] Methods for genetically deconstructing and parametrically encoding compositional paradigms, brushwork types, pattern motifs, and traditional color palettes in traditional Chinese visual art include: For each digital image of a work, compositional paradigm features are extracted. Specifically, the digital image is segmented into foreground and background regions to obtain foreground and background areas. The number of corresponding pixels in the background region and the digital image is counted to obtain the background pixel count and the image pixel count. The ratio of the background pixel count to the image pixel count is calculated to obtain the white space ratio. The white space ratio reflects the proportion of blank areas in the image. The specific implementation of foreground and background segmentation is a conventional technique in this field and will not be elaborated on here. Pixels in the foreground region are marked as foreground points, and the grayscale value and image coordinates of each foreground point are obtained. The image coordinates include horizontal and vertical coordinates. Using the grayscale value of each foreground point as a weight, the horizontal and vertical coordinates of each foreground point are weighted and averaged to obtain the horizontal and vertical components of the center of gravity. The horizontal and vertical components of the center of gravity together constitute the visual center of gravity coordinates, which are used to represent the spatial center of convergence of visual elements in the image. The foreground region is then connected... The process involves: performing a domain analysis to obtain the centroid coordinates of each connected domain; sorting the centroid coordinates of each connected domain in ascending order of their corresponding vertical coordinates; calculating the absolute value of the difference between the corresponding vertical coordinates of adjacent connected domains to form a layer spacing sequence; calculating the mean of the layer spacing sequence to obtain the average layer spacing, which reflects the density of layers in the depth direction of the image; calculating the standard deviation of the layer spacing sequence to obtain the layer spacing fluctuation, which reflects the uniformity of the layer distribution in the image; connected domain analysis is a well-known technique in this field, and the specific process will not be elaborated here; integrating the white space ratio, the horizontal component of the center of gravity, the vertical component of the center of gravity, the average layer spacing, and the layer spacing fluctuation to form a composition feature parameter set; concatenating the parameter values ​​in the composition feature parameter set into a one-dimensional vector and performing L2 norm normalization to obtain a composition gene vector; integrating the composition feature parameter set and the composition gene vector, and labeling the gene category as a composition paradigm to form a composition paradigm gene.

[0030] For each digital image of a work, brushstroke type features are extracted. Specifically, brushstroke detection is performed on the digital image of the work to obtain brushstroke regions and a set of brushstroke skeleton lines composed of multiple skeleton lines. The brushstroke region is the image area in the digital image containing brushstroke traces; the skeleton line is the center line of the brushstroke region along its direction. The specific implementation of brushstroke detection is a conventional technique in this field and will not be elaborated upon here. For each skeleton line in the set of brushstroke skeleton lines, the brushstroke width value at each sampling point is measured along the normal direction of the skeleton line to obtain a brushstroke width sequence. The stroke width value is the distance between the two stroke edges perpendicular to the direction of the skeleton line at the sampling point. The measurement method for stroke width is a conventional technique in this field and will not be elaborated upon here. The mean of all stroke width values ​​in the stroke width sequence corresponding to all skeleton lines is calculated to obtain the average stroke width. The average stroke width is used to reflect the overall thickness of the stroke. The standard deviation of all stroke width values ​​is calculated and then divided by the average stroke width to obtain the stroke width variation rate. The stroke width variation rate is used to reflect the richness of stroke thickness variation. For each... A skeleton line is used, and the local curvature value at each sampling point on the skeleton line is calculated. The method for calculating the local curvature value is a well-known technique in the field and will not be elaborated upon here. The average curvature of all local curvature values ​​of all skeleton lines is calculated, reflecting the degree of curvature of the brushstroke lines. The digital image of the artwork is converted to grayscale to obtain the grayscale value of each pixel within the brushstroke area. The average grayscale value of all pixels within the brushstroke area is calculated, yielding the average ink density. The average ink density reflects the overall density of the ink used. The brush... The standard deviation of the grayscale values ​​of all pixels within the touch area is calculated to obtain the ink density fluctuation. The ink density fluctuation reflects the amplitude of ink depth variation. The average stroke width, stroke width change rate, average line curvature, mean ink density, and ink density fluctuation are integrated to form a brush technique feature parameter set. The parameter values ​​in the brush technique feature parameter set are concatenated into a one-dimensional vector and L2 norm normalization is performed to obtain a brush technique gene vector. The brush technique feature parameter set and the brush technique gene vector are integrated, and the gene category is labeled as brush technique type to form a brush technique type gene.

[0031] For each digital image of a work, pattern motif features are extracted. Specifically, pattern region detection and segmentation are performed on the digital image of the work to extract the pattern regions. The specific implementation of pattern region detection and segmentation is a conventional technique in this field and will not be elaborated upon here. The pattern regions are then flipped horizontally to obtain a horizontal mirror image. This horizontal flip involves symmetrically mapping the horizontal coordinates of each pixel in the pattern region about the horizontal center of the image. The normalized cross-correlation coefficient between the pattern region and the horizontal mirror image is calculated to obtain the horizontal symmetry. The pattern regions are then flipped vertically to obtain a vertical mirror image. This vertical flip involves symmetrically mapping the vertical coordinates of each pixel in the pattern region about the vertical center of the image. The normalized cross-correlation coefficient between the pattern region and the vertical mirror image is calculated to obtain the vertical symmetry. The mean of the horizontal and vertical symmetry is calculated to obtain the pattern symmetry. The pattern symmetry is used to reflect... The symmetry and regularity of the pattern are assessed. The calculation method for the normalized cross-correlation coefficient is a well-known technique in the field, and the specific process will not be elaborated here. The number of corresponding pixels in the pattern area is counted to obtain the pattern pixel quantity. The ratio of the pattern pixel quantity to the image pixel quantity is calculated to obtain the pattern filling density. The pattern filling density is used to reflect the density of the pattern distribution in the image. Edge detection is performed on the pattern area to obtain the pixels located at the edge of the pattern area and mark them as pattern edge points. The ratio of the number of pattern edge points to the pattern pixel quantity is calculated to obtain the pattern edge density. The pattern edge density is used to reflect the geometric complexity of the pattern. The pattern symmetry, pattern filling density, and pattern edge density are integrated to form a pattern feature parameter set. The parameter values ​​in the pattern feature parameter set are concatenated into a one-dimensional vector and L2 norm normalization is performed to obtain the pattern gene vector. The pattern feature parameter set and the pattern gene vector are integrated, and the gene category is marked as a pattern motif to form a pattern motif gene.

[0032] For each digital image of the artwork, traditional color spectrum feature extraction is performed. Specifically, the digital image is converted from the RGB color space to the HSV color space to obtain the hue, saturation, and lightness values ​​corresponding to each pixel. Hue represents the color type attribute of the pixel, saturation represents the vividness of the color, and lightness represents the brightness of the color. The conversion method between RGB and HSV color spaces is a well-known technique in the field and will not be elaborated upon here. Pixels with saturation values ​​higher than a preset saturation lower limit and lightness values ​​higher than a preset lightness lower limit are selected and marked as valid color points. The saturation and lightness lower limits are preset by those skilled in the art based on color perception characteristics. A circular mean is calculated for the hue values ​​of all valid color points to obtain the principal hue angle. The principal hue angle reflects the overall color tendency of the image. A circular standard deviation is calculated for the hue values ​​of all valid color points to obtain the hue dispersion. The hue dispersion reflects the richness of the color types in the image. The circular mean and circular standard deviation calculations are for angle-type periodicity. The methods for calculating the mean and standard deviation are conventional techniques in this field and will not be elaborated upon here. The average saturation is calculated by averaging the saturation values ​​of all effective color points; this average saturation reflects the overall color vibrancy of the image. Based on the hue values ​​of each effective color point, effective color points with hue values ​​within a preset warm color range are marked as warm color points, and those with hue values ​​within a preset cool color range are marked as cool color points. The warm and cool color ranges are defined by those skilled in the art based on color... The theoretical settings are pre-defined; the ratio between the number of warm and cool color points is calculated to obtain the warm / cool ratio; the warm / cool ratio reflects the overall warm / cool color tendency of the image; the main hue angle, hue dispersion, average saturation, and warm / cool ratio are integrated to form a chromatographic feature parameter group; the parameter values ​​in the chromatographic feature parameter group are sequentially concatenated into a one-dimensional vector and subjected to L2 norm normalization to obtain a chromatographic gene vector; the chromatographic feature parameter group and the chromatographic gene vector are integrated, and the gene category is labeled as traditional chromatography to form a traditional chromatographic gene.

[0033] Methods for constructing traditional visual style gene libraries include: Compositional paradigm genes, brushwork type genes, pattern motif genes, and traditional color spectrum genes are collectively referred to as visual style genes. All visual style genes corresponding to the digital image of each artwork are compiled and associated with the corresponding unique artwork code and art category, forming the visual style genome of each traditional Chinese visual art artwork. The visual style genome is the collection of all visual style genes extracted from a single traditional Chinese visual art artwork. Each visual style gene is the smallest parameterized unit extracted from a traditional Chinese visual art artwork to represent a specific visual style characteristic. Each visual style gene contains a gene category, gene vector, and feature parameter set. The visual style genomes of all traditional Chinese visual art artworks are compiled to construct a traditional visual style gene library.

[0034] Step S3: Calculate the cross-modal similarity between each fingerprint vector and each visual style gene in the traditional visual style gene library, and determine the optimal visual style gene combination corresponding to each fingerprint vector according to the preset pairing rules.

[0035] Methods for calculating cross-modal similarity between each fingerprint vector and each visual style gene include: All visual style genes are obtained from a traditional visual style gene database, along with the corresponding gene category and gene vector for each visual style gene. A unified embedding space mapping is then performed on each fingerprint vector and each gene vector. Specifically, a unified embedding dimension and a cross-modal projection matrix set are pre-defined. The unified embedding dimension is the dimension of the common vector space used for cross-modal similarity calculation, and is pre-set by those skilled in the art based on the cross-modal representation accuracy requirements. The cross-modal projection matrix set includes the fingerprint projection matrix corresponding to each fingerprint vector and the gene projection matrix corresponding to each gene category. Each projection matrix is ​​defined by those skilled in the art. Based on the cross-modal correlation characteristics between musical features and visual styles, pre-setting is performed. For each fingerprint vector, the corresponding fingerprint projection matrix and the fingerprint vector are multiplied together, and the result is normalized using the L2 norm to obtain a standard fingerprint embedding vector. For each visual style gene, the corresponding gene projection matrix is ​​obtained according to the gene category. The gene projection matrix and the gene vector are multiplied together, and the result is normalized using the L2 norm to obtain a standard gene embedding vector. Both the standard fingerprint embedding vector and the standard gene embedding vector are normalized vectors with a dimension equal to the unified embedding dimension.

[0036] A pre-defined set of aesthetic semantic axes is provided, containing multiple aesthetic semantic axis vectors. Each aesthetic semantic axis vector is a unit vector with a dimension equal to the unified embedding dimension, used to represent an abstract aesthetic dimension direction shared by music and visual arts. The aesthetic semantic axes include, but are not limited to, tension axes, density axes, dynamic axes, and temperature axes. Among them, the tension axis is used to represent the aesthetic feature dimension from relaxed and soothing to tense and rousing; the density axis is used to represent the aesthetic feature dimension from sparse and simple to dense and rich; the dynamic axis is used to represent the aesthetic feature dimension from serene and calm to dynamic and changing; and the temperature axis is used to represent the aesthetic feature dimension from cool and elegant to passionate and vibrant. Each aesthetic semantic axis vector is pre-set by those skilled in the art based on traditional Chinese aesthetic theory and music aesthetic theory. Each standard fingerprint embedding vector is paired with each standard gene embedding vector to form multiple pairing combinations. Each pairing combination includes one standard fingerprint embedding vector and one standard gene embedding vector. For each pairing combination, cross-modal resonance similarity calculation based on aesthetic semantic axes is performed. Specifically, the dot product between the standard fingerprint embedding vector and each aesthetic semantic axis vector is calculated to obtain the audio axis activation intensity corresponding to each aesthetic semantic axis. Here, the dot product is the sum of the products of corresponding elements in the two vectors. The dot product between the standard gene embedding vector and each aesthetic semantic axis vector is also calculated to obtain the visual axis activation intensity corresponding to each aesthetic semantic axis. Here, the audio axis activation intensity and visual axis activation intensity are used to reflect the magnitude and direction of the projection components of the fingerprint vector and gene vector in the corresponding aesthetic semantic axis directions, respectively. For each aesthetic semantic axis, the corresponding cross-modal resonance coefficient is calculated. Specifically, for each aesthetic semantic axis, the product of the corresponding audio axis activation intensity and the visual axis activation intensity is calculated to obtain the axial activation product. The smaller of the absolute values ​​of the audio axis activation intensity and the visual axis activation intensity is taken to obtain the minimum activation amplitude. The minimum activation amplitude reflects the lowest activation level that audio and vision share on the corresponding aesthetic semantic axis. The product of the preset resonance amplification factor and the minimum activation amplitude is calculated, and then the natural exponential function value is taken to obtain the resonance amplification factor. The expression for the resonance amplification factor is as follows: In the formula, The resonance amplification factor, This is the resonance amplification factor. The minimum activation amplitude is set; the resonance amplification factor is preset by those skilled in the art based on the requirements of cross-modal correlation sensitivity; the resonance amplification factor is used to achieve a nonlinear amplification effect when both audio and visual have strong activation on the same aesthetic semantic axis; the product of the axial activation product and the resonance amplification factor is calculated to obtain the cross-modal resonance coefficient of the corresponding aesthetic semantic axis; the cross-modal resonance coefficient is used to measure the aesthetic resonance intensity of the fingerprint vector and the gene vector in the direction of the corresponding aesthetic semantic axis; when the activation intensity of the audio axis and the activation intensity of the visual axis have the same sign and both have large absolute values, the cross-modal resonance coefficient is positively amplified; when the two have opposite signs, the cross-modal resonance coefficient is negative, indicating that the aesthetic directions are contradictory; The audio and visual activation intensities of all aesthetic semantic axes in the same pairing are arranged according to the order of aesthetic semantic axes in the set to form audio activation pattern vectors and visual activation pattern vectors, respectively. The Pearson correlation coefficient between the audio and visual activation pattern vectors is calculated to obtain the cross-axis coordination factor. The calculation method of the Pearson correlation coefficient is a well-known technique in the field, and the specific process will not be elaborated here. The cross-axis coordination factor is used to reflect the overall consistency of the activation patterns of the fingerprint vector and the gene vector on each aesthetic semantic axis. The closer the cross-axis coordination factor is to positive one, the more coordinated the response patterns of the two are on each aesthetic dimension. Each aesthetic semantic axis is assigned a corresponding axis importance weight, which is pre-set by those skilled in the art based on the importance of each aesthetic dimension in the generation of traditional Chinese visual styles. The cross-modal resonance coefficients of all aesthetic semantic axes in the same pairing are weighted and summed based on the axis importance weights to obtain the basic resonance similarity. The product of a pre-set coordination amplification coefficient and a cross-axis coordination factor is calculated, and then one is added to obtain a coordination correction factor. The coordination amplification coefficient is pre-set by those skilled in the art based on the degree of influence of aesthetic coordination on matching quality. The coordination correction factor is used to correct the overall coordination of the basic resonance similarity. The product of the basic resonance similarity and the coordination correction factor is calculated to obtain the original cross-modal similarity. The original cross-modal similarity of all pairings is normalized to obtain the cross-modal similarity. The cross-modal similarity is used to measure the comprehensive matching degree between the fingerprint vector and the visual style gene in the multi-dimensional aesthetic semantic space.

[0037] Methods for determining the optimal visual style gene combination corresponding to each fingerprint vector include: A pre-defined cross-modal pairing rule table records the pairing relationships between each fingerprint type and each gene category. Fingerprint types include melody contours, rhythmic patterns, phonochromatic patterns, and emotional polarities; gene categories include compositional paradigms, brushstroke types, pattern motifs, and traditional color palettes. Each pairing relationship includes fingerprint type, gene category, applicable pairing marker, and preferred pairing quantity. The applicable pairing marker indicates whether the corresponding fingerprint type is allowed to be paired with the corresponding gene category, with a value of "applicable" or "not applicable." The preferred pairing quantity is the maximum number of visual style genes selected from the corresponding gene category. The cross-modal pairing rule table is pre-set by those skilled in the art based on their experience with the aesthetic correlation between musical characteristics and visual style elements. A cross-modal threshold is preset, which is pre-set by those skilled in the art according to the pairing quality requirements. For each fingerprint vector, based on the corresponding fingerprint type, all gene categories marked as applicable for pairing are obtained from the cross-modal pairing rule table to form a candidate gene category set. All visual style genes whose gene categories belong to the candidate gene category set are screened from the traditional visual style gene library to form a candidate gene set. For each fingerprint vector, the cross-modal similarity between it and each visual style gene in the corresponding candidate gene set is obtained in turn, and visual style genes with cross-modal similarity less than the cross-modal threshold are removed from the corresponding candidate gene set to obtain an effective candidate gene set. The effective candidate gene set is grouped according to gene category to obtain category gene groups corresponding to each gene category in the effective candidate gene set. The category gene group is a subset of all visual style genes belonging to the same gene category in the effective candidate gene set. For each category gene group, all visual style genes are sorted from largest to smallest according to cross-modal similarity, and the visual style genes with the highest ranking and number not exceeding the corresponding pairing preference number are selected as the preferred genes under the corresponding gene category. The preferred genes under each gene category in the effective candidate gene set corresponding to each fingerprint vector are summarized to form the optimal visual style gene combination corresponding to each fingerprint vector. The optimal visual style gene combination is the set of visual style genes with the highest cross-modal similarity selected from each gene category under the constraint of cross-modal pairing rules.

[0038] It should be understood that step S3, by constructing a set of aesthetic semantic axes, projects and compares the music fingerprint vector and the visual style gene vector along the axes in terms of abstract aesthetic dimensions such as tension, density, dynamics, and temperature. This breaks through the limitation of relying solely on the overall cosine similarity of the vectors for cross-modal matching, and achieves a refined measurement of the multi-dimensional aesthetic semantic relationship between music and vision. By introducing a nonlinear resonance amplification mechanism based on minimum activation amplitude, a superlinear resonance enhancement effect is generated when both audio and vision have strong responses in the same aesthetic dimension. This effectively strengthens the pairing signal with consistent aesthetic direction, while naturally suppressing weak association pairing with only unilateral activation. By calculating the cross-axis coordination factor to make overall coordination correction to the basic resonance similarity, it ensures that the final matching result not only matches in a single aesthetic dimension, but also maintains consistency in the overall outline of the multi-dimensional aesthetic mode. This makes the cross-modal pairing result more in line with the overall aesthetic principle of "spirit resonance" in traditional Chinese aesthetics.

[0039] Step S4: Based on the cross-modal similarity between each fingerprint vector and each visual style gene in the corresponding optimal visual style gene combination, and combined with the feature saliency corresponding to each fingerprint, dynamically calculate the expression weight and expression degree of each visual style gene to generate a style gene expression parameter set.

[0040] Methods for generating style gene expression parameter sets include: For each visual style gene in each optimal visual style gene combination, the corresponding expression excitation intensity is calculated sequentially. Specifically, the optimal visual style gene combinations corresponding to all fingerprint vectors are summarized, and all unique visual style genes are extracted to form a global candidate gene set. For each visual style gene in the global candidate gene set, all optimal visual style gene combinations corresponding to fingerprint vectors are traversed to determine all fingerprint vectors containing the corresponding visual style gene, and the number of corresponding fingerprint vectors is counted to obtain the fingerprint activation count. The fingerprint activation count reflects the extent to which visual style genes are simultaneously selected by different types of music features. For each fingerprint vector containing each visual style gene, the corresponding cross-modal similarity and the feature saliency of the corresponding fingerprint type are obtained, and the product of cross-modal similarity and feature saliency is calculated to obtain the single-source excitation value. The single-source excitation value reflects the weighted activation contribution of a single fingerprint vector to a single visual style gene. The method involves: summing all single-source activation values ​​corresponding to the same visual style gene to obtain the cumulative activation value of each visual style gene; pre-setting a co-amplification base, which is pre-set by those skilled in the art based on the strength of cross-modal synergistic effects; calculating the fingerprint activation count power of the co-amplification base to obtain the cross-source synergistic factor of each visual style gene; wherein, the cross-source synergistic factor is used to realize the superlinear amplification effect when the visual style gene is activated by multiple different types of music features simultaneously; when the fingerprint activation count is one, the cross-source synergistic factor is equal to the co-amplification base; when the fingerprint activation count is greater than one, the cross-source synergistic factor grows exponentially, reflecting the synergistic enhancement effect of multi-source activation; calculating the product of the cumulative activation value corresponding to the same visual style gene and the cross-source synergistic factor to obtain the expression activation intensity; wherein, the expression activation intensity is used to reflect the overall activation level of the visual style gene after comprehensively considering cross-modal matching degree, feature significance and multi-source synergistic effects;

[0041] The expression intensity of each visual style gene is nonlinearly activated to obtain the corresponding activation expression level. Specifically, an expression activation threshold and a co-sensitivity coefficient are preset. The expression activation threshold is used to control the critical activation intensity for the visual style gene to transition from a low expression state to a high expression state, and is preset by those skilled in the art according to the visual style expression clarity requirements. The co-sensitivity coefficient is used to control the steepness of the transition region between low and high expression, and is preset by those skilled in the art according to the visual style gene expression differentiation requirements. For each visual style gene, the corresponding expression activation intensity is substituted into the Hill activation function to obtain the activation expression level. The expression of the Hill activation function is: In the formula, To activate expression levels, To express the intensity of the incentive, To express the activation threshold, The co-sensitivity coefficient is used; the Hill activation function, borrowed from the Hill equation describing molecular binding dynamics in biochemistry, is used in this embodiment to achieve the nonlinear threshold response characteristics of visual style gene expression. When the expression stimulus intensity is much less than the expression activation threshold, the activation expression level approaches zero, indicating that the corresponding visual style gene is in an inhibited state; when the expression stimulus intensity is much greater than the expression activation threshold, the activation expression level approaches one, indicating that the corresponding visual style gene is in a saturated expression state. The larger the co-sensitivity coefficient, the steeper the activation transition, and the more significant the expression differentiation effect of visual style genes.

[0042] An aesthetic regulation matrix among genes was constructed and iteratively updated to obtain steady-state expression levels. Specifically, a homology synergy coefficient and a homology competition coefficient were preset, both of which were positive. The homology synergy coefficient quantifies the intensity of aesthetic synergy between visual style genes from different gene categories originating from the same traditional Chinese visual art work, and is preset by those skilled in the art based on the stylistic consistency rules within traditional art works. The homology competition coefficient quantifies the intensity of expression competition inhibition between visual style genes belonging to the same gene category but originating from different traditional Chinese visual art works, and is preset by those skilled in the art based on the exclusivity characteristics of homology visual style genes. For any two genes in the global candidate gene set... For each visual style gene, an aesthetic regulation coefficient is determined based on its unique work code and gene category. If two visual style genes have the same unique work code but different gene categories, the aesthetic regulation coefficient is set to the homologous synergy coefficient, indicating a mutually promoting aesthetic synergy relationship between them. If two visual style genes have the same gene category but different work codes, the aesthetic regulation coefficient is set to the opposite of the homologous competition coefficient, indicating a mutually inhibiting expression competition relationship between them. If two visual style genes have different work codes and gene categories, the aesthetic regulation coefficient is set to zero. The aesthetic regulation coefficients corresponding to the paired visual style genes in the global candidate gene set are integrated to form an aesthetic regulation matrix. The activation expression levels of all visual style genes are combined to form an expression level vector. Based on the stability and computational efficiency requirements of the regulatory dynamics, a pre-defined regulatory step size, convergence threshold, and maximum number of iterations are established. Iterative regulatory updates are then performed: matrix multiplication is performed on the aesthetic regulation matrix and the expression level vector to obtain the regulatory influence vector; the product of the regulatory step size and the regulatory influence vector is calculated to obtain the regulatory increment vector; the expression level vector and the regulatory increment vector are added element-wise, and the expression level vector is updated based on the calculation results; each element in the updated expression level vector is truncated, with elements less than zero set to zero and elements greater than one set to one; the sum of the absolute values ​​of the differences between corresponding elements in the expression level vector before and after the update is calculated to obtain the iterative change; if the iterative change is less than the convergence threshold or the maximum number of iterations is reached, the iteration stops, and the values ​​of each element in the expression level vector are taken as the steady-state expression level of the corresponding visual style gene; otherwise, the next iterative update continues. The steady-state expression level reflects the balanced expression state achieved by visual style genes after aesthetic synergy and competitive regulation among genes.

[0043] Based on the steady-state expression level, the expression weight and expression degree of each visual style gene are calculated. Specifically, the steady-state expression level of each visual style gene is taken as the corresponding expression degree. The expression degree controls the intensity scaling ratio of each parameter value in the corresponding feature parameter group of the visual style gene during rendering; the closer the expression degree is to one, the more fully the visual features of the visual style gene are presented. The global candidate gene set is grouped according to gene category, resulting in a global gene group corresponding to each gene category in the global candidate gene set. For each global gene group, all fingerprint types marked as applicable for the corresponding gene category are obtained from the cross-modal pairing rule table, and the feature saliency corresponding to each fingerprint type is obtained. The summation of all feature saliencies corresponding to the same global gene group is calculated to obtain the category importance of each gene category. The category importance is used to reflect... The importance of gene categories driven by musical features in the overall visual style generation is determined. The category importance of all gene categories is normalized to obtain the standard importance of each gene category. For each global gene group, the sum of the steady-state expression levels of all visual style genes within it is calculated to obtain the category expression sum. If the category expression sum is greater than zero, the ratio of the steady-state expression level of each visual style gene to the corresponding category expression sum is calculated to obtain the relative intensity within the category. If the category expression sum is equal to zero, the relative intensity within the category of each visual style gene within the global gene group is set to the reciprocal of the number of visual style genes in the corresponding global gene group. The product of the standard importance and the relative intensity within the category corresponding to the same visual style gene is calculated to obtain the expression weight of each visual style gene. The expression weight is used to determine the relative contribution ratio of visual style genes in the final image rendering.

[0044] For each visual style gene in the global candidate gene set, the corresponding gene category, feature parameter set, expression weight and expression degree are integrated to form a gene expression parameter record; the gene expression parameter records of all visual style genes are summarized to generate a style gene expression parameter set.

[0045] It should be understood that step S4 introduces a cross-source collaborative amplification mechanism, enabling visual style genes that are simultaneously activated by multiple different types of music features to gain an exponential expression enhancement advantage, thereby effectively distinguishing core style genes that are consistently recognized by multiple dimensions of music signals from marginal style genes that are only accidentally matched by a single feature; by introducing the Hill equation from the field of biochemistry into the expression control of visual style genes, a nonlinear threshold response characteristic with adjustable steepness is achieved, so that the expression state of visual style genes forms a clear differentiation boundary between inhibition and saturation, avoiding the problem of visual style blurring caused by a large number of visual style genes participating in rendering at similar intermediate expression levels at the same time; By constructing an aesthetic regulation matrix based on the dual rules of homologous synergy and similar competition and solving iterative dynamics, different categories of genes from the same traditional artwork tend to express synergistically to maintain the inherent consistency of style. At the same time, similar genes from different works form competitive inhibition to avoid style conflict. Ultimately, under the mutual regulation of aesthetics, each visual style gene converges to a steady-state expression level that conforms to the coordination law of traditional Chinese art style. This effectively simulates the visual organization law of "distinct primary and secondary, and interplay of emptiness and fullness" in the creation process of traditional Chinese art, thereby ensuring that the generated visual style image presents a harmonious and unified traditional aesthetic appearance in multiple dimensions such as composition, brushwork, pattern and color.

[0046] Step S5: Based on the style gene expression parameter set, drive the preset traditional visual style rendering model, and perform parameterized superposition rendering according to the expression weight and degree of each visual style gene to generate a Chinese traditional visual style image.

[0047] Methods for generating images in the traditional Chinese visual style include: From the style gene expression parameter set, obtain the gene category, feature parameter group, expression weight, and expression degree corresponding to each gene expression parameter record; for each gene expression parameter record, multiply each parameter value in the corresponding feature parameter group by the corresponding expression degree to obtain the expression scaling parameter group; where the expression scaling parameter group is used to represent the visual style feature parameters after expression degree adjustment, the closer the expression degree is to one, the more fully the corresponding visual style feature is presented, the closer the expression degree is to zero, the weaker the corresponding visual style feature is presented;

[0048] The rendering canvas size and background color value are preset. The rendering canvas size is the width and height in pixels of the output image, preset by those skilled in the art based on the output image resolution requirements. The background color value is the initial background color value of the rendering canvas, preset by those skilled in the art based on the color characteristics of traditional Chinese painting materials. The rendering canvas is initialized according to its size, and the pixel value of each pixel in the rendering canvas is set to the background color value. A carrying capacity map with the same size as the rendering canvas is initialized, and the carrying capacity value of each pixel in the carrying capacity map is set to one. The carrying capacity map is used to simulate the absorption and diffusion characteristics of ink information by traditional Chinese painting media (such as Xuan paper and silk), recording the remaining visual content carrying capacity of each pixel in the rendering canvas. This limits the superposition of visual information in different areas during image generation, naturally forming a density and white space structure characteristic of traditional Chinese painting. The carrying capacity value ranges from [missing value]. The higher the carrying capacity value, the greater the amount of subsequent visual content that the corresponding pixel can accept; the lower the carrying capacity value, the more the corresponding pixel has been fully occupied by the visual content of the previous rendering layer, and the weaker its ability to accept subsequent rendering layers. A preset rendering hierarchy order is defined, specifying the execution order of rendering layers corresponding to each gene category. The rendering hierarchy order is as follows: composition paradigm layer, brushstroke type layer, pattern motif layer, and traditional color spectrum layer. This rendering hierarchy order is pre-set by those skilled in the art based on the creative process of traditional Chinese painting. The traditional visual style rendering model includes category rendering sub-models corresponding to each gene category. These category rendering sub-models are pre-constructed parametric image generation models that generate corresponding category visual style images based on input feature parameter sets and canvas size. The category rendering sub-model corresponding to the composition paradigm is used to generate images based on the composition... The feature parameter set generates the spatial layout of the image; the category rendering sub-model corresponding to the brushstroke type is used to generate brushstroke texture effects based on the brushstroke feature parameter set; the category rendering sub-model corresponding to the pattern motif is used to generate decorative pattern patterns based on the pattern feature parameter set; the category rendering sub-model corresponding to the traditional color spectrum is used to generate color rendering effects based on the color spectrum feature parameter set; each category rendering sub-model includes, for example, a parametric drawing model based on procedural generation, an image synthesis model based on conditional generative adversarial networks, and an image rendering model based on style transfer networks, etc. The specific construction methods of each category rendering sub-model are conventional techniques in this field and will not be elaborated on here.

[0049] Following the rendering hierarchy, each gene category is sequentially treated as the current gene category, and layered rendering and load margin adjustment are overlaid. Specifically, for the current gene category, all gene expression parameter records belonging to the current gene category are obtained. For each visual style gene under the current gene category, the corresponding expression scaling parameter set and the rendering canvas size are input into the corresponding category rendering sub-model to generate a single gene rendering layer. The single gene rendering layer is image data with the same size as the rendering canvas, containing visual style content generated by the category rendering sub-model based on the input feature parameter set. The sum of the expression weights of all visual style genes under the current gene category is calculated to obtain the category overlay strength; where the category overlay strength is used to reflect the visual overlay intensity of the current gene category in the overall rendering; the ratio of the expression weight of each visual style gene to the category overlay strength is calculated to obtain the category standard weight; where the category standard weight is used to determine the relative contribution ratio of each visual style gene under the same gene category to the category fusion layer. For each pixel in the rendering canvas, the pixel values ​​of all single-gene rendering layers under the current gene category are weighted and summed based on the category standard weight to obtain the category fusion pixel value of each pixel; the category fusion pixel values ​​of all pixels are integrated to form a category fusion layer; the category fusion layer is a comprehensive visual layer after the rendering effects of all visual style genes under the current gene category are weighted and fused according to the category standard weight. A preset layer attenuation coefficient is used to control the attenuation rate of the carrying capacity margin after each rendering overlay. This coefficient is preset by those skilled in the art based on the decreasing visual carrying capacity between layers in traditional Chinese painting. The pixel value of each pixel in the rendering canvas is obtained, along with the carrying capacity margin value of the corresponding pixel in the carrying capacity margin map. The product of the carrying capacity margin value corresponding to the same pixel and the category overlay intensity is calculated to obtain the effective overlay coefficient for each pixel. If the effective overlay coefficient is greater than one, it is set to one. The effective overlay coefficient reflects the overlay ratio that the visual content of the current gene category can actually be accepted by the corresponding pixel under the carrying capacity margin constraint. Based on the effective overlay coefficient, the overlay pixel value of each pixel is calculated, and the pixel value of each pixel in the rendering canvas is updated to the corresponding overlay pixel value. The expression for the overlay pixel value is: In the formula, To overlay pixel values, For effective superposition coefficients, For pixel values, The pixel values ​​are fused for each category; based on the effective overlay coefficient, the update margin value for each pixel is calculated; the expression for the update margin value is: In the formula, To update the margin value, This is the bearing capacity margin value. The layer decay coefficient is used; if the update margin value is less than zero, the update margin value is set to zero; the load margin value of the corresponding pixel in the load margin map is updated to the update margin value; the update mechanism significantly reduces the load capacity of pixels that have been occupied by a large amount of visual content in subsequent layer rendering, while pixels with less visual content retain a higher load capacity, thus naturally forming a density hierarchy of visual content distribution in the overall image.

[0050] After layered rendering and load margin adjustment are completed for all gene categories, the pixel values ​​of all pixels in the rendering canvas are combined to form the final image data, which is then output as an image in the traditional Chinese visual style.

[0051] It should be understood that step S5, by introducing a carrying capacity map to simulate the limited adsorption and carrying capacity characteristics of traditional Chinese painting media for visual content, allows the rendering canvas to naturally reduce its capacity to accept subsequent layers of content as visual content accumulates during the layer-by-layer stacking process. This avoids the visual overload and layer confusion problems caused by the simple stacking of layers with fixed transparency in traditional digital image synthesis. By multiplying the carrying capacity value by the category stacking intensity to obtain the effective stacking coefficient at the pixel level, the stacking behavior of each pixel is simultaneously constrained by both the local carrying capacity state and the global category importance, thus achieving pixel-level adaptive rendering control. By adopting a content adaptive decay mechanism associated with the effective stacking coefficient in the carrying capacity update, areas that have been fully rendered automatically tend to become visually stable in subsequent layers, while areas that retain more carrying capacity can still accept rich rendering of subsequent layers. This naturally forms a visual density and layer distribution in the final image that conforms to the aesthetic principle of the interplay of emptiness and fullness in traditional Chinese painting.

[0052] This embodiment achieves a comprehensive characterization of music signals across multiple dimensions, including melody contour, rhythmic pattern, phonological spectrum, and emotional polarity, through multi-scale time-frequency analysis of music signals and extraction of four types of fingerprints: melody profile, rhythmic pattern, phonological spectrum, and emotional polarity. This overcomes the shortcomings of relying solely on a single musical feature dimension for visual generation, which results in a weak connection between the generated result and the musical content, thus enhancing the completeness of musical feature representation. Furthermore, by genetically deconstructing and parametrically encoding compositional paradigms, brushwork types, pattern motifs, and traditional phonological spectrums in traditional Chinese visual art, the complex stylistic elements of traditional visual art are deconstructed into measurable... This study utilizes modular, composable, minimally parameterized gene units to provide a structured data foundation for the precise mapping between musical features and visual styles. By constructing a set of aesthetic semantic axes and introducing a nonlinear resonance amplification mechanism based on minimum activation amplitude and cross-axis coordination factor correction, it performs axial projection and per-axis resonance measurement on musical fingerprint vectors and visual style gene vectors across multiple abstract aesthetic dimensions. This overcomes the limitation of relying solely on overall vector similarity for cross-modal matching, which fails to capture multidimensional aesthetic semantic associations. It achieves multidimensional aesthetic semantic matching between musical fingerprints and visual style genes, improving the interpretability and matching accuracy of cross-modal mapping. Furthermore, by introducing cross-source collaboration... The amplification mechanism enables visual style genes simultaneously activated by multiple musical features to achieve a superlinear expression enhancement advantage. Furthermore, the Hill equation from biochemistry is interdisciplinaryly introduced into the expression control of visual style genes to achieve a nonlinear threshold response with adjustable steepness. Simultaneously, by constructing an aesthetic regulation matrix based on both homologous synergy and homologous competition rules and performing iterative dynamic solutions, the visual style genes converge to a steady-state expression level that conforms to the harmony rules of traditional Chinese art styles under aesthetic mutual regulation, thus avoiding the visual conflict caused by the simultaneous superposition of multiple visual styles. Finally, a carrying capacity map is introduced to simulate the media of traditional Chinese painting. The limited adsorption and carrying capacity of visual content means that the superposition behavior of each pixel is simultaneously constrained by both the local carrying state and the global category importance. Through the content adaptive attenuation mechanism associated with the effective superposition coefficient, the fully rendered areas automatically tend towards visual stability. Ultimately, a visual density distribution that conforms to the aesthetic principle of the interplay between reality and illusion is naturally formed in the generated traditional Chinese visual style image. The whole system realizes intelligent generation of the entire chain from music signal to traditional Chinese visual style image, providing stable and reusable technical support for application scenarios such as digital exhibition of opera, folk instrumental music and multi-voice music, stage image generation, and cultural dissemination. Example 2:

[0053] Please see Figure 2As shown in the figure, the parts not described in detail in this embodiment are described in Embodiment 1. A traditional Chinese visual style generation system based on music features is provided, including a phonetic encoding module, a visual base construction module, a modality pairing module, a gene regulation module, and a style rendering module; each module is connected by wired and / or wireless means to realize data transmission between modules.

[0054] The audioprint encoding module is used to acquire music signals, perform multi-scale time-frequency analysis on the music signals, extract melody contour fingerprints, rhythm pattern fingerprints, audio chromatogram fingerprints and emotion polarity fingerprints respectively, encode each type of fingerprint into corresponding fingerprint vectors and calculate the feature salience in the music signal. The visual base construction module is used to collect traditional Chinese visual art works, and to perform genetic decomposition and parameter encoding of compositional paradigms, brushwork types, pattern motifs and traditional color spectrum in traditional Chinese visual art works, and to construct a traditional visual style gene library containing multiple visual style genes. The modality matching module is used to calculate the cross-modal similarity between each fingerprint vector and each visual style gene in the traditional visual style gene library, and to determine the optimal visual style gene combination corresponding to each fingerprint vector according to the preset matching rules. The gene regulation module is used to dynamically calculate the expression weight and expression level of each visual style gene based on the cross-modal similarity between each fingerprint vector and each visual style gene in the corresponding optimal visual style gene combination, and combined with the feature saliency corresponding to each fingerprint, to generate a set of style gene expression parameters. The style rendering module is used to drive a preset traditional visual style rendering model based on the style gene expression parameter set. It performs parameterized superposition rendering according to the expression weight and degree of each visual style gene to generate Chinese traditional visual style images. Example 3:

[0055] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable code that, when executed by the one or more processors, can perform a method for generating traditional Chinese visual styles based on musical features, as described above.

[0056] The method or system according to the embodiments of this application can also be implemented using the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or hard disk, may store the method for generating traditional Chinese visual styles based on musical features provided in this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is merely exemplary; when implementing different devices, one or more components in the electronic device shown in this application may be omitted according to actual needs. Example 4:

[0057] One embodiment of this application discloses a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions. When executed by a processor, the computer-readable instructions can perform a method for generating traditional Chinese visual styles based on musical features, as described above with reference to the accompanying drawings, according to an embodiment of this application. The storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0058] Furthermore, according to embodiments of this application, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, this application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be executed by a processor to perform instructions corresponding to the method steps provided in this application, such as a method for generating traditional Chinese visual styles based on musical features. When this computer program is executed by a central processing unit (CPU), it performs the functions defined in the method of this application.

[0059] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0060] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0061] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for generating traditional Chinese visual styles based on musical features, characterized in that, include: Step S1: Acquire music signal, perform multi-scale time-frequency analysis on music signal, extract melody contour fingerprint, rhythm pattern fingerprint, phonogram fingerprint and emotion polarity fingerprint respectively, encode each type of fingerprint into corresponding fingerprint vector and calculate the feature salience in music signal; Step S2: Collect traditional Chinese visual art works, and perform genetic decomposition and parameter encoding of the composition paradigms, brushwork types, pattern motifs and traditional color spectrum in traditional Chinese visual art works to construct a traditional visual style gene library containing multiple visual style genes; Step S3: Calculate the cross-modal similarity between each fingerprint vector and each visual style gene in the traditional visual style gene library, and determine the optimal visual style gene combination corresponding to each fingerprint vector according to the preset pairing rules; Step S4: Based on the cross-modal similarity between each fingerprint vector and each visual style gene in the corresponding optimal visual style gene combination, and combined with the feature saliency corresponding to each fingerprint, dynamically calculate the expression weight and expression degree of each visual style gene to generate a style gene expression parameter set. Step S5: Based on the style gene expression parameter set, drive the preset traditional visual style rendering model, and perform parameterized superposition rendering according to the expression weight and degree of each visual style gene to generate a Chinese traditional visual style image.

2. The method for generating traditional Chinese visual styles based on musical features according to claim 1, characterized in that, Methods for extracting melody contour fingerprints, rhythmic pattern fingerprints, phonogram fingerprints, and emotional polarity fingerprints include: A pre-defined multi-scale analysis window parameter set is used, including window lengths and window jump steps corresponding to three analysis scales: short, medium, and long time windows. Based on the window length and window jump steps for each analysis scale, the music signal is processed sequentially to obtain the time-spectrum matrix for each analysis scale. The time-spectrum matrix corresponding to the medium time window is selected as the melody analysis time-spectrum, and the melody analysis time-spectrum is analyzed to obtain pitch sequences and rhythm sequences, which are then integrated to form a melody contour fingerprint. The time-spectrum matrix corresponding to the short time window is selected as the rhythm analysis time-spectrum, and the rhythm analysis time-spectrum is analyzed to obtain... The average beat interval, beat interval variability, dominant rhythm cycle, and starting interval sequence are collected and integrated to form a rhythmic pattern fingerprint. The time spectrum matrix corresponding to the medium time window is selected as the time spectrum for timbre analysis. The time spectrum for timbre analysis is analyzed to obtain the Mel cepstrum mean vector, Mel cepstrum variability vector, average spectral centroid, and average spectral flux, and is integrated to form a chromatic fingerprint. The time spectrum matrix corresponding to the long time window is selected as the time spectrum for emotion analysis. The time spectrum for emotion analysis is analyzed to obtain the tonality polarity sequence, dynamic intensity sequence, overall emotion polarity value, and emotion variability, and is integrated to form an emotion polarity fingerprint.

3. The method for generating traditional Chinese visual styles based on musical features according to claim 2, characterized in that, Methods for constructing traditional visual style gene libraries include: We acquired data records of multiple traditional Chinese visual art works. Each record contains a unique work code, art category, and digital image of the work. For each digital image, we sequentially extracted features of composition paradigm, brushwork type, pattern motif, and traditional color spectrum, resulting in composition paradigm genes, brushwork type genes, pattern motif genes, and traditional color spectrum genes. The composition paradigm gene includes a set of compositional feature parameters and a composition gene vector; the brushwork type gene includes a set of brushwork feature parameters and a brushwork gene vector; the pattern motif gene includes a set of pattern feature parameters and a pattern gene vector; and the traditional color spectrum gene includes a set of color spectrum feature parameters and a color spectrum gene vector. Compositional paradigm genes, brushwork type genes, pattern motif genes, and traditional color spectrum genes are collectively referred to as visual style genes. All visual style genes corresponding to the digital image of each work are summarized and associated with the corresponding unique work code and art category to form the visual style genome of each Chinese traditional visual art work. The visual style genomes of all Chinese traditional visual art works are summarized to construct a traditional visual style gene library.

4. The method for generating traditional Chinese visual styles based on musical features according to claim 3, characterized in that, Methods for calculating cross-modal similarity between each fingerprint vector and each visual style gene include: All visual style genes are obtained from the traditional visual style gene library, and the gene category and gene vector corresponding to each visual style gene are obtained; a unified embedding space mapping is performed on each fingerprint vector and each gene vector to obtain the standard fingerprint embedding vector and the standard gene embedding vector. A pre-defined set of aesthetic semantic axes is established, containing multiple aesthetic semantic axis vectors. Each standard fingerprint embedding vector is paired with each standard gene embedding vector to form multiple pairing combinations. For each aesthetic semantic axis, the corresponding audio axis activation intensity, visual axis activation intensity, and cross-modal resonance coefficient are calculated for each pairing combination. The audio axis activation intensity and visual axis activation intensity of all aesthetic semantic axes in the same pairing combination are arranged according to the order of aesthetic semantic axes in the set of aesthetic semantic axes to form audio activation mode vectors and visual activation mode vectors, respectively. The Pearson correlation coefficient between the audio activation mode vectors and the visual activation mode vectors is calculated to obtain the cross-axis coordination factor. Each aesthetic semantic axis is assigned a corresponding axis importance weight. Based on the axis importance weight, the cross-modal resonance coefficients of all aesthetic semantic axes in the same pairing are weighted and summed to obtain the basic resonance similarity. According to the basic resonance similarity and the cross-axis coordination factor, the original cross-modal similarity of each pairing is calculated, and the original cross-modal similarity is normalized to obtain the cross-modal similarity.

5. The method for generating traditional Chinese visual styles based on musical features according to claim 4, characterized in that, Methods for determining the optimal visual style gene combination corresponding to each fingerprint vector include: A pre-defined cross-modal pairing rule table records the pairing relationships between each fingerprint type and each gene category; each pairing relationship includes fingerprint type, gene category, applicable pairing marker, and preferred pairing quantity; For each fingerprint vector, based on the corresponding fingerprint type, all gene categories marked as applicable for pairing are obtained from the cross-modal pairing rule table to form a candidate gene category set; all visual style genes whose gene categories belong to the candidate gene category set are screened from the traditional visual style gene library to form a candidate gene set; for each fingerprint vector, the cross-modal similarity between it and each visual style gene in the corresponding candidate gene set is obtained in turn, and visual style genes whose cross-modal similarity is less than a preset cross-modal threshold are removed from the corresponding candidate gene set to obtain an effective candidate gene set; The effective candidate gene set is grouped according to gene category to obtain the category gene group corresponding to each gene category in the effective candidate gene set; for each category gene group, all visual style genes are sorted from largest to smallest according to cross-modal similarity, and the visual style genes with the highest ranking and number not exceeding the corresponding preferred pairing are selected as the preferred genes under the corresponding gene category; the preferred genes under each gene category in the effective candidate gene set corresponding to each fingerprint vector are summarized to form the optimal visual style gene combination corresponding to each fingerprint vector.

6. The method for generating traditional Chinese visual styles based on musical features according to claim 5, characterized in that, Methods for generating style gene expression parameter sets include: For each visual style gene in each optimal visual style gene combination, the corresponding expression excitation intensity is calculated sequentially; the expression excitation intensity of each visual style gene is subjected to nonlinear activation transformation to obtain the corresponding activation expression level; an aesthetic regulation matrix between genes is constructed, and the activation expression level of each visual style gene is iteratively regulated and updated to obtain the steady-state expression level, and the steady-state expression level of each visual style gene is taken as the corresponding expression degree. The optimal visual style gene combination corresponding to all fingerprint vectors is summarized, and all unique visual style genes are extracted to form a global candidate gene set. The global candidate gene set is then grouped according to gene category, resulting in a global gene group for each gene category. For each global gene group, all fingerprint types marked as applicable for the corresponding gene category are obtained from the cross-modal pairing rule table, and the feature saliency for each fingerprint type is acquired. Based on the feature saliency of each global gene group, the standard importance of each gene category is calculated. Based on the steady-state expression level of all visual style genes corresponding to each global gene group, the total expression of each category is calculated, and the relative intensity within each category is determined based on the total expression. The product of the standard importance and the relative intensity within each category is calculated to obtain the expression weight of each visual style gene. For each visual style gene in the global candidate gene set, the corresponding gene category, feature parameter set, expression weight and expression degree are integrated to form a gene expression parameter record; the gene expression parameter records of all visual style genes are summarized to generate a style gene expression parameter set.

7. The method for generating traditional Chinese visual styles based on musical features according to claim 6, characterized in that, Methods for obtaining steady-state expression levels include: The system predefines homology and competition coefficients. For any two different visual style genes in the global candidate gene set, if the unique work codes of the two visual style genes are the same but the gene categories are different, the aesthetic regulation coefficient is set to the homology coefficient. If the gene categories of the two visual style genes are the same but the unique work codes are different, the aesthetic regulation coefficient is set to the negative of the competition coefficient. If the unique work codes and gene categories of the two visual style genes are both different, the aesthetic regulation coefficient is set to zero. The aesthetic regulation coefficients corresponding to the pairwise paired visual style genes in the global candidate gene set are integrated to form an aesthetic regulation matrix. The activation expression levels of all visual style genes are combined to form an expression level vector. A preset control step size, convergence threshold, and maximum number of iterations are established. Iterative control updates are then performed: matrix multiplication is performed on the aesthetic control matrix and the expression level vector to obtain a control influence vector; based on the control step size and the control influence vector, a control increment vector is calculated; the expression level vector and the control increment vector are added element-wise, and the expression level vector is updated based on the calculation results, with the iterative change calculated; if the iterative change is less than the convergence threshold or the maximum number of iterations is reached, the iteration stops, and the values ​​of each element in the expression level vector are taken as the steady-state expression level of the corresponding visual style gene; otherwise, the next iterative update continues.

8. The method for generating traditional Chinese visual styles based on musical features according to claim 7, characterized in that, Methods for generating images in the traditional Chinese visual style include: From the style gene expression parameter set, obtain the gene category, feature parameter group, expression weight, and expression degree corresponding to each gene expression parameter record; for each gene expression parameter record, multiply each parameter value in the corresponding feature parameter group by the corresponding expression degree to obtain the expression scaling parameter group; preset the rendering canvas size and background color value; initialize the rendering canvas according to the rendering canvas size, and set the pixel value of each pixel in the rendering canvas to the background color value; initialize a load margin map with the same size as the rendering canvas, and set the load margin value of each pixel in the load margin map to one; The rendering hierarchy is preset, which specifies the execution order of the rendering layers corresponding to each gene category. The traditional visual style rendering model contains category rendering sub-models corresponding to each gene category. According to the rendering hierarchy, each gene category is used as the current gene category, and layered rendering and load margin adjustment are superimposed. After all gene categories have completed layered rendering and load margin adjustment, the pixel values ​​of all pixels in the rendering canvas are combined to form the final image data, which is then output as a traditional Chinese visual style image.

9. The method for generating traditional Chinese visual styles based on musical features according to claim 8, characterized in that, The methods for performing layered rendering and load capacity adjustment include: For each visual style gene under the current gene category, the corresponding expression scaling parameter set and the rendering canvas size are input into the corresponding category rendering sub-model to generate a single-gene rendering layer; the sum of the expression weights of all visual style genes under the current gene category is calculated to obtain the category superposition intensity; based on the expression weights and category superposition intensity, the category standard weight of each visual style gene is calculated; for each pixel in the rendering canvas, the pixel values ​​of all single-gene rendering layers under the current gene category at the corresponding pixel are weighted and summed based on the category standard weight to obtain the corresponding category fusion pixel value; Obtain the pixel value of each pixel in the rendering canvas and the corresponding pixel's carrying capacity value in the carrying capacity map; calculate the product of the carrying capacity value corresponding to the same pixel and the category stacking intensity to obtain the effective stacking coefficient of each pixel; calculate the stacked pixel value of each pixel based on the effective stacking coefficient and the category fusion pixel value, and update the pixel value of each pixel in the rendering canvas to the corresponding stacked pixel value; calculate the update carrying capacity value of each pixel based on the effective stacking coefficient and the preset level attenuation coefficient; update the carrying capacity value of the corresponding pixel in the carrying capacity map to the update carrying capacity value.

10. A system for generating traditional Chinese visual styles based on musical features, implementing the method for generating traditional Chinese visual styles based on musical features as described in any one of claims 1-9, characterized in that, include: The audioprint encoding module is used to acquire music signals, perform multi-scale time-frequency analysis on the music signals, extract melody contour fingerprints, rhythm pattern fingerprints, audio chromatogram fingerprints and emotion polarity fingerprints respectively, encode each type of fingerprint into corresponding fingerprint vectors and calculate the feature salience in the music signal. The visual base construction module is used to collect traditional Chinese visual art works, and to perform genetic decomposition and parameter encoding of compositional paradigms, brushwork types, pattern motifs and traditional color spectrum in traditional Chinese visual art works, and to construct a traditional visual style gene library containing multiple visual style genes. The modality matching module is used to calculate the cross-modal similarity between each fingerprint vector and each visual style gene in the traditional visual style gene library, and to determine the optimal visual style gene combination corresponding to each fingerprint vector according to the preset matching rules. The gene regulation module is used to dynamically calculate the expression weight and expression level of each visual style gene based on the cross-modal similarity between each fingerprint vector and each visual style gene in the corresponding optimal visual style gene combination, and combined with the feature saliency corresponding to each fingerprint, to generate a set of style gene expression parameters. The style rendering module is used to drive a preset traditional visual style rendering model based on the style gene expression parameter set. It performs parameterized superposition rendering according to the expression weight and degree of each visual style gene to generate Chinese traditional visual style images.