Cross-modal AI-based traditional art gene decoding method and system
By constructing a multimodal dataset and employing technical means, and through the patented improved visual encoder and self-attention mechanism of CLIP, combined with a high-frequency fidelity generation method, the technical problems existing in the prior art have been solved. This has enabled an accurate mapping from the visual features of traditional Chinese painting to the semantics of traditional music, and has resolved the issues of pentatonic scale deviation and loss of cultural semantics. The generated music is superior to existing models in terms of emotional consistency and aesthetic integrity.
Patent Information
- Application Number
- CN202511260311.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for generating Chinese-style music suffer from problems such as high deviation of the pentatonic scale, serious loss of cultural semantics, monotonous traditional timbre, and difficulty in quantifying aesthetic elements, and lack an effective evaluation system.
A multimodal dataset of Chinese painting, music, and text is constructed. A visual encoder improved by CLIP-ViT and a self-attention mechanism are used to map visual tokens to the latent space of music. High-frequency fidelity generative adversarial network is combined to generate traditional Chinese music that conforms to the pentatonic scale. A pentatonic scale constraint matrix and a multi-codebook loss calculation mechanism are introduced.
It achieves a precise mapping from the visual features of traditional Chinese painting to the semantics of traditional music, captures the intrinsic connection between the rhythm of brush and ink and the structure of traditional music, and generates music that is superior to existing models in terms of emotional consistency and aesthetic integrity, with rich timbre features that conform to the timbre features of traditional musical instruments.
Smart Images

Figure CN120998160A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of cross-technology of artificial intelligence and digital media art, and particularly relates to a Chinese art gene decoding method and system based on cross-modal AI. BACKGROUND
[0002] Cross-modal art generation, as a frontier direction in the cross field of artificial intelligence and digital humanities, is significantly promoting the digital protection and innovative transformation of traditional cultural heritage. This technology is committed to realizing the intelligent dialogue between visual art and auditory art, and generating traditional music with a similar artistic conception by decoding the aesthetic genes of Chinese painting "writing spirit with form". It not only opens up a new path for the multi-sensory expression of Chinese aesthetics, but also becomes a touchstone for testing key technologies such as cross-modal representation learning and cultural adaptability generation.
[0003] The current mainstream technology presents three evolution paths: the framework based on text mediation (such as Art2Mus) uses ImageBind to build a unified embedding space, and achieves a user score of 4.2 / 5 in the emotional resonance index, but the semantic deviation rate caused by double conversion reaches 32.7%, resulting in a serious loss of abstract aesthetics such as the "white space artistic conception" of Chinese painting; the end-to-end waveform generation scheme (represented by CPTGZ) uses a latent diffusion model to realize the direct generation of Chinese painting to guzheng music, with a classical flavor retention rate of 89%, but the fidelity of details such as the ink wash technique is only 72%, and the timbre expression is limited to a single instrument; the culture-enhanced architecture (such as the Chinese-style video-music system of Cao Muchi) introduces Stable Audio coding technology, which improves the cultural fit score by 37%, but the emotional consistency decreases by 28% when static Chinese painting is input, revealing the inherent limitations of dynamic information dependence.
[0004] These technology routes are jointly faced with three core constraints: at the data level, there is a lack of directly paired data of Chinese painting-music available globally, and professional annotations such as "ink color density-timbre thickness" are lacking in general data sets (such as AudioSet); at the cultural adaptation level, the deviation of the pentatonic scale of the western basic model (such as MusicGen) when generating Chinese-style music reaches 0.62, making it difficult to capture the internal relationship between the spread rhythm and the ink wash technique; at the evaluation system level, traditional audio indicators (FAD, KL divergence) cannot quantify the cross-modal aesthetic correspondence between "ink and five colors" and "sound and twelve pitches", and it is urgent to establish an evaluation dimension that integrates art theory. SUMMARY
[0005] To solve the above technical problems, the application provides a Chinese art gene decoding method based on cross-modal AI, comprising the following steps:
[0006] Step S1: based on the national painting-music data pair, the national painting-text data pair and the text-music data pair, a national painting-music-text multi-modal data set is constructed;
[0007] Step S2: the national painting image is input into the visual encoder improved based on CLIP-ViT, passes through the normalization module, the position coding module and the Transformer encoder, and outputs the 512-dimensional visual Token sequence;
[0008] Step S3: the visual Token sequence and the emotion label from the user or the multi-modal data set by default Input the cross-modal adapter, adopt the self-attention mechanism to directly map the visual Token to the music hidden space, and obtain the music embedding vector fused with the visual content and the emotion information ;
[0009] Step S4: the , the user parameter is input into the improved high-frequency fidelity generative adversarial network, and the Chinese traditional music audio conforming to the five sound scale is generated.
[0010] Beneficial effects:
[0011] 1. The national art gene decoding method and system based on cross-modal AI provided by the application solve the technical problems that the existing western basic model has high five sound scale deviation and serious cultural semantic loss when generating Chinese music, realize accurate mapping from national painting visual features to traditional music semantics, and effectively capture the internal correlation mechanism of national painting brush rhythm and traditional music scatter structure by constructing a national painting-music-text multi-modal data set, adopting the architecture of deep fusion of CLIP visual encoder and MusicGen audio generation model, and combining adaptive visual conditioner technology.
[0012] 2. The application overcomes the difficulty that abstract aesthetic elements such as “white space artistic conception” and “brush rhythm” in national painting are difficult to quantitatively express by using a hierarchical adaptation mechanism and a multi-codebook loss calculation mechanism, realizes multi-level feature extraction and fusion from local brush strokes to overall composition, and makes the generated music better than the existing end-to-end generation model in terms of emotional consistency and aesthetic integrity.
[0013] 3. The application introduces a five sound scale constraint matrix and a high-frequency fidelity generative adversarial network, solves the problems of single tone color and large non-traditional scale interference in traditional music generation, and generates music that retains the tone color characteristics of traditional instruments such as guqin and pipa at the waveform level, providing a high-quality and interactive intelligent solution for the inheritance of national art in the digital age. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 It is a national art gene decoding method flowchart based on cross-modal AI.
[0015] Figure 2 The overall architecture of the method of the present application is shown in the figure;
[0016] Figure 3 The structure block diagram of a national art gene decoding system based on cross-modal AI according to the present application is shown in the figure. DETAILED DESCRIPTION
[0017] In order to make the objectives, technical solutions and advantages of the present application clearer and more comprehensible, the present application will be further described in detail below in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0018] Embodiment one
[0019] As shown in the figure, the national art gene decoding method based on cross-modal AI provided by the embodiment of the present application comprises the following steps: Figure 1
[0020] Step S1: based on the national painting-music data pair, the national painting-text data pair and the text-music data pair, a multi-modal data set of national painting-music-text is constructed;
[0021] Step S2: input the national painting image into the visual encoder based on the improved CLIP-ViT, pass through the normalization module, the position coding module and the Transformer encoder, and output the 512-dimensional visual Token sequence;
[0022] Step S3: input the visual Token sequence and the emotion label from the user or the multi-modal data set default into the cross-modal adapter, adopt the self-attention mechanism to directly map the visual Token to the music hidden space, and obtain the music embedding vector fused with the visual content and the emotion information
[0023] Step S4: input the , user parameters into the improved high-frequency fidelity generative adversarial network, and generate Chinese traditional music audio conforming to five sound levels.
[0024] In one embodiment, the above step S1: based on the national painting-music data pair, the national painting-text data pair and the text-music data pair, a multi-modal data set of national painting-music-text is constructed, specifically comprising:
[0025] Step S11: Chinese painting-music data pair: key frames are extracted from the relevant videos of Chinese painting appreciation and documentary as Chinese painting images, and the audio segments at the corresponding time points are synchronously extracted; songs labeled with Chinese style and ancient style tags and their album cover images are obtained from a music platform;
[0026] In the embodiment of the application, key frames are extracted from B station related Chinese painting appreciation and documentary videos as Chinese painting images, and audio segments at the corresponding time points are synchronously extracted; at the same time, songs labeled with "Chinese style", "ancient style" and the like and their album cover images are obtained from a cool dog music platform to supplement the source of Chinese painting images. In the embodiment of the application, 326 pairs of Chinese painting-music data pairs are collected;
[0027] Step S12: Chinese painting-text data pair: basic attribute information of Chinese painting works is obtained from authoritative platforms; a large language model and expert auxiliary labeling are used to analyze and label professional aesthetic features of each painting, including artistic conception, color features and ink technique;
[0028] In the embodiment of the application, basic attribute information of Chinese painting works, such as painting name, author, era, school, etc., is obtained from authoritative platforms such as Yachang Art Platform, Palace Famous Paintings, China Art Museum, etc. The Kimi AI large language model and expert auxiliary labeling are used to analyze and label the artistic conception, color features, ink technique and other professional aesthetic features of each painting. In the embodiment of the application, 11,067 pairs of Chinese painting-text data pairs are collected;
[0029] Step S13: text-music data pair: Chinese traditional style music segments with detailed labels are crawled from an audio material platform, and corresponding audio files are synchronously downloaded to establish the association between text description and music features;
[0030] Chinese traditional style music segments with detailed labels are crawled from an audio material platform such as Aigewang, and corresponding audio files are synchronously downloaded. These labels are used to establish the association between text description and music features; in the embodiment of the application, 1,030 pairs of text-music data pairs are collected;
[0031] Step S14: aligning the Chinese painting-music data pair, the Chinese painting-text data pair and the text-music data pair to construct a multi-modal data set of Chinese painting-music-text.
[0032] In one embodiment, step S2 above: input the Chinese painting image into the visual encoder improved based on CLIP-ViT, pass through the normalization module, the position encoding module and the Transformer encoder, and output the 512-dimensional visual Token sequence, which specifically includes:
[0033] Step S21: First, the Chinese painting image is input into the normalization module, and is uniformly scaled to 224x224 pixels. For the low saturation characteristics of ink, the normalization layer is optimized to retain the changes in ink density, lightness and dryness:
[0034]
[0035] wherein, is the original pixel value, is the normalized pixel value, is the data set mean, is the data set standard deviation;
[0036] Step S22: The normalized image is input into the position encoding module, segmented into 16x16 pixel blocks, linearly projected into 768-dimensional pixels, and the learnable position encoding is added to capture the composition rule, so that the model can explicitly learn and utilize the unique spatial layout information of Chinese painting:
[0037]
[0038] wherein, is the weight parameter, is the image width, is the image height, is the position coordinate;
[0039] Step S23: Finally, the image features after position encoding are input into the 12-layer Transformer encoder (attention head number=12) for feature extraction, and the visual Token sequence is obtained
[0040] The Transformer encoder is fine-tuned in stages, first pre-trained on ImageNet, and then fine-tuned on the Chinese painting data set.
[0041] In one embodiment, the above step S3: the visual Token sequence and the emotion label from the user or the multi-modal data set default is input into the cross-modal adapter, and the self-attention mechanism is used to directly map the visual Token to the music hidden space to obtain the music embedding vector fused with the visual content and emotion information , specifically comprising:
[0042] Step S31: The visual Token sequence is projected to the music hidden space through a two-layer MLP and a nonlinear alignment head of LayerNorm to obtain the visual feature vector with semantic alignment capability :
[0043]
[0044] ;
[0045] ;
[0046] wherein, , is a first linear layer weight matrix for linearly transforming the input 512-dimensional visual features and performing activation operation to map to a 1024-dimensional hidden space, is a bias vector, and GELU is a Gaussian Error Linear Unit activation function;
[0047] LayerNorm is a layer normalization operation, and Dropout is a random deactivation operation, with a deactivation probability; the current layer is normalized and randomly deactivated to stabilize training and prevent overfitting;
[0048] is a second linear layer weight matrix for further mapping the 1024-dimensional to a music hidden space dimension , is a MusicGen hidden layer dimension, is a second linear layer bias vector;
[0049] Step S32: converting the emotion label into an emotion feature vector using a multi-layer perception (MLP) and fusing it with :
[0050] ;
[0051] ;
[0052] wherein, is an element-wise multiplication operation, and MaxPool is a maximum pooling operation for extracting key features to extract the most significant information from the fused features, is a music embedding vector fused with visual content and emotion information;
[0053] Step S33: parallel processing of detail features, local features, and global features of the Chinese painting through an attention mechanism, wherein the detail features are low-level features including edge features and texture features; the local features are medium-range semantic units including objects and figures; and the global features are overall semantic information including scene categories; using a cross-modal contrastive loss to optimize alignment, to reduce the distance between relevant visual-music feature pairs and to increase the distance between irrelevant visual-music feature pairs.
[0054] In one embodiment, the above step S4: the , the user parameter is input into the improved high-frequency fidelity generation adversarial network, and the Chinese traditional music audio conforming to the five sound scale is generated, specifically comprising:
[0055] Step S41: the The key-value pair is input into the Transformer decoder cross-attention layer:
[0056] ;
[0057] Wherein, is a query vector, is a current generated audio token hidden state, which is calculated and updated by the Transformer decoder itself, is a learnable weight matrix for projecting the audio hidden state to the query space for attention;
[0058] is a key vector generated by the visual condition, is a learnable weight matrix for projecting the visual condition to the key space;
[0059] is a value vector generated by the visual condition, is another learnable weight matrix responsible for projecting the visual condition to the value space;
[0060] is the dimension of the key vector;
[0061] Step S42: using the multi-codebook strategy (8 codebooks x 1024 tokens), embedding the five sound scale constraint matrix in the output layer, suppressing non-traditional scale token probability, and generating process with as a strong condition guide, a multi-codebook loss function is constructed based on cross entropy Optimize the model parameters, and the loss function is used to calculate the average cross entropy loss on all codebooks, encouraging the model to generate music segments conforming to the five sound law:
[0062] ;
[0063] Wherein, is the number of codebooks, each codebook is responsible for modeling part of the audio semantics, is the effective token set of the th codebook, is the codebook in the batch at time a real token, for input Chinese painting, a visual condition vector;
[0064] Step S43: based on the condition feature generate high-quality audio waveform, get Chinese traditional music audio conforming to five sound scale, and perform timbre mixing and equalization post-processing.
[0065] The present application improves the harmonic structure of traditional musical instruments such as Guqin and Pipa on the high-frequency fidelity generative adversarial network HiFi-GAN, so that it can better model traditional timbre; then use the improved HiFi-GAN, with condition feature, generate high-quality Chinese traditional music audio conforming to five sound scale based on user input parameters (such as control parameters: temperature 0.3-1.0, Top-K / Top-P value). Then perform timbre mixing (balance different instrument timbres) and equalization (adjust frequency response) post-processing, further improve the naturalness and traditional charm of the generated music.
[0066] As shown in Figure 2 , it is the overall architecture schematic diagram of the method of the present application.
[0067] Example two
[0068] As shown in Figure 3 , the present application provides a national art gene decoding system based on cross-modal AI, which comprises the following modules:
[0069] The multi-modal data set construction module 51 is used to construct a multi-modal data set of Chinese painting-music-text based on Chinese painting-music data pairs, Chinese painting-text data pairs and text-music data pairs;
[0070] The visual encoder module 52 is used to input the Chinese painting image into the visual encoder based on the improved CLIP-ViT, pass through the normalization module, the position encoding module and the Transformer encoder, and output the 512-dimensional visual Token sequence;
[0071] The cross-modal adapter module 53 is used to input the visual Token sequence and the emotion label from the user or the multi-modal data set default into the cross-modal adapter, and adopt the self-attention mechanism to directly map the visual Token to the music hidden space, so as to obtain the music embedding vector fusing visual content and emotion information;
[0072] The music generation module 54 is used to input the music embedding vector In the improved high-frequency fidelity generation adversarial network with user parameter input, the generated Chinese traditional music audio conforms to five sound orders.
[0073] In addition, in order to comprehensively verify the semantic fit degree, aesthetic consistency and user satisfaction of the application in the image-music cross-modal generation process, an innovative subjective evaluation mechanism is constructed, which is divided into non-expert evaluation and expert evaluation two modules, forming a double-level, multi-dimensional systematic evaluation framework.
[0074] 1. Non-expert evaluation module
[0075] This module is aimed at the general public, starting from three aspects of emotional perception, visual-auditory consistency and version iteration preference, and setting up the following three sub-tasks:
[0076] (1) Music-text matching degree test: This task aims to test the emotional matching effect between the generated music and the semantic text. The experiment randomly shows 10 Chinese descriptions covering different emotions such as "happy", "bold", "calm", "sad" to users, and each text corresponds to 4 pieces of music (from CoDi, Mumu, MusicGen and the VMM model of the application), a total of 40 pieces of music. Users need to score the emotional fit degree of each "text-music" combination based on their auditory perception on a 1 to 5 Likert scale.
[0077] (2) Image-music matching degree test: To evaluate the aesthetic consistency of the system in the image to music generation path, 5 Chinese paintings are selected, and music is generated by CoDi, Mumu, NVMM (traditional image to text to music method) and the VMM model of the application, a total of 20 pieces of audio. Users need to score the "picture and music whether the artistic conception is consistent" according to their intuitive feeling after watching the picture and listening to the audio on a 1 to 5 Likert scale.
[0078] (3) Image-music version comparison test: This task focuses on the evolution effect of system version. 6 paintings are shown, each with 12 pieces of music generated by V0 version (basic model) and V1 version (introducing text guided optimization), and users choose the version with better subjective feeling from "music overall quality" and "image-music matching degree" after comparing the two versions of generated results.
[0079] 2. Expert evaluation module
[0080] On the basis of the non-expert module, the application further introduces professional aesthetic evaluation standards, and an evaluation team composed of experts from the fields of artificial intelligence art, music theory and Chinese painting appreciation scores the output music from a professional dimension.
[0081] (4) Multi-model pure music generation evaluation: Take 3 typical Chinese painting works as input, and generate 1 piece of music by small model, medium model and large model respectively, a total of 9 audio files. Experts score from 1 to 5 based on the auditory characteristics of the music itself without viewing the image from the following three dimensions:
[0082] Overall music quality (melody fluency, rhythm stability, timbre level, etc.);
[0083] Style richness (whether to reflect traditional music style, multi-instrument interaction, etc.);
[0084] Music correctness (whether the music has no obvious errors or unnatural elements).
[0085] The innovation of this subjective evaluation system lies in: combining non-expert user perception and expert professional evaluation to build a cross-level aesthetic evaluation closed loop; multi-task setting covers semantic matching, visual-auditory translation consistency and version evolution feedback; for the first time, "multi-model pure music quality evaluation" is introduced in the image-music task, which can provide feedback basis for model size selection and training path.
[0086] Through the double-layer subjective evaluation system, the application can not only quantitatively measure the cross-modal generation effect, but also provides an experimental paradigm and practical path for establishing a unified perceptual aesthetic evaluation standard for future cultural adaptability artificial intelligence models.
[0087] A Chinese art gene decoding device based on cross-modal AI, comprising one or more electronic devices, wherein the one or more electronic devices are used to implement a method of decoding Chinese art gene based on cross-modal AI.
[0088] An electronic device comprising: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a method of decoding Chinese art gene based on cross-modal AI.
[0089] A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to implement a method of decoding Chinese art gene based on cross-modal AI.
[0090] A non-transitory computer-readable storage medium having stored thereon a computer program that, when executed by a processor, implements a method of decoding Chinese art gene based on cross-modal AI.
[0091] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Therefore, the application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for decoding national art genes based on cross-modal AI, characterized in that, The method comprises the steps of: Step S1: based on the Chinese painting-music data pair, the Chinese painting-text data pair and the text-music data pair, a multi-modal data set of Chinese painting-music-text is constructed; Step S2: input the Chinese painting image into the visual encoder based on the improved CLIP-ViT, pass through the normalization module, the position encoding module and the Transformer encoder, and output the 512-dimensional visual Token sequence; Step S3: inputting the visual Token sequence and the emotion label from the user or the multi-modal data set by default The input cross-modal adapter adopts a self-attention mechanism to directly map the visual Token to a music hidden space, obtaining a music embedding vector that fuses visual content and emotion information ; Step S4: generating , the user parameter is input into the improved high-frequency fidelity generative adversarial network to generate Chinese traditional music audio conforming to five sound orders. 2.The cross-modal AI-based national art gene decoding method according to claim 1, wherein, The step S1: based on the Chinese painting-music data pair, the Chinese painting-text data pair and the text-music data pair, a multi-modal data set of Chinese painting-music-text is constructed, specifically comprising: Step S11: Chinese painting-music data pair: extract key frames as Chinese painting images from relevant videos of Chinese painting appreciation and documentary, and synchronously extract audio segments at the corresponding time points; obtain songs labeled with Chinese style and ancient style tags and album cover images from a music platform; Step S12: Chinese painting-text data pair: obtain basic attribute information of Chinese painting works from an authoritative platform; use a large language model and expert auxiliary labeling to analyze and label professional aesthetic features of each painting, including: artistic conception, color feature, ink technique; Step S13: text-music data pair: crawl Chinese traditional style music segments with detailed labels from an audio material platform, and synchronously download the corresponding audio files, to establish the association between the text description and the music features; Step S14: align the Chinese painting-music data pair, the Chinese painting-text data pair and the text-music data pair, and construct a multi-modal data set of Chinese painting-music-text. 3.The cross-modal AI-based national art gene decoding method according to claim 2, wherein, The step S2: input the Chinese painting image into the visual encoder based on the improved CLIP-ViT, pass through the normalization module, the position encoding module and the Transformer encoder, and output the 512-dimensional visual Token sequence, specifically comprising: Step S21: first input the Chinese painting image into the normalization module, and uniformly scale it to 224x224 pixels; for the low saturation characteristics of ink, the normalization layer is optimized to retain the changes of ink density, dryness and wetness: ; wherein, is the original pixel value, is the normalized pixel value, is the dataset mean, is the dataset standard deviation; Step S22: input the normalized image into the position encoding module, divide it into 16x16 pixel blocks, linearly project it to 768-dimensional pixels, and add learnable position encoding to capture the composition rule, so that the model can explicitly learn and use the spatial layout information specific to Chinese painting: ; wherein, is a weight parameter, is an image width, is an image height, is a position coordinate; Step S23: Finally, the position-coded image features are fed into a 12-layer Transformer encoder for feature extraction, resulting in a sequence of visual tokens . 4.The cross-modal AI-based national art gene decoding method according to claim 3, wherein, The step S3: the visual Token sequence and the emotion label from the user or the multi-modal data set by default An input cross-modal adapter adopts a self-attention mechanism to directly map the visual Token to a music hidden space, to obtain a music embedding vector that fuses visual content and emotion information , specifically comprising: Step S31: aligning the visual Token sequence through a two-layer MLP and a nonlinear LayerNorm head projecting to a music latent space to obtain a visual feature vector with semantic alignment capability : ; ; ; wherein, , is a first linear layer weight matrix for mapping the input 512-dimensional visual features to a 1024-dimensional hidden space, is a bias vector, and GELU is a Gaussian Error Linear Unit activation function. LayerNorm is a layer normalization operation, Dropout is a random inactivation operation, is the inactivation probability; is a second linear layer weight matrix for mapping the 1024 dimensional further to the music latent space dimension , is a MusicGen hidden layer dimension, is a second linear layer bias vector; Step S32: Use a multilayer perceptron (MLP) to assign sentiment labels Convert to sentiment feature vector and with To merge: ; ; wherein, is an element-wise multiplication operation, MaxPool is a max-pooling operation to extract key features, is a music embedding vector that fuses visual content and emotional information. Step S33: the detail features, local features and global features of the Chinese painting are processed in parallel through the attention mechanism, wherein the detail features are low-level features including edge features and texture features; the local features are semantic units in a medium range, including objects and figures; and the global features are overall semantic information, including scene categories; cross-modal contrast loss is used for optimization to narrow the distance between related visual-music feature pairs and widen the distance between unrelated visual-music feature pairs. 5.The cross-modal AI-based national art gene decoding method according to claim 4, wherein, The step S4 comprises: The method is applied to the improved high-frequency fidelity generative adversarial network, and Chinese traditional music audio conforming to five sound orders is generated. Step S41: inputting the key-value pairs into the Transformer decoder cross-attention layer: inputting the key-value pairs into the Transformer decoder cross-attention layer: ; wherein, is the query vector, is the current generated audio token hidden state, computed and updated by the Transformer decoder itself, is a learnable weight matrix that projects the audio hidden state into the query space for attention; a key vector generated for the visual condition, is a learnable weight matrix used to project the visual condition into the key space; a value vector generated for the visual condition, is another learnable weight matrix responsible for projecting the visual condition into the value space; is the key vector dimension; Step S42: using a multi-codebook strategy, embedding a five-sound scale constraint matrix in the output layer, suppressing non-traditional scale Token probability, guiding in the generation process with strong conditions, building a multi-codebook loss function based on cross entropy as a strong condition Optimize model parameters: ; wherein, is the number of codebooks, each codebook responsible for modeling a part of the audio semantics, is the effective token set for the th codebook, is the batch of codebooks at time , the real token, is the input Chinese painting, is the visual condition vector; Step S43: through the improved high frequency fidelity generation adversarial network, based on Generate high-quality audio waveform, get the Chinese traditional music audio conforming to the five sound scale, and carry out timbre mixing and equalization post-processing.
6. A cross-modal AI-based national art gene decoding system, characterized in that, The method comprises the following modules: A multi-modal data set construction module is used to construct a multi-modal data set of Chinese painting-music-text based on the Chinese painting-music data pair, the Chinese painting-text data pair and the text-music data pair; a visual encoder module, configured to input the Chinese painting image into a visual encoder improved based on CLIP-ViT, pass through a normalization module, a position encoding module and a Transformer encoder, and output a visual Token sequence with a dimension of 512; a cross-modal adapter module to map the visual token sequence to a music embedding vector that incorporates both visual content and sentiment information an input cross-modal adapter that employs a self-attention mechanism to map visual tokens directly to a music latent space ; The music generation module is configured to generate , input the user parameters into an improved high-frequency fidelity generative adversarial network, and generate Chinese traditional music audio conforming to five sound scales. 7.A cross-modal AI-based national art gene decoding apparatus, characterized by, one or more electronic devices configured to implement the method of any one of claims 1 to 5.
8. An electronic device, comprising: comprising: one or more processors; a memory storing one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement the method of any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that, a non-transitory computer-readable medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 5.
10. A non-transitory computer-readable storage medium, comprising: a computer program stored on a computer readable medium, which when executed by a processor, causes the processor to perform the method of any one of claims 1 to 5.
Citation Information
Cited By
Generative semantic communication system for 3D content generation
CN121547151A
A generative semantic communication system for 3D content generation
CN121547151B
Metainformation-driven synthesis training method, system and device for traditional Chinese painting large model and storage medium
CN121685749A