Multi-modal art generation and capitalization method based on emotion recognition
Through a multimodal art generation and assetization system based on emotion recognition, user emotions are transformed into multimodal art expression and digital assets, and the problem of lack of structured output in the existing technology is solved, and users' subjective emotional participation and emotional value conversion ability are improved.
Patent Information
- Application Number
- CN202510644586.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-29
AI Technical Summary
The existing emotion recognition technology lacks structured, multimodal visual, audible, and interactive output, and the user's subjective emotional participation is low, making it difficult to form a personalized link of "inner emotions-external content".
A multimodal art generation and assetization system based on emotion recognition is designed. Through user input modules, emotion recognition and coordinate mapping modules, strategy matching engines, multimodal generation modules, visualization modules and emotional asset packaging modules, user emotions are transformed into multimodal art expressions and digital assets, supporting the generation of images, music, accessories, etc., and optimizing the generation strategy through user feedback.
It realizes multimodal artistic expression and digital assetization of user emotions, improves the freedom of expression and the ability to transform emotional value, and forms personalized content output.
Smart Images

Figure FT_1
Abstract
Description
1. Technical Field
[0001] The present invention relates to the fields of artificial intelligence, human-computer interaction, art generation and digital assets, and in particular to a multimodal art generation and assetization method and system based on emotion recognition, which can be applied to scenarios such as personalized art creation, psychological healing, digital collections and lifestyle recommendations. 2. Background Technology
[0002] Emotion recognition technology is currently being used in areas such as health monitoring and personalized recommendations, but most systems only provide label-based output and lack structured, multimodal visual, audible, and interactive output. Furthermore, while art generation and NFT digital asset technologies are rapidly developing, user engagement with subjective emotions is low, making it difficult to establish a personalized connection between internal emotions and external content. Therefore, there is an urgent need for a multimodal generation and asset system based on emotional coordinates that can transform user emotions into images, music, accessories, and even lifestyle recommendations through structured mapping, thereby enhancing freedom of expression and the ability to transform emotional value. 3. Summary of the Invention
[0003] The present invention provides a system and method for converting a user's multi-source emotional input into multimodal artistic expression and digital assets. The system includes a user input module, an emotion recognition and coordinate mapping module, a strategy matching engine, a multimodal generation module, a visualization module, an emotion asset packaging module, and a user feedback mechanism. The system supports multi-source emotional input such as images, keywords, music, behavior descriptions, expressions, and living environments. After mapping to emotional coordinates, it calls style, color, music, makeup, accessories, and life matching templates to generate personalized picture cards, music clips, art paintings, outfit / makeup / space suggestions, and other content, which can be packaged as NFT assets. The final user behavior data is fed back to the strategy system to achieve a closed loop of content optimization. IV. Description of the Figures
[0004] The accompanying figure is a flow chart of the overall structure of the system described in the present invention, including: ① user input stage, ② emotion recognition and coordinate mapping, ③ generation strategy matching, ④ multimodal generation, ⑤ product visualization, ⑥ emotional asset generation (optional), and ⑦ user feedback closed loop. V. Specific Implementation Methods
[0005] Users can upload a selfie, select emotional keywords, and describe their behavior. The system will identify their primary emotion, such as "calm + joyful," and automatically match it to coordinates, including impressionist painting styles, soft tones, a medium-tempo piano music template, a sketch of a water droplet-shaped accessory, and light-colored, transparent outfit suggestions. These are then displayed as "emotional artworks" within the interface. Users can choose whether to generate these artworks as NFTs for collection, display, or social sharing. The system will record user clicks, collections, and preferred creations, continuously optimizing the generation strategy and emotion matching model.
Claims
1. A multimodal art generation and assetization method based on emotion recognition, characterized by: The steps include: (1) receiving emotion-related information input by the user, including images, keywords, music clips, expression data, behavior descriptions, images of living environments, and the user's favorite colors or paintings. The information can be uploaded by the user or selected from the system database; (2) identifying the main emotion of the above information and mapping it to a position in a preset two-dimensional or polar coordinate system; (3) Matching content strategies including color style, music style template, art style, accessory structure, makeup suggestions, clothing suggestions, space decoration suggestions, and lifestyle accessories matching suggestions in the strategy library according to the coordinate position; (4) outputting at least one multimodal artistic expression such as an emotion card, a music clip, an art painting, a jewelry pattern, a DIY jewelry bag, or a lifestyle suggestion through a multimodal generation module according to the matching strategy; (5) Visually display the above generated content to form a unified emotional work interface; (6) Optionally package the emotional works into digital emotional assets to generate NFTs or other tradable digital forms; (7) Record user behavior preferences, collection behavior and other data, and use them to optimize generation algorithms and recommendation strategies.
2. The method according to claim 1, wherein the emotion coordinate system is a two-dimensional plane or a polar coordinate diagram, and the coordinate axes represent pleasure, arousal or other combinations of emotion dimensions. The method according to claim 1 , wherein the music style template supports a user to specify a composer or a music period.
4. The method according to claim 1, wherein the generated jewelry design pattern can be used for 3D modeling or DIY material package output.
5. The method according to claim 1, wherein the life suggestions include makeup suggestions, clothing suggestions, space decoration styles or life accessories combination plans.
6. The method according to claim 1, wherein the digital emotional assets include NFT images, music, jewelry samples, life advice cards, etc.