Art non-abandoned item inheritance and interactive learning method based on cross-modal generation algorithm

By constructing a multi-source heterogeneous feature map information database and a cross-modal generation algorithm, combined with virtual reality technology, the problems of data correlation and style inaccuracy in art-related intangible cultural heritage projects have been solved, achieving accurate inheritance and innovative design, and improving learning efficiency and cultural authenticity.

CN120994066APending Publication Date: 2025-11-21XIAN PEIHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511131666.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate visual, auditory, and tactile data from intangible cultural heritage projects in the fine arts, resulting in the loss of the connection between patterns and symbols and the cultural semantics behind them. This leads to stylistic inaccuracies in the generated works, and the interactive learning system lacks real-time feedback and adaptive mechanisms, making it difficult to achieve accurate inheritance.

Method used

Construct a multi-source heterogeneous feature map information database, achieve semantic alignment of visual, text and action features through cross-modal generation algorithms, combine virtual reality technology for real-time feedback and correction, develop a two-way interactive learning system, and use the Style-Conditional Diffusion model to regulate the weight of regional culture, so as to realize the authentic transmission of cultural genes and innovative design.

Benefits of technology

It has enabled the precise inheritance and innovative design of intangible cultural heritage projects in the field of fine arts, improved the authenticity of culture, shortened the learning cycle, improved learning efficiency, and supported the rapid iteration of the cultural and creative industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994066A_ABST
    Figure CN120994066A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal generation algorithm-based inheritance and interactive learning method for art non-perpetual items, and the method comprises the steps: collecting the information data of art non-perpetual items, forming multi-source heterogeneous data, carrying out the cleaning, classification and labeling of the data, forming a semantic-associated data set, and carrying out the recognition of the multi-source heterogeneous data. The data are classified into a static pattern library, a dynamic process library and a context knowledge library, the three information libraries form a standard multi-source heterogeneous characteristic spectrum information library of non-perpetual items, and non-perpetual characteristics of different modalities are fused by using an improved cross-style generation algorithm so as to form a standard multi-source heterogeneous characteristic spectrum information library of the non-perpetual items; generating a plurality of scheme works with diversified styles, which not only can retain traditional culture genes, but also can adjust the innovation degree; and learning and design reconstruction are carried out by using virtual-real fusion learning design and a cross-style generation algorithm, and fusion learning and creation of artistic contents are carried out by combining historical creation data of the user. The method disclosed by the invention has important practical significance for diversified presentation and inheritance learning of non-perpetual culture.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of cultural heritage learning, and specifically relates to a method for inheriting and interactive learning of art non-heritage projects based on a cross-modal generation algorithm. BACKGROUND

[0002] At present, the digital inheritance of art non-heritage still faces three dilemmas: (1) In the aspect of preservation, traditional technologies mainly focus on the two-dimensional scanning and storage of pattern designs. Existing data are mainly in a single mode (such as images or texts), which only records the information of inheritors and simple patterns (styles, colors), without correlating and integrating the data of visual (work pattern), auditory (instructor explanation) and tactile (process operation feedback), and without comprehensively recording and preserving the stories and origins of patterns and the cultural semantics contained therein, resulting in the loss or failure of matching the relevance of pattern symbols and the folk narratives and regional resources behind them; (2) In the aspect of creation, existing art generation models (such as GAN and Diffusion Model) have not been deeply bound with regional cultural feature libraries, and their style migration is easily disturbed by general data sets, resulting in the phenomenon of "style misalignment" such as pattern structure dislocation and color symbol misreading in generated works, which is difficult to meet the core demand of faithful transmission of non-heritage cultural genes; (3) In the aspect of inheritance, existing interactive learning systems mainly adopt a one-way content output mode, lack of analysis of brush stroke mechanics, detection of technique compliance and self-adaptive feedback mechanism for user creation track, making it difficult to effectively transfer the implicit knowledge of "correspondence between mind and hand" in traditional crafts. In order to realize accurate inheritance, it is necessary to convert the process details into quantifiable and traceable digital assets through sensors and multi-modal data collection technology, and to establish a quantitative standard system of cultural genes.

[0003] In view of the above problems, the present technology builds a multi-source heterogeneous feature map information library, aligns the visual features such as patterns with multi-modal data such as oral texts and process video, and constructs a closed-loop transmission path of virtual and real fusion, in order to solve the bottleneck of implicit experience transmission of art non-heritage. And the improved Style-Conditional Diffusion model is used to realize the accurate regulation of regional cultural weight coefficient, to ensure that the generated works meet the needs of artistic innovation and strictly abide by the boundaries of cultural genes on the basis of non-heritage inheritance protection; at the same time, a two-way interactive learning system is developed, which realizes real-time correction and adaptive inheritance of technique track through real-time brush stroke track analysis and style gradient feedback, including cross-modal feedback mechanism, to build an immersive learning closed loop of "observation-practice-correction".

[0004] Taking woodblock New Year pictures as an example, the text description of the "picture draft" of the woodblock New Year pictures is not corresponded with the video of the actual carving track, so that the learners cannot understand the conversion logic from "draft design" to "carving practice". The function of the present application is to build a multi-source feature map database for the intangible heritage of woodblock New Year pictures, which is divided into two sub-databases of Tianjin Yangliuqing and Suzhou Taohuawu, and each sub-database contains accurate pattern, picture, inheritance population text, process production video, interview, folk customs and other multi-source data information. When people want to watch and learn, the pattern elements (such as the dress of the ladies in Yangliuqing New Year pictures and the auspicious pattern of Taohuawu New Year pictures) and color composition in the image are extracted, and the semantic information of the related text mode (such as the picture title, historical anecdotes and folk meaning) and audio / video mode (such as the production process explanation and cultural background explanation) is analyzed. With the help of a deep learning model, the image features and text features are mapped to a unified semantic space, so that the "bat pattern" is associated with the semantics of "good luck" in the text and "harmonious symbol" in the audio, and finally the corresponding relationship of the mutual mapping is formed under the semantic concepts of "auspicious culture" and "folk symbol". In simple terms, the picture (such as the bat and peony patterns in the picture), the text description (such as "the bat represents good luck" and "the peony symbolizes wealth and honor"), and the audio explanation (such as the introduction of the meaning of New Year pictures) are "translated" into the same "language channel" by technical means, so that the meaning of the picture and the meaning of the text and the audio can be matched. By looking at the picture, you can understand what the text is talking about, and by listening to the sound, you can understand what the picture is about. In this way, the meaning of different forms of information is matched, and the demand for comprehensive deep learning is met.

[0005] With the help of the improved Style-Conditional Diffusion model, the woodblock New Year pictures can retain tradition while playing new tricks. By adjusting the weight coefficients and characteristic proportions of different regional cultures of New Year pictures in different regions (such as Yangliuqing and Taohuawu), new works can be created with new creative ideas while retaining the original cultural flavor of New Year pictures. At the same time, the interactive learning system acts as an "intelligent teacher" that can quickly analyze the learner's technical problems by capturing the stroke track in real time during carving, combining style gradient feedback with image, tactile and other cross-modal feedback. For example, if the angle of the knife is found to be incorrect or the force is found to be inappropriate, the system will immediately remind you to correct it through picture prompts and touch adjustments, so that you can learn and improve at the same time, quickly master the traditional carving techniques, and help the woodblock New Year pictures to be better inherited.

[0006] The system is highly integrated and extracted by computer science technology, which not only solves the key problems of cultural context rupture and style expression disorder in digitalization of intangible cultural heritage, but also converts static protection into dynamic inheritance through generation algorithm technology, helps to save and learn accurate artistic materials, provides controllable innovation space for the contemporary transformation of traditional crafts, and has important practical significance for the inheritance and interactive learning of art intangible cultural heritage. SUMMARY

[0007] The purpose of the present application is to provide an art intangible cultural heritage project inheritance and interactive learning method based on a cross-modal generation algorithm, which solves the problem of inaccurate art intangible cultural heritage artistic inheritance and the limitation of intangible cultural heritage interactive appreciation and learning.

[0008] The technical solution adopted by the present application is an art intangible cultural heritage project inheritance and interactive learning method based on a cross-modal generation algorithm, and the specific operation steps are as follows: Step 1: Collect information data of art intangible cultural heritage project embroidery, including different types of pattern images, process operation videos, inheritance population audio, regional folk text, etc., form multi-source heterogeneous data of intangible cultural heritage objects, and clean the data; Step 2: Classify and label the effective data formed in step 1 to form a data set associated with semantics, and according to the common retrieval standard, the data of the intangible cultural heritage project is layered according to the inheritance difficulty, pattern type, cultural implication and process technique requirement, and the different data types obtained are converted into a unified format; Step 3: The data processed in step 2 is subjected to secondary screening and analysis and is classified into three kinds of atlas information library, namely static pattern library, dynamic process library and context knowledge base, that is, the standard multi-source heterogeneous feature atlas information library of the intangible cultural heritage project; the data in each library represents different modalities; Step 4: When the intangible cultural heritage project needs to be displayed in a cultural creative way, an improved cross-style generation algorithm can be used to fuse the intangible cultural heritage features in step 3 modalities, based on the standard database to develop commercial design and generate a number of diversified style scheme works that can both preserve traditional cultural genes and adjust the degree of innovation; Step 5: Using VR, AR and other virtual-real fusion technologies, learners can learn and design reconstruction. On the one hand, the acquired implicit experience modalities of the inheritors are converted into explicit quantifiable data, and the learner's action process in the physical canvas space coordinates is real-time mapped with the standard virtual limb movement of the inheritors, helping the learner to correct cognitive and practical operation errors in time; on the other hand, after sufficient virtual experience, the learner can carry out fusion learning and automatic creation of intangible cultural heritage art content, such as derivative design of intangible cultural heritage project patterns, to reconstruct highly personalized design works.

[0009] The application also has the characteristics that Step 1: obtaining more than 1 TB of multi-source heterogeneous data by a multi-sensor acquisition device, and forming a semantic association data set by cultural experts in collaboration with labeling; the multi-sensor acquisition device includes a high-definition scanner, a motion capture system, and an environmental recording device; The high-definition scanner has a resolution of ≥600DPI and is used to collect two-dimensional images of representative patterns in non-heritage projects; The motion capture system uses a Vicon optical motion capture device; the hand movement trajectory of the inheritor during production is recorded, wherein the sampling frequency is ≥120Hz; The sampling rate of the environmental recording device is 44.1kHz, and the time length is ≥50h.

[0010] Step 2 is as follows: The data after cleaning in step 1 is labeled: the pattern image is labeled with structural feature codes, i.e. PSD / EPS / JPG values, and color codes, i.e. HSV / RGB values; the process operation video is labeled with key action nodes and time sequence relationships, i.e. MP4 / MOV values; the inheritor's spoken audio is processed by word segmentation, and is labeled with dialects, professional terms, experiential descriptions, and process details, i.e. WAV values; the regional folk custom text is classified and labeled as customs, songs, and historical records; the above contents are associated and mapped with terms-patterns-process, image noise points are processed by a denoising algorithm, action trajectory frequency domain features are extracted by Fourier transform, spoken audio is transcribed into text using speech recognition technology, and all are stored in JSON format to form a semantic association data set.

[0011] When recording the hand and arm movement trajectory of the inheritor during production in the process video, an optical motion capture instrument is used to record the position change and orientation data of the inheritor's ten fingers, two wrists, and arms; the obtained data is processed and divided in time and space to obtain the operation technique data of the art category non-heritage project, which serves as a standard technique action database of the inheritor. Step 4 is as follows: Full use is made of the complementarity between different modalities to integrate information extracted from different modalities into a stable cross-modal representation, and a more accurate standard fitting dynamic equation is obtained:

[0012] Wherein, Z is a cross-modal fusion representation library, 、 、 is a learnable cross-modal projection matrix; represents a matrix group formed by multiple combinations among vision, text, and action; v is a visual feature vector; t is a text feature vector; is an action feature vector; The cross-modal style controllable generation algorithm is used for generating works that fuse non-heritage cross-modal features and style adjustable, that is, after fusing non-heritage features of different modalities, generating diversified style works that retain traditional cultural genes and can adjust the degree of innovation. The algorithm is based on the Style-Conditional Diffusion diffusion model architecture, and through model iteration logic and non-heritage scene fusion derivation, the generated non-heritage cross-modal style controllable diffusion update equation is as follows:

[0013] Traditional cultural feature encoding: ; Innovation guiding feature encoding: ; Wherein, represents n the noise image input at time t; is the image sampled in accordance with the normal distribution during the generation process; is the noise coefficient less than 1 in the diffusion process; is the innovation adjustment parameter, ; the parameter directly controls the balance between cultural preservation and innovation degree; is the noise prediction of the preserved traditional cultural features, is the noise prediction of the introduced innovation elements, and two independent noise predictors are used to process cultural features and innovation guidance respectively; is the process of converting non-heritage traditional cultural elements into a numerical vector, is the process of converting user innovation intention and modern elements into a vector.

[0014] The cross-style generation algorithm weight adjustment formula of step 4 is as follows:

[0015] Wherein, is the input parameter, when λ=1, it completely depends on the database feature generation, when λ=0.1, it allows to introduce 5%-10% of innovation disturbance; y is the final generated feature symbol vector group, is the cultural gene vector of the basic pattern, is the innovation style representation vector Diffusion (Z, W) output finally based on the non-heritage cross-modal style controllable generation algorithm, through diffusion model iteration, and the introduction of innovation elements, Taking embroidery pattern inheritance as an example, is the innovation factor of non-heritage activation, and the traditional pattern (such as ancient painting / old embroidery piece features) is the "foundation", which is adjusted through λ and Fusion, mixed new ideas (such as cyberpunk color, 3D texture) generated "new non heritage style seeds", in the development and design can control the innovation degree of cultural and creative (small lambda close to traditional, large lambda more like "new national tide").

[0016] Step 5 is as follows: Step 5.1: use the wearable device to enter the standard multi-source heterogeneous feature spectrum information library of the non heritage project obtained in step 1 to learn, the learner selects to learn related non heritage knowledge and skill action, and establishes the learner's cognition through database quantitative learning; Step 5.2: the wearable device system compares with the standard action in real time; If the learning action is different from the standard action in the action learning, the wearable device system reminds to correct the learner to adjust repeatedly until it is consistent with the standard; Step 5.3: select the related type data sample in the standard multi-source heterogeneous feature spectrum information library of the non heritage project, and the learner can further cultural creative design; The cultural creative design includes selecting and combining related patterns to generate works with accurate communication significance; Step 5.4: the cross style generation algorithm weight adjustment equation is used to correct the content of the creation work, such as the application error of dynasty, pattern and technique in the work, and the correction process is synchronized in the interactive learning stage, the purpose is to make it conform to the accurate expression of non heritage project, if the creation work generated after adjustment conforms to the characteristics of the learned non heritage, the work is output, if it does not conform, continue to adjust the algorithm weight until the work that conforms to the characteristics of inheritance is generated; The wearable device includes: AR glasses (worn on the head, can see virtual patterns and prompts), tactile gloves (worn on the hands, will have vibration or resistance when touching virtual objects), electronic paintbrush (can sense the force and angle of drawing), sensing canvas (a layer of sensor under the canvas, can know the specific position and force of the learner on the canvas).

[0017] The learning process is as follows: first, through the previously constructed standard multi-source heterogeneous feature spectrum information library of non heritage project, the learner selects to learn related non heritage knowledge and standard action, and establishes the learner's cognition through database quantitative learning, such as when the master paper cutting, the finger turns 15 degrees, the system sets this angle as "correct"; Then, when the learner holds the electronic paintbrush to draw on the sensing canvas, the system will compare with "standard action" in real time; Finally, in the action learning, if the learner appears the force is not right, such as too forceful when carving wood board, the gloves will vibrate to remind; When the angle is not right, the AR glasses will draw a correct virtual track line in front of the learner, so that the learner adjusts repeatedly until it is consistent with the standard.

[0018] Creation process: The AR glasses will bring up a database selection, showing many intangible cultural heritage related patterns, such as peony patterns and auspicious cloud patterns. Learners can use gestures to zoom in, zoom out, or rotate to select; they can also select a basic pattern and use gestures to pinch and pull the pattern, such as pinching the petals to make them more pointed. Under the fusion of virtual and real elements, they can select the system, output commands, and mix the modified pattern with the original intangible cultural heritage pattern, such as combining peony flowers with modern geometric shapes. Learners can create works at will, and save them when satisfied, or export them as images or 3D models.

[0019] The beneficial effects of this invention are: This invention presents a method for the inheritance and interactive learning of intangible cultural heritage projects in the arts, based on cross-modal generation algorithms. Through the deep integration of cross-modal generation algorithm technology with the principles of intangible cultural heritage art, and leveraging an innovative system of "cultural gene analysis - intelligent creation assistance - virtual-real interactive inheritance," it can provide a breakthrough solution for current intangible cultural heritage projects in the arts. Based on a multi-source heterogeneous feature map generation algorithm, it ensures the digital fidelity of core elements of traditional techniques and supports the compliant recombination and innovation of intangible cultural heritage elements, generating thousands of culturally accurate design schemes daily, providing a sustainable content supply for the cultural and creative industries. The virtual-real fusion learning system, through AR dynamic demonstrations, tactile feedback, and real-time trajectory correction, reduces the mastery period of complex processes and breaks through the geographical limitations of the traditional apprenticeship system, enabling large-scale inheritance across generations and regions. This has significant practical implications for the diversified presentation and inheritance learning of intangible cultural heritage. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the process of the inheritance and interactive learning method for intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithm of the present invention; Figure 2 This is a test sample diagram of the motion capture technology for the inheritance and interactive learning method of art-related intangible cultural heritage projects based on cross-modal generation algorithm of this invention; Figure 3 This invention is a three-dimensional evaluation matrix table for the inheritance and interactive learning method of art-related intangible cultural heritage projects based on cross-modal generation algorithms; Figure 4 This is a schematic diagram of the cross-modal feature space projection matrix of the method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithm of the present invention; Figure 5 This is a schematic diagram of the architecture of the intangible cultural heritage core feature mapping library for the inheritance and interactive learning method of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithm of this invention; Figure 6 This is a data reconstruction model diagram of the present invention, which is based on a cross-modal generation algorithm for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts. Detailed Implementation

[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0022] Example 1 This invention relates to a method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms, such as... Figure 1 As shown, this method is an innovative approach that integrates the construction of a multi-source heterogeneous feature map information database of intangible cultural heritage, the development of a cross-style generation algorithm for cross-modal feature fusion, an interactive learning method, and virtual reality technology. In terms of constructing the multi-source heterogeneous feature map information database of intangible cultural heritage: (1) First, when collecting cross-modal data, the selected type is determined, and then two-dimensional images of more than 10 representative patterns of the type (such as animal and plant totems and geometric patterns) are collected by a high-definition scanner (resolution ≥ 600 DPI); at the same time, the audio of oral explanation of the craft is collected by an environmental recording device (sampling rate 44.1 kHz, duration ≥ 50 hours) to form the original dataset (total ≥ 1TB). A motion capture system (such as Vicon optical motion capture equipment) was used to record the hand movement trajectory of the inheritor during the production process (sampling frequency ≥120Hz). During the operation, the positional changes and orientation of the upper limbs such as fingers, wrists, forearms, and upper arms were recorded to test human motion parameters. As shown in Figure 2, a total of 18 test points were set for the three joints of the experimental sample person, including the left and right palms and finger joints, including the shoulder, elbow, and wrist. The experimenter was asked to perform the corresponding operation according to the craft technique, and the test points measured the data at each position. Then, based on the test point data, a dynamic trajectory model of the experimenter's continuous single-arm movement was drawn. Any abnormal or missing points, i.e. invalid data, were linearly reconstructed and cleaned in the later stage, thus obtaining very accurate and effective test point data. The obtained test data with measurement errors was filled and modified, and a new data model was reconstructed, including special finger movements, arm activities, and dynamic three-dimensional models. The obtained spatiotemporal data was analyzed and divided to obtain a database of the physical movements of the characters in the art techniques of intangible cultural heritage projects. The main parameters of this database include age, gender, height, finger posture and special habits, the size of the arm extension space and the size of the movement. According to these selection criteria, the above basic elements were selected and the parameters were recorded. Then, other parameters were measured and statistically analyzed for this group of people using motion capture and other devices. (2) For example Figure 3As shown, a three-dimensional evaluation matrix table was constructed. In the data processing after obtaining the samples, cultural experts and technicians collaborated to complete cross-modal semantic annotation. Color codes (HSV / RGB values) and structural features (such as the number of petals and the curvature of lines) were annotated for the selected pattern images; key action nodes (such as "starting stroke - moving stroke - ending stroke") and temporal relationships were annotated for the technique videos; word segmentation was performed on the spoken text to establish a "term-pattern-craft" association mapping (such as "outlining" corresponding to the drawing technique of a specific descriptive area); image noise was processed by denoising algorithms (such as median filtering); frequency domain features of action trajectories were extracted using Fourier transform; and audio was transcribed into text using speech recognition technology (such as Baidu Voice API). The data was uniformly stored in JSON format to form a semantic association dataset.

[0023] Based on motion capture technology, the experimenter completed the same operation five times. Using SPASS software for data processing, the test data was organized and compiled to obtain accurate and effective spatiotemporal data during the test. After removing invalid experimental data points, the spatiotemporal data was stored in C3D and RPD formats, and then imported into Motion Builder software to output BVH / TXT format files. Different variable parameters were set as needed to ensure the singularity of influencing factors, and their changing patterns were analyzed to form a simulation model using computer technology. The test point model database was then refined and corrected. Finally, according to the type, technique, and process of intangible cultural heritage in the arts, different data areas were divided, and a complete multi-source heterogeneous feature map information database of intangible cultural heritage was constructed with the help of blockchain technology.

[0024] Developing a cross-modal feature fusion cross-style generation algorithm involves fusing visual features (pattern images), textual features (craft descriptions), and action features (technique trajectories) of intangible cultural heritage projects in the field of fine arts across modalities. This achieves a controllable balance between "cultural authenticity" and "innovative design," and solves problems such as "style inaccuracy" and "cultural context break" in existing models.

[0025] Cross-modal feature fusion in fine arts intangible cultural heritage projects involves converting different data (such as text, images, and actions) into a form that computers can understand and process. This is based on cross-modal alignment mapping, which maps these feature data to a unified vector space. Specifically, it constructs a core feature mapping library for intangible cultural heritage through cross-modal alignment technology, achieving vectorized association between pattern, action, and text features. Figure 4 As shown, it can fully utilize the complementarity between different modes, integrating information extracted from different modes into a stable cross-modal representation. This yields a more accurate and standardized fitting dynamic equation:

[0026] Where Z is the cross-modal fusion representation library, Wv, Wt, and Wa are learnable cross-modal projection matrices (representing matrix groups formed by multiple combinations of visual, text, and action elements, respectively); v is the visual feature vector (such as various image features extracted by ResNet, such as texture and composition; statistical features of traditional color systems, such as dominant color distribution and embedding representation of color symbolism); t is the text feature vector (such as word vectors generated by BERT, such as the structure, symmetry, and repetition patterns of traditional patterns, extracted through specially designed convolutional kernels or geometric features); and a is the action feature vector (such as trajectory features extracted by LSTM, such as dynamic information like gestures and techniques).

[0027] After data construction, the cross-modal style-controllable generation algorithm is used to generate works that integrate cross-modal features of intangible cultural heritage and have adjustable styles. This involves fusing features from different modalities of intangible cultural heritage to generate diverse styles that retain traditional cultural genes while allowing for adjustable levels of innovation. The algorithm is built on a Style-Conditional Diffusion model architecture. Through model iteration logic and intangible cultural heritage scene fusion derivation, the generated cross-modal style-controllable diffusion update equation for intangible cultural heritage is as follows:

[0028] Traditional cultural feature coding: ; Innovation-guided feature coding: ; in, express n The noisy image input at any given time; These are images obtained by sampling according to a normal distribution during the generation process; It is the noise figure of the diffusion process, which is less than 1. It is an innovative adjustment parameter. ;pass The parameters directly control the trade-off between cultural preservation and innovation. It is a noise prediction method that preserves traditional cultural characteristics. It is a noise prediction that incorporates innovative elements, using two independent noise predictors to handle cultural characteristics and innovation guidance respectively; It is the process of converting intangible cultural heritage elements into numerical vectors. It is the process of converting user innovative intentions and modern elements into vectors. This formula can be directly used as the mathematical basis for algorithm implementation, while maintaining sufficient expressiveness to handle the needs of cross-modal integration and stylistic innovation of intangible cultural heritage.

[0029] Based on the above model, the actual needs of "cultural authenticity" and "innovative integration" are balanced by the input parameter λ. For example, if the user inputs text descriptions such as "modern minimalist style" or "incorporating technological elements", we use style equations to express the user's intent while preserving the semantics of traditional culture.

[0030] The formula for adjusting style weights across styles is as follows:

[0031] in, (when When it relies entirely on database feature generation (fidelity mode), It allows for the introduction of 5%-10% innovation perturbation (fusion mode). y is the final generated feature symbol vector set. It is the cultural gene vector of the basic pattern (from the database). The final output is the innovative style representation vector Diffusion(Z, W). The innovation vector Diffusion(Z, W) is generated based on the above diffusion model. Based on the controllable cross-modal style generation algorithm of intangible cultural heritage, it incorporates innovative elements through diffusion model iteration, and finally outputs the innovative style representation vector Diffusion(Z, W). Taking the inheritance of embroidery patterns as an example, These are the innovative elements for revitalizing intangible cultural heritage, including traditional patterns. (Such as the characteristics of ancient paintings / old embroidery pieces) serve as the "base," adjusted and... By integrating and blending new creative ideas (such as cyberpunk color schemes and 3D textures) to generate "seeds of new intangible cultural heritage styles," the level of innovation in cultural and creative products can be controlled during the development and design process (by adjusting the level of innovation). Approaching tradition, increase (Approaching the new national trend).

[0032] Through the above methods, this invention provides a method for the inheritance and interactive learning of intangible cultural heritage projects in the arts based on cross-modal generation algorithms. It captures and records the upper limb movements of selected intangible cultural heritage inheritors, obtaining processed dynamic data chains. By recording multiple data contents from visual, textual, and action perspectives, it forms a unified cultural gene knowledge graph with cross-modal fusion. Through improved cross-style algorithms, it ultimately integrates the data flow and constructs the most accurate and standardized multi-source heterogeneous feature graph information database for the inheritance project. This facilitates the improvement of the accuracy of the long-term cultural inheritance of intangible cultural heritage projects in the arts. It also allows for experiential learning and interaction with the generation algorithm through specific indexing and reading. Furthermore, in-depth professional research explores the essence of the inherited culture, helping to clarify the connotation of the inheritance of intangible cultural heritage projects in the arts, which has important practical guiding significance for the cultural dissemination value they carry.

[0033] Example 2 The present invention provides a method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms. This method includes constructing a multi-source heterogeneous feature map information database of intangible cultural heritage, developing a cross-style generation algorithm for cross-modal feature fusion, and integrating interactive learning and virtual reality technology for the integrated learning and creation of artistic content. Specifically: Step 1: Collect information and data on intangible cultural heritage projects in the field of fine arts, including different types of pattern images, videos of craft operations, audio recordings of inheritors' oral accounts, and regional folk custom texts, to form multi-source heterogeneous data of intangible cultural heritage objects, and clean the data; Step 2: Classify and label the data generated in Step 1 to form a semantically related dataset. Then, stratify the data of the intangible cultural heritage project according to the difficulty of inheritance, the types of patterns, the cultural connotations, and the requirements of craftsmanship techniques. Finally, convert the different data types obtained into a unified format. Step 3: The data processed in Step 2 is further filtered and analyzed, and classified into the static pattern library, dynamic process library, and contextual knowledge library. These three types of graph information libraries constitute the standard multi-source heterogeneous feature graph information library of this intangible cultural heritage project. The data in each library represents different modalities. Step 4: Using an improved cross-style generation algorithm, the intangible cultural heritage features in different modalities from Step 3 are fused to generate several diverse style schemes that can both preserve the genes of traditional culture and adjust the degree of innovation. Step 5: Utilize virtual-real fusion learning design and cross-style generation algorithms for learning and design reconstruction, and combine user historical creation data for the fusion learning and creation of artistic content.

[0034] Example 3 Based on Example 2, in step 1, at least 1TB of multi-source heterogeneous data is acquired through a multi-sensor acquisition device, and a semantic association dataset is formed by collaborative annotation by cultural experts; the multi-sensor acquisition device includes a high-definition scanner, a motion capture system, and an environmental recording device. The high-definition scanner has a resolution of ≥600 DPI and is used to collect two-dimensional images of representative patterns from intangible cultural heritage projects. The motion capture system uses Vicon optical motion capture equipment to record the hand movements of the inheritor during the production process, with a sampling frequency of ≥120Hz. The environmental recording device has a sampling rate of 44.1 kHz and a duration of ≥50 h.

[0035] The cleaned data from step 1 is annotated as follows: Pattern images are annotated with structural feature codes (PSD / EPS / JPG values) and color codes (HSV / RGB values); Key action nodes and temporal relationships in the process operation videos are annotated with MP4 / MOV values; Oral audio recordings of inheritors are segmented and annotated with dialects, professional terms, experiential descriptions, and process details (WAV values); Regional folk texts are categorized and labeled as customs, folk songs, and unofficial histories; A terminology-pattern-process association mapping is established for the above content; image noise is processed using denoising algorithms; Fourier transform is used to extract frequency domain features of action trajectories; and speech recognition technology is used to transcribe the oral audio into text, which is then uniformly stored in JSON format to form a semantically related dataset.

[0036] Example 4 Based on Example 3, step 4 is as follows: By fully leveraging the complementarity between different modalities, information extracted from different modalities is integrated into a stable cross-modal representation, resulting in a more accurate and standardized fitting dynamic equation:

[0037] Where Z is the cross-modal fusion representation library. , , It is a learnable cross-modal projection matrix; each represents a matrix group formed by multiple combinations of visual, text, and action elements; v is the visual feature vector; t It is a text feature vector; It is an action feature vector; By fusing intangible cultural heritage features from different modalities, a cross-style generation algorithm is developed to generate diverse works that retain traditional cultural genes while allowing for adjustable levels of innovation. The improved cross-style generation algorithm's diffusion equation is as follows:

[0038] Traditional cultural feature coding: ; Innovation-guided feature coding: ; in, express n The noisy image input at any given time; These are images obtained by sampling according to a normal distribution during the generation process; It is the noise figure of the diffusion process, which is less than 1. It is an innovative adjustment parameter. ;pass The parameters directly control the trade-off between cultural preservation and innovation. It is a noise prediction method that preserves traditional cultural characteristics. It is a noise prediction that incorporates innovative elements, using two independent noise predictors to handle cultural characteristics and innovation guidance respectively; It is the process of converting intangible cultural heritage elements into numerical vectors. It is the process of converting users' innovative intentions and modern elements into vectors.

[0039] The weight adjustment formula for the cross-style generation algorithm in step 4 is as follows:

[0040] in, For input parameters, ,when It relies entirely on database features for generation, when It allows for the introduction of 5%-10% innovation perturbation; y is the final generated feature symbol vector set. It is the cultural gene vector of the basic pattern. It is the final output vector representing the innovative style; like Figures 5-6 As shown, step 5 is as follows: Step 5.1: Using wearable devices, learners access the multi-source heterogeneous feature map information database of the intangible cultural heritage project obtained in Step 1. Learners select to learn relevant intangible cultural heritage knowledge and skills. Through quantitative learning in the database, learners establish their cognition. Step 5.2: The wearable device system compares the learned movement with the standard movement in real time; if a difference is found between the learned movement and the standard movement during the movement learning process, the wearable device system will remind the learner to make corrections and then make repeated adjustments until it matches the standard. Step 5.3: Select relevant data samples from the standard multi-source heterogeneous feature map information database obtained for this intangible cultural heritage project, and learners can then carry out cultural and creative design; the cultural and creative design includes selecting relevant patterns to combine and generate works with accurate dissemination significance; Step 5.4: Correct the content of the created works by adjusting the weight equation of the cross-style generation algorithm. For example, if there are errors in the application of dynasties, patterns, techniques, etc. in the works, the correction process is synchronized in the interactive learning stage. The purpose is to make it conform to the accurate expression of the intangible cultural heritage project. If the created works generated after adjustment conform to the characteristics of the learned intangible cultural heritage, the works are output. If they do not conform, the algorithm weights are adjusted multiple times until a works that conform to the characteristics of inheritance are generated. Step 5.5: If the learner is satisfied with the generated work, output the work; if not, repeat step 5.3 until satisfied.

[0041] Example 5 To verify the feasibility of this invention in practice, it was applied to the national intangible cultural heritage "Miao paper-cutting art," with the Miao paper-cutting project in southeastern Guizhou Province selected as a pilot case. This art suffers from several problems: a disconnect or loss of folk narratives (such as the fertility symbolism of the "Butterfly Mother" totem) from the techniques; structural misalignments in generated patterns (such as a 30% deviation in the number of petals); misinterpretations of colors (such as red, symbolizing celebration, being mistakenly used for funerals); a six-month apprenticeship system requiring mastery of the basics; and a 40% dropout rate among cross-regional learners due to a lack of real-time feedback. It also faces issues such as fragmented digitalization, stylistic inaccuracies, and low transmission efficiency. This embodiment uses a real-world scenario, combined with specific data, experimental comparisons, and in-depth analysis, to demonstrate the effectiveness of this invention in data fidelity, innovative generation, and interactive learning. The implementation period was six months, involving 10 inheritors, 50 learners, and a technical support team.

[0042] The steps and methods of this implementation case are as follows: (1) Based on the cross-modal data acquisition module: 10 representative patterns of Miao paper-cutting (such as "butterfly totem" and "geometric meander") were selected, and 2,000 different visual images were collected using a high-definition scanner to form a static pattern library; the oral process was recorded synchronously (sampling rate 44.1kHz, total duration 60 hours) to record the folk context, covering folk stories and key points of techniques; the upper limb operation of 5 inheritors was recorded using a motion capture system, and dynamic data was collected at 18 key points such as shoulder, elbow, and wrist, including finger posture and arm range, to ensure that the data covers parameters such as age, gender, and habits, forming an initial dataset (total 1.2TB). The cross-modal data sources include images, audio, and motion files (BVH format). (2) Data analysis and map construction module: The three-dimensional evaluation matrix was used to perform cross-modal semantic alignment on the collected data; and cultural experts were organized to process the data, mapping the patterns, audio, and motion into a unified information library to form a multi-source heterogeneous feature map information library. (3) Learning interaction and innovation generation application stage: The subjects used AR glasses, tactile pens and other devices for phased training. Under the demonstration of virtual tutors, the novice subjects had a high acceptance of Miao paper-cutting art and achieved an accuracy rate of 40% in completing the paper-cutting operation. The skilled subjects even used the information database to generate new patterns, input user intentions (such as "integrating modern elements"), and adjusted the innovation coefficient, outputting more than 500 effective and authentic design schemes per day.

[0043] Example 6 To verify the effectiveness of the present invention, multiple cross-experiments were conducted on 50 learners and 200 pattern samples. The traditional method (based on scanning and basic model) and the method of the present invention were compared. Based on system logs, expert evaluation and user feedback, the data comparison in Table 1 was obtained.

[0044] Table 1. Comparative Experimental Data on Cross-Modal Generation Algorithm for Learning and Inheriting Miao Paper-cutting Art

[0045] This demonstrates that, in terms of cultural fidelity, the accuracy of pattern structure assessments based on the image library jumped from 68% to 96%, confirming the accurate association between patterns and folk customs; in terms of generation efficiency, the average daily output of design schemes increased from 80 to 1200, supporting the rapid iteration of the cultural and creative industries; in terms of learning efficiency for test subjects, the average learning time was shortened from 6 months to 4 weeks, thanks to the real-time correction system (feedback latency <100ms); the error rate of personnel data collection was significantly reduced after the experiment, invalid points were reduced by 90% after motion capture cleaning, and trajectory accuracy reached 98%; color misreading decreased from 25% to 3%; and in terms of user learning feedback, the retention rate was 95%, with learners reporting that "immersive guidance improved efficiency and reduced frustration."

[0046] Table 2

[0047] The above case demonstrates the high feasibility of this invention in art-related intangible cultural heritage projects: the implementation process is simple and efficient, and the results show a 41.2% improvement in cultural fidelity, an 83.3% increase in learning efficiency, and a 1400% expansion in generation capabilities, completely overcoming the problems of fragmented storage, stylistic inaccuracies, and inefficient transmission. This method can be extended to other intangible cultural heritages (such as Suzhou embroidery or Thangka), promoting living heritage and cultural revival. It can be piloted in more regions to accelerate the digital innovation of intangible cultural heritage and ensure the accuracy of intangible cultural heritage transmission and the effectiveness of learning.

[0048] The following table compares the method of this invention with traditional methods in the inheritance and interactive learning of intangible cultural heritage projects in the arts: Table 3 Comparison of learning outcomes between the present invention and traditional learning methods

[0049] The learning speed data comes from the tracking results of 50 learners in Example 6 (the traditional method averages 4 months, while this method averages 6 weeks); Error correction rate: Compared with the error rate based on motion capture, the traditional method has an operation error of 35%, while this method reduces the error to 4% through real-time feedback; Creation threshold quantification: The traditional mode requires mastery of two professional software programs (such as AI + PS), while this method requires no software background; Work innovation statistics: In the traditional mode, only 30 out of 200 samples contain improved designs, while in this method, 864 out of 1200 generated solutions contain innovative elements.

[0050] It should be noted that the present invention is not limited to the specific embodiments described above. For those skilled in the art, several improvements or substitutions can be made based on the concept of the present invention, and all such improvements or substitutions should be considered to fall within the scope of the present invention.

Claims

1. A method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms, characterized in that, We will construct a multi-source heterogeneous feature map information database for intangible cultural heritage projects, develop cross-style generation algorithms that integrate cross-modal feature fusion, and combine interactive learning and virtual reality technologies to integrate and create artistic content.

2. The method for the inheritance and interactive learning of intangible cultural heritage projects in the arts based on cross-modal generation algorithms according to claim 1, characterized in that, The specific operating steps are as follows: Step 1: Collect information and data on intangible cultural heritage projects in the field of fine arts, including different types of pattern images, videos of craft operations, audio recordings of inheritors' oral accounts, and regional folk custom texts, to form multi-source heterogeneous data of intangible cultural heritage objects, and clean the data; Step 2: Classify and label the data generated in Step 1 to form a semantically related dataset. Then, stratify the data of the intangible cultural heritage project according to the difficulty of inheritance, the types of patterns, the cultural connotations, and the requirements of craftsmanship techniques. Finally, convert the different data types obtained into a unified format. Step 3: The data processed in Step 2 is further filtered and analyzed, and classified into the static pattern library, dynamic process library, and contextual knowledge library. These three types of graph information libraries constitute the standard multi-source heterogeneous feature graph information library of this intangible cultural heritage project. The data in each library represents different modalities. Step 4: Using an improved cross-style generation algorithm, the intangible cultural heritage features in different modalities from Step 3 are fused to generate several diverse style schemes that can both preserve the genes of traditional culture and adjust the degree of innovation. Step 5: Utilize virtual-real fusion learning design and cross-style generation algorithms for learning and design reconstruction, and combine user historical creation data for the fusion learning and creation of artistic content.

3. The method for the inheritance and interactive learning of intangible cultural heritage projects in the arts based on cross-modal generation algorithms according to claim 2, characterized in that, Step 1 involves acquiring no less than 1TB of multi-source heterogeneous data using multi-sensor acquisition devices, and then having cultural experts collaboratively annotate the data to form a semantic association dataset. The multi-sensor acquisition devices include a high-definition scanner, a motion capture system, and an environmental recording device. The high-definition scanner has a resolution of ≥600 DPI and is used to collect two-dimensional images of representative patterns from intangible cultural heritage projects. The motion capture system uses Vicon optical motion capture equipment to record the hand movements of the inheritor during the production process, with a sampling frequency of ≥120Hz. The environmental recording device has a sampling rate of 44.1 kHz and a duration of ≥50 h.

4. The method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms according to claim 3, characterized in that, Step 2 is as follows: A three-dimensional evaluation matrix is ​​constructed, and the data processed in step 2 is subjected to secondary filtering and analysis to complete cross-modal semantic annotation, as follows: The structural features of the pattern images are labeled, namely PSD / EPS / JPG values; and the color codes are labeled, namely HSV / RGB values. The key action nodes and time sequence relationships of the process operation videos are labeled, namely MP4 / MOV values. The spoken audio is segmented to establish a term-pattern-process association mapping. Image noise is processed by denoising algorithms, and frequency domain features of action trajectories are extracted by Fourier transform. The spoken audio is transcribed into text using speech recognition technology and stored in JSON format to form a semantically related dataset.

5. The method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms according to claim 3, characterized in that, When recording the hand and arm movements of inheritors in craft videos, optical motion capture instruments are used to record the positional changes and orientation data of the inheritors' ten fingers, two wrists, and upper and lower arms. The acquired data is processed and divided into time and space to obtain operational technique data for intangible cultural heritage projects in the arts, which serves as a database of standard technique movements for inheritors.

6. The method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms according to claim 1, characterized in that, Step 4 is as follows: By fully leveraging the complementarity between different modalities, information extracted from different modalities is integrated into a stable cross-modal representation, resulting in a more accurate and standardized fitting dynamic equation: Where Z is the cross-modal fusion representation library. , , It is a learnable cross-modal projection matrix; each represents a matrix group formed by multiple combinations of visual, text, and action elements; v is the visual feature vector; t It is a text feature vector; It is an action feature vector; By fusing intangible cultural heritage features from different modalities, a cross-style generation algorithm is developed to generate diverse works that retain traditional cultural genes while allowing for adjustable levels of innovation. The improved cross-style generation algorithm's diffusion equation is as follows: Traditional cultural feature coding: ; Innovation-guided feature coding: ; in, express n The noisy image input at any given time; These are images obtained by sampling according to a normal distribution during the generation process; It is the noise figure of the diffusion process, which is less than 1. It is an innovative adjustment parameter. ;pass The parameters directly control the trade-off between cultural preservation and innovation. It is a noise prediction method that preserves traditional cultural characteristics. It is a noise prediction that incorporates innovative elements, using two independent noise predictors to handle cultural characteristics and innovation guidance respectively; It is the process of converting intangible cultural heritage elements into numerical vectors. It is the process of converting users' innovative intentions and modern elements into vectors.

7. The method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms according to claim 5, characterized in that, The weight adjustment formula for the cross-style generation algorithm in step 4 is as follows: in, For input parameters, ,when It relies entirely on database features for generation, when It allows for the introduction of 5%-10% innovation perturbation; y is the final generated feature symbol vector set. It is the cultural gene vector of the basic pattern. It is the final output vector representing the innovative style.

8. The method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms according to claim 7, characterized in that, Step 5 is as follows: Step 5.1: Using wearable devices, learners access the multi-source heterogeneous feature map information database of the intangible cultural heritage project obtained in Step 1. Learners select to learn relevant intangible cultural heritage knowledge and skills. Through quantitative learning in the database, learners establish their cognition. Step 5.2: The wearable device system compares the learned movement with the standard movement in real time; if a difference is found between the learned movement and the standard movement during the movement learning process, the wearable device system will remind the learner to make corrections and then make repeated adjustments until it matches the standard. Step 5.3: Select relevant data samples from the standard multi-source heterogeneous feature map information database obtained for this intangible cultural heritage project, and learners can then carry out cultural and creative design; the cultural and creative design includes selecting relevant patterns to combine and generate works with accurate dissemination significance; Step 5.4: Correct the content of the created works by adjusting the weight equation of the cross-style generation algorithm. For example, if there are errors in the application of dynasties, patterns, techniques, etc. in the works, the correction process is synchronized in the interactive learning stage. The purpose is to make it conform to the accurate expression of the intangible cultural heritage project. If the created works generated after adjustment conform to the characteristics of the learned intangible cultural heritage, the works are output. If they do not conform, the algorithm weights are adjusted multiple times until a works that conform to the characteristics of inheritance are generated. Step 5.5: If the learner is satisfied with the generated work, output the work; if not, repeat step 5.3 until satisfied.

9. The method for the inheritance and interactive learning of intangible cultural heritage projects in the field of fine arts based on cross-modal generation algorithms according to claim 8, characterized in that, The wearable devices include AR glasses, haptic gloves, and electronic pen.