Cultural and creative product interaction method and system based on multi-impression experience fusion

By constructing a multimodal interaction model that combines visual, auditory, and tactile information, the problem of multisensory experience integration and insufficient user immersion in the display and interaction technology of cultural and creative products has been solved, achieving a flexible and natural interaction method and a deep immersive experience.

CN121502191APending Publication Date: 2026-02-10SHANXI DEEPIN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511517137.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing cultural and creative product display and interaction technologies are insufficient in terms of multi-sensory experience integration, flexibility and naturalness of interaction methods, and improvement of user immersion. Existing technologies mainly focus on the visual level, and the integration of other sensory experiences is relatively limited, affecting the flexibility and naturalness of interaction.

Method used

By acquiring multimodal cultural and creative training data, an initial multimodal interaction model is constructed, and a deep neural network architecture is used for training. The model is iteratively optimized to generate interactive responses, and visual, auditory, and tactile information is combined to enhance the user interaction experience.

Benefits of technology

It achieves a deep integration of multi-sensory experiences, enhances the flexibility of interaction methods and the user's immersion, and meets the needs of the cultural and creative industries for efficient and immersive interactive systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502191A_ABST
    Figure CN121502191A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cultural and creative product interaction, in particular to a cultural and creative product interaction method and system based on multi-impression experience fusion, and the method comprises the steps: obtaining multi-modal cultural and creative training data, extracting a verification data set and a residual training data set, constructing an initial multi-modal interaction model, and optimizing the model through iterative training and performance evaluation. And finally, generating a target multi-modal interaction model to realize interaction response of multi-sensory fusion. According to the method, multi-sensory experience fusion, interaction flexibility and user immersion in cultural and creative product display and interaction can be effectively improved, and the defects of multi-sensory collaboration and user experience in the prior art are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of cultural and creative product interaction, specifically a cultural and creative product interaction method and system based on multi-sensory experience fusion. BACKGROUND

[0002] With the rapid development of cultural and creative industries, cultural and creative product interaction methods and systems based on multi-sensory experience fusion have gradually become a research hotspot. This type of technology combines multiple sensory experiences (such as vision, hearing, touch, etc.) and interaction methods to provide users with immersive product display and interactive experiences, thereby enhancing the appeal of cultural and creative products and the effectiveness of cultural dissemination. However, existing related technologies still have room for improvement in terms of multi-sensory fusion, depth and diversity of interactive experience, and user experience optimization.

[0003] After searching, a cultural and creative product display method and display device based on AR (CN117576355B) was disclosed on April 19, 2024. This patent builds a three-dimensional model of cultural and creative products and displays the model on an interactive plane using augmented reality (AR) technology, achieving interactive control based on calibration points and improving the positional accuracy and interactive convenience of AR display. However, this technical solution mainly focuses on visual augmented reality display and has limited fusion design for other sensory experiences (such as hearing and touch), and the user's sense of immersion and depth of interaction need to be further improved. In addition, its interactive method relies on calibration point settings, which may affect the flexibility and naturalness of interaction in complex scenarios.

[0004] After searching, a protection method for traditional sports non-material cultural heritage (CN110069829B) was disclosed on October 18, 2022. This patent uses motion capture technology and database construction, combined with an interactive learning system and 3D printing technology, to achieve digital protection and activation display of traditional sports non-material cultural heritage. However, this technical solution focuses on data collection and reconstruction of cultural heritage and lacks sufficient attention to multi-sensory experience fusion. The interactive method mainly relies on traditional learning systems, and immersive and diversified interactive methods need to be strengthened. The needs of users for cultural and creative product interaction experience are not fully met.

[0005] The above problems show that existing cultural and creative product display and interaction technologies still have room for improvement in terms of multi-sensory experience fusion, flexibility and naturalness of interactive methods, and improvement of user immersion. Therefore, the present application provides a cultural and creative product interaction method and system based on multi-sensory experience fusion, aiming to integrate visual, auditory, tactile, and other sensory experiences, design flexible and natural interaction methods, and optimize user experience, thereby meeting the needs of the cultural and creative industry for efficient and immersive interactive systems. SUMMARY

[0006] The application provides a cultural and creative product interaction method and system based on multi-sensory experience fusion, which mainly aims to solve the problems of the existing cultural and creative product display and interaction technology in multi-sensory experience fusion, interaction mode flexibility, and user immersion improvement.

[0007] To achieve the above-mentioned purpose, the application provides a cultural and creative product interaction method based on multi-sensory experience fusion, which includes: obtaining multi-modal cultural and creative training data, wherein the multi-modal cultural and creative training data includes visual material training data, auditory material training data, tactile feedback training data, and interaction behavior annotation data; extracting a validation data set from the multi-modal cultural and creative training data according to a preset validation data ratio, obtaining a remaining training data set; constructing an initial multi-modal interaction model, wherein the initial multi-modal interaction model is based on a deep neural network architecture design; extracting an iterative training data set from the remaining training data set according to a preset training data ratio and an initial sampling weight, obtaining an updated training data set; training the initial multi-modal interaction model using the iterative training data set, obtaining an iterative multi-modal interaction model; evaluating the performance of the iterative multi-modal interaction model using the validation data set, obtaining an evaluation error distribution and a maximum evaluation error; determining whether the maximum evaluation error is less than a preset error tolerance threshold; if the maximum evaluation error is not less than the error tolerance threshold, adjusting the remaining training data set, the initial multi-modal interaction model, and the initial sampling weight using the updated training data set, the iterative multi-modal interaction model, and the evaluation error distribution, respectively, and returning to the step of extracting the iterative training data set from the remaining training data set according to the preset training data ratio and the initial sampling weight; if the maximum evaluation error is less than the error tolerance threshold, taking the iterative multi-modal interaction model as a target multi-modal interaction model; receiving multi-modal cultural and creative real-time data, generating an interaction response using the target multi-modal interaction model according to the multi-modal cultural and creative real-time data, and obtaining a response output; extracting core interaction information from the multi-modal cultural and creative real-time data according to the response output.

[0008] Optionally, the extracting a verification data set from the multi-modal cultural and creative training data according to a preset verification data proportion, to obtain a remaining training data set, comprises: sequentially and randomly extracting multi-modal cultural and creative sample data from the multi-modal cultural and creative training data; performing sample statistics on the multi-modal cultural and creative sample data to obtain a random sample number; determining whether the random sample number is less than a total sample number corresponding to the verification data proportion; if the random sample number is less than the total sample number, returning to the step of sequentially and randomly extracting multi-modal cultural and creative sample data from the multi-modal cultural and creative training data; if the random sample number is not less than the total sample number, aggregating all multi-modal cultural and creative sample data to obtain an initial verification data set; identifying an interactive behavior annotation distribution of the initial verification data set, and determining whether the interactive behavior annotation distribution covers a preset complete annotation range; if the interactive behavior annotation distribution does not cover the complete annotation range, returning to the step of sequentially and randomly extracting multi-modal cultural and creative sample data from the multi-modal cultural and creative training data; if the interactive behavior annotation distribution covers the complete annotation range, taking the initial verification data set as the verification data set; removing the verification data set from the multi-modal cultural and creative training data to obtain the remaining training data set.

[0009] Optionally, the constructing an initial multi-modal interactive model comprises: setting a visual input node, an auditory input node and a tactile input node to form a multi-modal input layer node set; setting an interactive response output node to form an output layer node; and generating the initial multi-modal interactive model by using a pre-constructed deep neural network architecture according to the multi-modal input layer node set and the output layer node.

[0010] Optionally, the extracting an iterative training data set from the remaining training data set according to a preset training data proportion and an initial sampling weight, to obtain an updated training data set, comprises: calculating a visual material sample number, an auditory material sample number and a tactile feedback sample number by using the following formula according to the training data proportion and the initial sampling weight: wherein, represents an i-th material sample number, represents a total sample number corresponding to the training data proportion, and represents a weight coefficient of the i-th material in the initial sampling weight; extracting visual material sample data, auditory material sample data and tactile feedback sample data from the remaining training data set according to the visual material sample number, the auditory material sample number and the tactile feedback sample number, to obtain the iterative training data set; and removing the iterative training data set from the remaining training data set to obtain the updated training data set.

[0011] Optionally, training the initial multimodal interaction model using the iterative training dataset to obtain an iterative multimodal interaction model includes: performing image segmentation on the visual material sample data to obtain segmented image data; normalizing the segmented image data to obtain standard visual data; extracting features from the standard visual data using a pre-built convolutional neural network to obtain visual feature vectors; performing spectral transformation on the auditory material sample data to obtain spectral data; performing noise reduction on the spectral data to obtain standard auditory data; performing sequence modeling on the standard auditory data using a pre-built recurrent neural network to obtain auditory feature vectors; performing signal quantization on the tactile feedback sample data to obtain quantized tactile data; smoothing the quantized tactile data to obtain standard tactile data; extracting features from the standard tactile data using a pre-built autoencoder to obtain tactile feature vectors; sequentially extracting single-example visual features, single-example auditory features, and single-example tactile features from the visual feature vector, auditory feature vector, and tactile feature vector respectively; and performing sequence modeling on the single-example visual features and single-example auditory features. Features and single-instance tactile features are input into the initial multimodal interaction model from the visual input node, auditory input node, and tactile input node, respectively, to obtain iterative output results; interactive behavior annotations corresponding to the single-instance visual features, single-instance auditory features, and single-instance tactile features are obtained; the target value range of the interactive behavior annotations is extracted from the pre-constructed interactive behavior value range table, and the center value is extracted from the target value range; the single-instance error between the iterative output result and the center value is calculated, and it is determined whether the single-instance error is less than the error tolerance threshold; if the single-instance error is not less than the error tolerance threshold, the model is considered to be in a state of flux. If the error tolerance threshold is not met, the parameters of the initial multimodal interaction model are adjusted according to the singleton error until the singleton error is less than the error tolerance threshold. It is then determined whether the extraction of singleton visual features, singleton auditory features, and singleton tactile features from the visual feature vector, auditory feature vector, and tactile feature vector has been completed. If extraction is not completed, the process returns to the steps described above, where singleton visual features, singleton auditory features, and singleton tactile features are extracted sequentially from the visual feature vector, auditory feature vector, and tactile feature vector, respectively. If extraction is completed, an iterative multimodal interaction model is obtained.

[0012] Optionally, the step of using the verification dataset to evaluate the performance of the iterative multimodal interaction model and obtain the evaluation error distribution and maximum evaluation error includes: classifying the verification dataset to obtain visual verification sample data, auditory verification sample data, and tactile verification sample data; converting the visual verification sample data, auditory verification sample data, and tactile verification sample data into visual verification features, auditory verification features, and tactile verification features, respectively; sequentially extracting single-example visual verification features, single-example auditory verification features, and single-example tactile verification features from the visual verification features, auditory verification features, and tactile verification features, respectively; inputting the single-example visual verification features, single-example auditory verification features, and single-example tactile verification features into the iterative multimodal interaction model to obtain verification output results; and determining whether the visual verification features, auditory verification features, and tactile verification features have completed single-example visual verification and single-example auditory verification. Extraction of features and single-example tactile verification features; if extraction is not completed, return to the steps above of sequentially extracting single-example visual verification features, single-example auditory verification features, and single-example tactile verification features from the visual verification features, auditory verification features, and tactile verification features respectively; if extraction is completed, summarize the verification output results to obtain a verification output result set; classify the verification output result set according to the interactive behavior annotation data to obtain multiple sets of verification result subsets; sequentially extract verification result subsets from the multiple sets of verification result subsets, and identify the target value range center value corresponding to the verification result subset; calculate the error set between the verification result subset and the target value range center value, calculate the error mean based on the error set, and obtain the error mean distribution; perform statistical analysis on the error mean distribution to obtain the evaluation error distribution, extract the maximum error mean from the error mean distribution, and use the maximum error mean as the maximum evaluation error.

[0013] Optionally, the step of extracting core interactive information from the multimodal cultural and creative real-time data based on the response output includes: extracting visual keywords, auditory key segments, and tactile key signals from a pre-constructed interactive behavior-multimodal mapping table based on the response output; extracting core visual elements, core auditory segments, and core tactile signals from the real-time visual data, real-time auditory data, and real-time tactile data based on the visual keywords, auditory key segments, and tactile key signals, respectively; and summarizing the core visual elements, core auditory segments, and core tactile signals to obtain core interactive information.

[0014] Optionally, before extracting visual keywords, auditory key segments, and tactile key signals from the pre-constructed interaction behavior-multimodal mapping table based on the response output, the method further includes: sequentially extracting interaction behavior categories from the interaction behavior annotation data; receiving visual keywords, auditory key segments, and tactile key signals set by the user for the interaction behavior categories; and constructing an interaction behavior-multimodal mapping table based on the interaction behavior categories and the visual keywords, auditory key segments, and tactile key signals.

[0015] To achieve the above objectives, the present invention also provides an interactive system for cultural and creative products based on multi-sensory experience fusion, comprising: an initial multimodal interaction model training module, used to acquire multimodal cultural and creative training data, wherein the multimodal cultural and creative training data includes: visual material training data, auditory material training data, tactile feedback training data, and interactive behavior annotation data; extracting a verification dataset from the multimodal cultural and creative training data according to a preset verification data ratio to obtain a remaining training dataset; constructing an initial multimodal interaction model, wherein the initial multimodal interaction model is based on a deep neural network architecture design; extracting an iterative training dataset from the remaining training dataset according to a preset training data ratio and initial sampling weights to obtain an updated training dataset; training the initial multimodal interaction model using the iterative training dataset to obtain an iterative multimodal interaction model; and an iterative multimodal interaction model evaluation module, used to evaluate the iterative multimodal interaction model using the verification dataset. A performance evaluation is performed to obtain the evaluation error distribution and the maximum evaluation error. A target multimodal interaction model acquisition module is used to determine whether the maximum evaluation error is less than a preset error tolerance threshold. If the maximum evaluation error is not less than the error tolerance threshold, the remaining training dataset, the initial multimodal interaction model, and the initial sampling weights are adjusted using the updated training dataset, the iterative multimodal interaction model, and the evaluation error distribution, respectively. The process then returns to the steps described above, where the iterative training dataset is extracted from the remaining training dataset according to the preset training data ratio and the initial sampling weights. If the maximum evaluation error is less than the error tolerance threshold, the iterative multimodal interaction model is used as the target multimodal interaction model. A core interaction information extraction module is used to receive real-time multimodal cultural and creative data, generate an interactive response using the target multimodal interaction model based on the real-time multimodal cultural and creative data, and obtain a response output. Core interaction information is extracted from the real-time multimodal cultural and creative data based on the response output.

[0016] To address the aforementioned issues, the present invention also provides an electronic device, comprising: a memory storing at least one instruction; and a processor executing the instruction stored in the memory to implement the aforementioned interactive method for cultural and creative products based on the fusion of multiple sensory experiences.

[0017] To address the aforementioned issues, the present invention also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the above-described interactive method for cultural and creative products based on the fusion of multiple sensory experiences.

[0018] To address the problems described in the background section, this invention first trains an initial multimodal interaction model using an iterative training dataset, and then evaluates the performance of the iterative multimodal interaction model using a validation dataset. Once the evaluation conditions are met, an interactive response is generated using the target multimodal interaction model based on real-time multimodal cultural and creative data, and core interactive information is extracted from the real-time multimodal cultural and creative data. Before training the initial multimodal interaction model, it is necessary to acquire multimodal cultural and creative training data and extract a validation dataset from it to obtain a remaining training dataset. This remaining training dataset can then be used to train the initial multimodal interaction model. During training, the initial multimodal interaction model is trained in multiple batches. First, an iterative training dataset is extracted from the remaining training dataset according to a preset training data ratio and initial sampling weights. This iterative training dataset is then used to train the initial multimodal interaction model, resulting in an iterative multimodal interaction model. To verify the training effect, the performance of the iterative multimodal interaction model is evaluated using the validation dataset to obtain the evaluation error distribution and the maximum evaluation error. When the maximum evaluation error is not less than the error tolerance threshold, the remaining training dataset, the initial multimodal interaction model, and the initial sampling weights are adjusted using the updated training dataset, the iterative multimodal interaction model, and the evaluation error distribution, and the iterative multimodal interaction model is retrained for the next batch. When the maximum evaluation error is less than the error tolerance threshold, the iterative multimodal interaction model is used as the target multimodal interaction model. At this point, interactive responses can be generated using the target multimodal interaction model based on the real-time multimodal cultural and creative data, and core interactive information can be extracted. Therefore, this invention can address the shortcomings of existing cultural and creative product display and interaction technologies in terms of multisensory experience integration, interaction flexibility, and enhanced user immersion. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the interactive method for cultural and creative products based on the fusion of multiple sensory experiences in an embodiment of the present invention, which shows the overall process from obtaining multimodal cultural and creative training data to extracting core interactive information.

[0020] Figure 2 This is a schematic diagram of the initial multimodal interaction model construction and training process in an embodiment of the present invention, which focuses on the process of model construction, data sampling and iterative training.

[0021] Figure 3This is a structural block diagram of a cultural and creative product interactive system based on multi-sensory experience integration in an embodiment of the present invention, showing the composition of each functional module of the system and their interrelationships.

[0022] The attached figures are labeled as follows: 1. Visual input node; 2. Auditory input node; 3. Tactile input node; 4. Interactive response output node; 5. Initial multimodal interaction model; 6. Validation dataset; 7. Iterative training dataset; 8. Target multimodal interaction model; 9. Core interaction information. Detailed Implementation

[0023] This invention provides a method and system for interactive cultural and creative products based on multi-sensory experience fusion. Its core lies in achieving multi-sensory interactive functions for cultural and creative products through the acquisition, processing, and model training of multimodal data. The following is in conjunction with the appendix... Figure 1 , Figure 2 and Figure 3 The accompanying reference numerals are explained in detail.

[0024] exist Figure 1 The document demonstrates the overall process from acquiring multimodal cultural and creative product training data to extracting core interactive information. First, multimodal training data needs to be acquired, including visual material training data, auditory material training data, haptic feedback training data, and interactive behavior annotation data. Visual material training data primarily comes from images or videos of cultural and creative products; auditory material training data includes background sound effects or voice narration; and haptic feedback training data involves force feedback signals when users interact with the products. Interactive behavior annotation data records user behavior patterns during interaction with the products, such as clicks, touches, or swipes. This data is stored in the electronic device's memory and read by the processor for subsequent processing.

[0025] To ensure the reliability of the training data, after acquiring the multimodal cultural and creative training data, a validation dataset (6) needs to be extracted from it according to a preset validation data ratio. This process is as follows: Figure 1 As shown, firstly, multimodal cultural and creative sample data is randomly extracted sequentially from the multimodal cultural and creative training data, and these sample data are statistically analyzed to obtain the number of random samples. Then, it is determined whether the number of random samples is less than the total number of samples corresponding to the proportion of validation data. If the number of random samples is insufficient, extraction continues until the condition is met. Once the number of random samples reaches the requirement, all extracted multimodal cultural and creative sample data is summarized to form the initial validation dataset. Next, the interactive behavior annotation distribution of the initial validation dataset needs to be identified, and it needs to be determined whether it covers the preset complete annotation range. If it does not cover the complete annotation range, sample data is re-extracted; if it does cover the range, the initial validation dataset is used as validation dataset 6 and removed from the multimodal cultural and creative training data to obtain the remaining training dataset.

[0026] After completing data segmentation, an initial multimodal interaction model 5 is constructed. For example... Figure 2 As shown, the construction process of the initial multimodal interaction model 5 includes setting up a visual input node 1, an auditory input node 2, and a tactile input node 3 to form a multimodal input layer node set. Simultaneously, an interactive response output node 4 is set up to form an output layer node. Based on this, the initial multimodal interaction model 5 is generated using a pre-built deep neural network architecture. Visual input node 1 receives visual material training data, auditory input node 2 receives auditory material training data, and tactile input node 3 receives tactile feedback training data. These input nodes are interconnected through the deep neural network architecture, ultimately transmitting the processing results to the interactive response output node 4.

[0027] After the initial multimodal interaction model 5 is constructed, iterative training dataset 7 needs to be extracted from the remaining training dataset according to the preset training data ratio and initial sampling weights. For example... Figure 2 As shown, first according to the formula : The number of visual material samples, auditory material samples, and tactile feedback samples are calculated, where represents the number of samples of the i-th type of material, represents the total number of samples corresponding to the proportion of training data, and represents the weight coefficient of the i-th type of material in the initial sampling weights. Then, based on the calculation results, visual material sample data, auditory material sample data, and tactile feedback sample data are extracted from the remaining training dataset to form iterative training dataset 7. After removing iterative training dataset 7 from the remaining training dataset, the updated training dataset is obtained.

[0028] The process of training the initial multimodal interaction model 5 using the iterative training dataset 7 is as follows: Figure 2As shown, the visual material sample data is first segmented to obtain segmented image data, which is then normalized to generate standard visual data. A pre-built convolutional neural network is then used to extract features from the standard visual data, resulting in visual feature vectors. For auditory material sample data, spectral transformation is first performed to generate spectral data, which is then denoised to generate standard auditory data. A pre-built recurrent neural network is then used to perform sequence modeling on the standard auditory data, resulting in auditory feature vectors. For tactile feedback sample data, signal quantization is first performed to generate quantized tactile data, which is then smoothed to generate standard tactile data. A pre-built autoencoder is then used to extract features from the standard tactile data, resulting in tactile feature vectors. After feature extraction, single-example visual features, single-example auditory features, and single-example tactile features are sequentially extracted from the visual feature vector, auditory feature vector, and tactile feature vector, respectively. These features are then input into the initial multimodal interaction model 5 from visual input node 1, auditory input node 2, and tactile input node 3, respectively, to obtain iterative output results. Subsequently, the interactive behavior annotations corresponding to the single-example visual features, single-example auditory features, and single-example tactile features are obtained, and the target value center value is extracted from the pre-constructed interactive behavior value range table. By calculating the single-example error between the iterative output result and the target value center value, it is determined whether the single-example error is less than the error tolerance threshold. If the single-example error is not less than the error tolerance threshold, the parameters of the initial multimodal interaction model 5 are adjusted according to the single-example error until the single-example error is less than the error tolerance threshold. The above steps are repeated until all features are extracted, and finally the iterative multimodal interaction model is obtained.

[0029] After completing iterative training, the performance of the iterative multimodal interaction model needs to be evaluated using validation dataset 6. For example... Figure 1As shown, the validation dataset 6 is first classified to obtain visual validation sample data, auditory validation sample data, and tactile validation sample data. These data are then converted into visual validation features, auditory validation features, and tactile validation features, respectively, and single-example visual validation features, single-example auditory validation features, and single-example tactile validation features are extracted sequentially. These features are then input into the iterative multimodal interaction model to obtain validation output results. After all features are extracted, the validation output results are summarized to generate a validation output result set. The validation output result set is classified according to the interaction behavior annotation data to obtain multiple subsets of validation results. The target value center value is extracted from each subset of validation results, and the error set between the validation result subset and the target value center value is calculated. Statistical analysis of the error set yields the evaluation error distribution and the maximum evaluation error. If the maximum evaluation error is not less than the error tolerance threshold, the remaining training dataset, the initial multimodal interaction model, and the initial sampling weights are adjusted using the updated training dataset, iterative multimodal interaction model, and evaluation error distribution, and the training steps described above are returned. If the maximum evaluation error is less than the error tolerance threshold, then the iterative multimodal interaction model will be used as the target multimodal interaction model.

[0030] After training the target multimodal interaction model 8, it can receive real-time multimodal cultural and creative data and generate interactive responses. For example... Figure 1 As shown, firstly, a response output is generated using the target multimodal interaction model 8 based on the multimodal cultural and creative real-time data, and then core interaction information 9 is extracted from the multimodal cultural and creative real-time data based on the response output. Before extracting the core interaction information 9, interaction behavior categories need to be extracted sequentially from the interaction behavior annotation data, and the visual keywords, auditory key segments, and tactile key signals set by the user for the interaction behavior categories are received. An interaction behavior-multimodal mapping table is constructed based on the interaction behavior categories and the visual keywords, auditory key segments, and tactile key signals. Subsequently, visual keywords, auditory key segments, and tactile key signals are extracted from the interaction behavior-multimodal mapping table based on the response output, and core visual elements, core auditory segments, and core tactile signals are extracted from the real-time visual data, real-time auditory data, and real-time tactile data, respectively. Finally, the core visual elements, core auditory segments, and core tactile signals are summarized to obtain the core interaction information 9.

[0031] In practical applications, this invention can be implemented using electronic devices. For example... Figure 3As shown, the electronic device includes a memory and a processor. The memory stores multimodal cultural and creative product training data, an initial multimodal interaction model 5, a validation dataset 6, an iterative training dataset 7, and a target multimodal interaction model 8. The processor executes the instructions stored in the memory to implement an interactive method for cultural and creative products based on multi-sensory experience fusion. In addition, the system also includes an initial multimodal interaction model training module, an iterative multimodal interaction model evaluation module, a target multimodal interaction model acquisition module, and a core interaction information extraction module. These modules work together to complete the entire process from data acquisition to core interaction information extraction.

[0032] This invention achieves multi-sensory interactive functions for cultural and creative products through the above-mentioned method, and solves the shortcomings of existing technologies in terms of multi-sensory experience integration, interaction flexibility and user immersion enhancement.

[0033] To enable those skilled in the art to fully understand and implement this invention, the specific implementation principle of this invention will be further explained below in conjunction with a specific application scenario.

[0034] In practical applications, when users interact with cultural and creative products through electronic devices, they first need to input multimodal training data containing visual, auditory, and tactile information into the system. For example, in a virtual reality (VR)-based cultural and creative display system, a user might want to experience the process of making a traditional handicraft. In this case, the system's memory stores relevant visual material training data (such as 3D model images of the handicraft), auditory material training data (such as the sounds of tools used during the making process), and tactile feedback training data (such as force feedback signals when the user simulates touching the material). This data is read and processed by the processor to construct an initial multimodal interaction model.

[0035] In constructing the initial multimodal interaction model 5, visual input node 1, auditory input node 2, and tactile input node 3 are first set up to form a multimodal input layer node set. Then, an interaction response output node 4 is set up as the model's output layer node. Based on this, the initial multimodal interaction model 5 is generated using a pre-built deep neural network architecture. Visual input node 1 receives image data from the 3D model of the handicraft, auditory input node 2 receives audio data such as tool tapping sounds, and tactile input node 3 receives tactile feedback signals from the user's simulated operation. These input nodes are interconnected through the deep neural network architecture, and the processing results are ultimately transmitted to the interaction response output node 4 to generate the user's interaction response.

[0036] Next, the system extracts iterative training dataset 7 from the remaining training dataset according to the preset training data ratio and initial sampling weights. For example, assuming the training data ratio is set to 80%, and the initial sampling weights for visual, auditory, and tactile data are 0.5, 0.3, and 0.2 respectively, the number of visual material samples, auditory material samples, and tactile feedback samples can be calculated using the formula \(N_i=T\timesW_i\). The system extracts the corresponding sample data according to the calculation results to form iterative training dataset 7. Subsequently, iterative training dataset 7 is removed from the remaining training dataset to obtain an updated training dataset, which is used for subsequent model training.

[0037] When training the initial multimodal interaction model 5 using the iterative training dataset 7, the system first performs image segmentation on the visual material sample data to obtain segmented image data, and then normalizes it to generate standard visual data. Next, a pre-built convolutional neural network is used to extract features from the standard visual data to obtain visual feature vectors. For auditory material sample data, the system first converts it into spectral data and performs noise reduction to generate standard auditory data. Subsequently, a pre-built recurrent neural network is used to perform sequence modeling on the standard auditory data to obtain auditory feature vectors. For tactile feedback sample data, the system first performs signal quantization to generate quantized tactile data, and then smooths it to generate standard tactile data. Finally, a pre-built autoencoder is used to extract features from the standard tactile data to obtain tactile feature vectors. These feature vectors are input from visual input node 1, auditory input node 2, and tactile input node 3, respectively, into the initial multimodal interaction model 5 to generate iterative output results.

[0038] To ensure model performance, the system uses validation dataset 6 to evaluate the performance of the iterative multimodal interaction model. First, the system classifies validation dataset 6 to obtain visual, auditory, and tactile validation sample data. Then, these data are converted into visual, auditory, and tactile validation features, respectively, and single-example visual, auditory, and tactile validation features are extracted sequentially. These features are input into the iterative multimodal interaction model to generate validation output results. After extracting all features, the system aggregates the validation output results to generate a validation output result set, and classifies the result set according to the interaction behavior annotation data, obtaining multiple subsets of validation results. By statistically analyzing the error between each subset of validation results and the target value center value, the system obtains the evaluation error distribution and the maximum evaluation error. If the maximum evaluation error is not less than the error tolerance threshold, the system adjusts the remaining training dataset, the initial multimodal interaction model, and the initial sampling weights using updated training datasets, iterative multimodal interaction models, and the evaluation error distribution, and then returns to the above training steps. If the maximum evaluation error is less than the error tolerance threshold, then the iterative multimodal interaction model will be used as the target multimodal interaction model.

[0039] After training the target multimodal interaction model 8, the system can receive real-time multimodal cultural and creative data and generate interactive responses. For example, when a user interacts with handicrafts through a VR device, the system generates a response output based on the real-time multimodal cultural and creative data using the target multimodal interaction model 8. Before generating the response output, the system needs to extract interaction behavior categories sequentially from the interaction behavior annotation data and receive the visual keywords, auditory key segments, and tactile key signals set by the user for the interaction behavior categories. Based on the interaction behavior categories and these keywords, segments, and signals, the system constructs an interaction behavior-multimodal mapping table. Subsequently, the system extracts visual keywords, auditory key segments, and tactile key signals from the interaction behavior-multimodal mapping table based on the response output, and extracts core visual elements, core auditory segments, and core tactile signals from the real-time visual data, real-time auditory data, and real-time tactile data, respectively. Finally, the system summarizes these core elements to obtain core interaction information 9.

[0040] In actual operation, the processor in the electronic device executes the instructions stored in the memory to realize the above-mentioned interactive method for cultural and creative products based on the fusion of multi-sensory experiences. For example, when a user simulates touching the surface of a handicraft through VR gloves, the tactile input node 3 receives the force feedback signal collected by the glove sensor and inputs it into the target multimodal interaction model 8. The model generates corresponding interactive responses based on the tactile signals, such as adjusting the visual display of the surface material of the handicraft or playing relevant sound effects, thereby enhancing the user's immersion. At the same time, the initial multimodal interaction model training module, the iterative multimodal interaction model evaluation module, the target multimodal interaction model acquisition module, and the core interaction information extraction module in the system work together to complete the entire process from data acquisition to core interaction information extraction.

[0041] Through the above steps, this invention achieves multi-sensory interactive functionality for cultural and creative products. For example, when displaying traditional handicrafts, users can not only visually observe their appearance details, but also auditorily experience the sound atmosphere of the production process, and tactilely simulate their texture. This multi-sensory interactive method significantly enhances the user's immersion and interaction depth, overcoming the shortcomings of existing technologies in multi-sensory experience integration, interaction flexibility, and user immersion enhancement.

Claims

1. A method for interactive cultural and creative products based on the integration of multi-sensory experiences, characterized in that, The method includes: Acquire multimodal cultural and creative training data, wherein the multimodal cultural and creative training data includes visual material training data, auditory material training data, tactile feedback training data, and interactive behavior annotation data; The remaining training dataset is obtained by extracting the verification dataset from the multimodal cultural and creative training data according to the preset verification data ratio. Construct an initial multimodal interaction model, wherein the initial multimodal interaction model is designed based on a deep neural network architecture; Based on the preset training data ratio and initial sampling weights, an iterative training dataset is extracted from the remaining training dataset to obtain an updated training dataset; The initial multimodal interaction model is trained using the iterative training dataset to obtain an iterative multimodal interaction model; The performance of the iterative multimodal interaction model was evaluated using the validation dataset to obtain the evaluation error distribution and the maximum evaluation error. Determine whether the maximum evaluation error is less than a preset error tolerance threshold.

2. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 1, characterized in that, The step of determining whether the maximum evaluation error is less than a preset error tolerance threshold includes... If the maximum evaluation error is not less than the error tolerance threshold, then the remaining training dataset, the initial multimodal interaction model and the initial sampling weights are adjusted by updating the training dataset, iterating the multimodal interaction model and the evaluation error distribution, respectively, and the steps of extracting the iterative training dataset from the remaining training dataset according to the preset training data ratio and the initial sampling weights are returned. If the maximum evaluation error is less than the error tolerance threshold, then the iterative multimodal interaction model is taken as the target multimodal interaction model. Receive multimodal cultural and creative real-time data, generate an interactive response based on the target multimodal interaction model using the multimodal cultural and creative real-time data, and obtain the response output; Based on the response output, core interactive information is extracted from the multimodal cultural and creative real-time data.

3. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 1, characterized in that, The step of extracting a verification dataset from the multimodal cultural and creative training data according to a preset verification data ratio to obtain the remaining training dataset includes: Multimodal cultural and creative sample data are randomly extracted sequentially from the multimodal cultural and creative training data; The multimodal cultural and creative sample data were statistically analyzed to obtain the number of random samples; Determine whether the number of random samples is less than the total number of samples corresponding to the proportion of verification data; If the number of random samples is less than the total number of samples, then return to the steps described above of randomly extracting multimodal cultural and creative sample data from the multimodal cultural and creative training data. If the number of random samples is not less than the total number of samples, then all multimodal cultural and creative sample data are aggregated to obtain the initial verification dataset; Identify the distribution of interactive behavior annotations in the initial verification dataset, and determine whether the distribution of interactive behavior annotations covers the preset complete annotation range; If the distribution of the interactive behavior annotations does not cover the complete annotation range, then return to the steps described above of randomly extracting multimodal cultural and creative sample data from the multimodal cultural and creative training data. If the distribution of the interactive behavior annotations covers the complete annotation range, then the initial verification dataset will be used as the verification dataset. The validation dataset is removed from the multimodal cultural and creative training data to obtain the remaining training dataset.

4. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 3, characterized in that, The construction of the initial multimodal interaction model includes: Visual input node (1), auditory input node (2) and tactile input node (3) are set to form a multimodal input layer node set; Set the interactive response output node (4) to form the output layer node; Based on the set of multimodal input layer nodes and output layer nodes, an initial multimodal interaction model is generated using a pre-built deep neural network architecture (5).

5. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 4, characterized in that, The step of extracting iterative training data from the remaining training dataset according to a preset training data ratio and initial sampling weights to obtain an updated training dataset includes: Based on the training data ratio and initial sampling weights, the number of visual material samples, auditory material samples, and tactile feedback samples are calculated using the following formula: Where, represents the number of samples of the i-th type of material, represents the total number of samples corresponding to the proportion of training data, and represents the weight coefficient of the i-th type of material in the initial sampling weight; Based on the number of visual material samples, the number of auditory material samples, and the number of tactile feedback samples, visual material sample data, auditory material sample data, and tactile feedback sample data are extracted from the remaining training dataset to obtain the iterative training dataset (7). The iterative training dataset (7) is removed from the remaining training dataset to obtain the updated training dataset.

6. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 5, characterized in that, The step of training the initial multimodal interaction model using the iterative training dataset to obtain the iterative multimodal interaction model includes: The visual material sample data is subjected to image segmentation processing to obtain segmented image data; The segmented image data is normalized to obtain standard visual data; The standard visual data is used to extract features using a pre-constructed convolutional neural network to obtain a visual feature vector; The auditory material sample data is subjected to spectral conversion to obtain spectral data; The spectral data is subjected to noise reduction processing to obtain standard auditory data; The standard auditory data is sequence-modeled using a pre-constructed recurrent neural network to obtain auditory feature vectors; The tactile feedback sample data is subjected to signal quantization processing to obtain quantized tactile data; The quantized tactile data is smoothed to obtain standard tactile data; The standard tactile data is used to extract features using a pre-built autoencoder to obtain a tactile feature vector; Single-example visual features, single-example auditory features, and single-example tactile features are extracted sequentially from the visual feature vector, auditory feature vector, and tactile feature vector, respectively. The single-instance visual features, single-instance auditory features, and single-instance tactile features are input from the visual input node (1), auditory input node (2), and tactile input node (3) respectively into the initial multimodal interaction model (5) to obtain the iterative output results; Obtain the interactive behavior annotations corresponding to the single-instance visual features, single-instance auditory features, and single-instance tactile features; Extract the target value range of the interactive behavior annotation from the pre-constructed interactive behavior value range table, and extract the center value from the target value range; Calculate the singleton error between the iterative output result and the center value, and determine whether the singleton error is less than the error tolerance threshold; If the singleton error is not less than the error tolerance threshold, then the parameters of the initial multimodal interaction model (5) are adjusted according to the singleton error until the singleton error is less than the error tolerance threshold. Determine whether the visual feature vector, auditory feature vector, and tactile feature vector have completed the extraction of single-example visual features, single-example auditory features, and single-example tactile features; If the extraction is not completed, return to the steps described above for extracting single-example visual features, single-example auditory features, and single-example tactile features from the visual feature vector, auditory feature vector, and tactile feature vector respectively. If the extraction is complete, an iterative multimodal interaction model is obtained.

7. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 6, characterized in that, The performance evaluation of the iterative multimodal interaction model using the validation dataset, to obtain the evaluation error distribution and the maximum evaluation error, includes: The verification dataset (6) is classified to obtain visual verification sample data, auditory verification sample data and tactile verification sample data; The visual verification sample data, auditory verification sample data, and tactile verification sample data are converted into visual verification features, auditory verification features, and tactile verification features, respectively. Single-case visual verification features, single-case auditory verification features, and single-case tactile verification features are extracted sequentially from the visual verification features, auditory verification features, and tactile verification features, respectively. The single-instance visual verification feature, single-instance auditory verification feature, and single-instance tactile verification feature are respectively input into the iterative multimodal interaction model to obtain the verification output results; Determine whether the visual verification features, auditory verification features, and tactile verification features have completed the extraction of single-case visual verification features, single-case auditory verification features, and single-case tactile verification features; If the extraction is not completed, return to the steps described above for extracting single-example visual verification features, single-example auditory verification features, and single-example tactile verification features from the visual verification features, auditory verification features, and tactile verification features respectively. If the extraction is complete, the verification output results are summarized to obtain the verification output result set; The verification output result set is classified according to the interaction behavior annotation data to obtain multiple sets of verification result subsets; Extract the verification result subsets sequentially from the multiple sets of verification result subsets, and identify the target value domain center value corresponding to the verification result subsets; Calculate the error set between the subset of verification results and the center value of the target value range, and calculate the mean error based on the error set to obtain the mean error distribution; Statistical analysis is performed on the error mean distribution to obtain the evaluation error distribution. The maximum error mean is extracted from the error mean distribution and used as the maximum evaluation error.

8. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 1, characterized in that, The step of extracting core interactive information from the multimodal cultural and creative real-time data based on the response output includes: Based on the response output, visual keywords, auditory key segments, and tactile key signals are extracted from a pre-constructed interaction behavior-multimodal mapping table. Based on the visual keywords, auditory key segments, and tactile key signals, core visual elements, core auditory segments, and core tactile signals are extracted from the real-time visual data, real-time auditory data, and real-time tactile data, respectively. The core visual elements, core auditory segments and core tactile signals are summarized to obtain the core interactive information (9).

9. The interactive method for cultural and creative products based on the fusion of multiple sensory experiences as described in claim 8, characterized in that, Before extracting visual keywords, auditory key segments, and tactile key signals from a pre-constructed interaction behavior-multimodal mapping table based on the response output, the method further includes: The interaction behavior categories are extracted sequentially from the interaction behavior annotation data; Receive visual keywords, auditory key segments, and tactile key signals set by the user for the interactive behavior category; An interactive behavior-multimodal mapping table is constructed based on the interactive behavior categories and the visual keywords, auditory key segments, and tactile key signals.

10. An interactive system for cultural and creative products based on the integration of multi-sensory experiences, characterized in that, The system includes: An initial multimodal interaction model training module is used to acquire multimodal cultural and creative training data, wherein the multimodal cultural and creative training data includes visual material training data, auditory material training data, tactile feedback training data, and interactive behavior annotation data; according to a preset verification data ratio, a verification dataset (6) is extracted from the multimodal cultural and creative training data to obtain the remaining training dataset; an initial multimodal interaction model (5) is constructed, wherein the initial multimodal interaction model (5) is designed based on a deep neural network architecture; according to a preset training data ratio and initial sampling weights, an iterative training dataset (7) is extracted from the remaining training dataset to obtain an updated training dataset; the initial multimodal interaction model (5) is trained using the iterative training dataset (7) to obtain an iterative multimodal interaction model; The iterative multimodal interaction model evaluation module is used to evaluate the performance of the iterative multimodal interaction model using the validation dataset (6) to obtain the evaluation error distribution and the maximum evaluation error. The target multimodal interaction model acquisition module is used to determine whether the maximum evaluation error is less than a preset error tolerance threshold. If the maximum evaluation error is not less than the error tolerance threshold, the remaining training dataset, the initial multimodal interaction model (5), and the initial sampling weights are adjusted by updating the training dataset, iterating the multimodal interaction model, and the evaluation error distribution, respectively, and the steps of extracting the iterative training dataset (7) from the remaining training dataset according to the preset training data ratio and the initial sampling weights are returned. If the maximum evaluation error is less than the error tolerance threshold, the iterative multimodal interaction model is used as the target multimodal interaction model (8). The core interactive information extraction module is used to receive multimodal cultural and creative real-time data, generate interactive responses using the target multimodal interactive model (8) based on the multimodal cultural and creative real-time data, and obtain response output; and extract core interactive information (9) from the multimodal cultural and creative real-time data based on the response output.

Citation Information

Patent Citations

  • A method for protecting intangible cultural heritage related to traditional sports

    CN110069829B

  • A method and device for displaying cultural and creative products based on AR

    CN117576355B