Intelligent Museum Virtual Interaction Method and System for Integrating Multimodal Data
By combining self-attention mechanism with clustering, combined with cross-attention mechanism to optimize feature cluster matching between modals, the deviation problem of multimodal data fusion and alignment in the virtual display of the Smart Museum is solved, and a higher quality virtual interaction experience is achieved.
Patent Information
- Application Number
- CN202510339343.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Prior Art In virtual displays of smart museums, cross-modal attention mechanisms may lead to excessive crossover, affecting the quality of multimodal data fusion and alignment.
The self-attention mechanism is used to analyze the feature set of each modal, cluster it through local correlation indicators, and optimize the feature cluster matching between modals with the cross attention mechanism, reduce excessive crossover effects, and improve the quality of data fusion.
Through the combination of self-attention mechanism and clustering, the precise fusion of local features is captured, and the cross-attention mechanism is used to correct the matching between modes, improving the alignment accuracy of multimodal data and the authenticity of virtual interaction experience.
Smart Images

Figure CN119848793B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data fusion, and specifically relates to a virtual interaction method and system for a smart museum that fuses multi-modal data. Background Art
[0002] As a core node for cultural display and dissemination in a smart city, the smart museum is facing new challenges and requirements. With the development of technology, the smart museum is no longer just an extension of the digital museum, but is based on multi-modal perception "data" to establish a comprehensive and in-depth interconnection system, eliminate information silos, and realize a systematic collaborative working mode among people, people and things, and things and things. The development model of the smart museum has changed from technology-driven to demand-driven and business-led, forming a joint force by reconstructing the correlation relationship of each element of the museum, and strengthening the coordination of museum services, protection and management work. In addition, the smart museum also provides a two-way and multi-dimensional information interaction channel among "things, people, and data", realizing intelligent adaptive control and optimization of museum services, protection and management.
[0003] There are two main technical points in the multi-modal data fusion of the smart museum, namely alignment and fusion. The existing technology usually obtains the multi-modal data that needs to be displayed for the operation of the smart museum, specifically image data, text data and voice data, and uses machine learning and deep learning methods to obtain the global and local semantic features of the image data, the global and local semantic features of the text data, and the global and local semantic features of the voice data, and uses deep learning to achieve multi-modal alignment, automatically discovering and learning the corresponding relationships between different modalities.
[0004] However, currently, the cross-modal attention-based multi-modal data fusion makes cross-modal attention judgments on the overall data. When applied to a smart museum, during the virtual display and interaction of the museum, usually the exhibits are displayed, and at the same time, multi-angle demonstrations of the exhibits are carried out. Therefore, using the existing cross-modal attention mechanism may lead to an excessive cross situation when performing the cross-modal attention mechanism, resulting in a certain degree of fusion and alignment deviation, and affecting the final data fusion quality. Summary of the Invention
[0005] In order to solve the technical problem that in the existing technology, an excessive cross situation occurs when performing the cross-modal attention mechanism, resulting in a certain degree of fusion and alignment deviation and affecting the final data fusion quality, the purpose of the present invention is to provide a virtual interaction method and system for a smart museum that fuses multi-modal data. The specific technical solutions adopted are as follows:
[0006] According to one aspect of an embodiment of the present invention, a smart museum virtual interaction method for fusing multi-modal data is provided. The method includes:
[0007] Obtain the feature vectors in each type of modal data and form a feature set. The modalities include images, texts, and audios.
[0008] Based on the self-attention mechanism, analyze the degree of association between the feature vectors in the feature set of each modality and other local feature vectors, and obtain the local association index of each feature vector; cluster the feature vectors in each modality according to the local association index to obtain feature clustering clusters.
[0009] Based on the cross-attention mechanism, analyze the feature vector correlation matching degree between each feature clustering cluster in each modality and each feature clustering cluster in other modalities, and determine the most matching clustering cluster of each feature clustering cluster.
[0010] Perform multi-modal fusion alignment according to the most matching clustering clusters of the feature clustering clusters between different modalities.
[0011] In the above solution, considering the global alignment difference of the cross-attention mechanism when processing multi-modal data, first, the self-attention mechanism is combined with clustering division to calculate the self-attention mechanism for the feature set in a single modality. Through the attention focus degree association of local feature vectors, the change of the content features of the identified items is reflected, and clustering division is performed to better capture local details and improve the precise fusion of local features in different modalities. At the same time, there will be an excessive cross-influence of redundant information parts between multi-modal data. The cross-attention mechanism is used to further correct and match the feature clustering clusters between different modalities to obtain the most matching clustering clusters to optimize the information alignment between modalities, improve the accuracy of fine-grained information matching, and use the matching relationship between different modalities for multi-modal data fusion to reduce inconsistency. The present invention combines the self-attention mechanism with clustering segmentation to improve the precise recognition of local features, uses the cross-modal attention mechanism to correct and match the clusters of different modalities, reduces the situation of excessive crossing, improves the quality of data fusion, and provides a richer and more realistic virtual interaction experience.
[0012] Further, the method for obtaining the local association index includes:
[0013] In the feature set of each modality, obtain the self-attention score of each feature vector and each other feature vector through the self-attention mechanism, and the self-attention weight of each feature vector in the feature set where it is located.
[0014] For any feature vector in each modal feature set, use the feature vectors within a preset local range in the feature set where the feature vector is located as the local feature vectors of the feature vector; obtain the local correlation degree of the feature vector according to the relative degree between the distribution of the self-attention scores of the feature vector and the local feature vectors and the self-attention weights;
[0015] Combine the local correlation degree of the feature vector and the self-attention weights to obtain the local association index of the feature vector.
[0016] In the above solution, by means of the self-attention scores representing the relationship with other feature vectors and the features representing the contribution to the whole, combined with the limitation of the local range, the local association index of the feature vector is obtained, and the self-attention weights are adjusted to obtain the local association index, which more precisely quantifies the importance and mutual relationship of the feature vector in its local context, can capture more detailed correlation relationships, reduce the computational redundancy of global analysis from local relationships, improve the robustness of the feature vector in correlation analysis, and make the subsequent feature clustering division more accurate.
[0017] Furthermore, the method for obtaining the local correlation degree includes:
[0018] Successively use each local feature vector as the target vector, calculate the projection of the self-attention weight of the target vector on the self-attention weight of the feature vector to obtain the relative attention degree of the target vector;
[0019] Take the product of the self-attention scores of the feature vector and the target vector and the relative attention degree as the correlation degree between the feature vector and the target vector;
[0020] Take the mean value of the correlation degrees between the feature vector and all local feature vectors as the local correlation degree of the feature vector.
[0021] Furthermore, the method for obtaining the feature clustering clusters includes:
[0022] For any modality, take the set composed of the local correlation degrees corresponding to the feature vectors in the modality as the association set; determine the number of clusters for the association set of the modality by the elbow method;
[0023] Cluster the feature vectors according to the magnitudes of the local correlation degrees of the feature vectors in the modality to obtain the number of clusters of clustering clusters as the feature clustering clusters of the modality.
[0024] In the above solution, considering that the change of the subject will simultaneously cause a large change in the correlation of local data features, determine the number of clustering segments through the change of the local correlation degree, and cluster the feature vectors, so that each clustering cluster corresponds to more accurate local detailed features, providing a data basis for subsequent local fine matching.
[0025] Further, the method for obtaining the most matching clustering cluster includes:
[0026] Obtain the cross-attention scores between each feature vector in each modality and each feature vector in other modalities through the cross-attention mechanism, as well as the cross-attention weights of each feature vector in the feature set of other modalities; for any one modality, sequentially take each other modality as the cross-target modality;
[0027] For any feature vector in this modality, obtain the matching degree between this feature vector and each feature clustering cluster in the cross-target modality according to the cross-attention scores and cross-attention weights between this feature vector and the feature vectors in each feature clustering cluster in the cross-target modality, and the cross-attention weights of the feature vectors in each feature clustering cluster in the cross-target modality in this modality;
[0028] Combine the matching degrees between all feature vectors in the feature clustering cluster where this feature vector is located and each feature clustering cluster in the cross-target modality to obtain the clustering matching index between the feature clustering cluster where this feature vector is located and each feature clustering cluster in the cross-target modality; based on the clustering matching indexes between each feature clustering cluster in this modality and the feature clustering clusters in the cross-target modality, determine the most matching clustering cluster of each feature clustering cluster in this modality.
[0029] In the above solution, by combining the cross-attention mechanism with the matching of feature clustering clusters, analyze the cross-attention scores and cross-attention weights between feature vectors in each modality and feature clustering clusters in other modalities, and analyze the mutual correlation between feature vectors and feature clustering clusters. From the correlation between single feature vectors and feature clustering clusters in each modality, and then comprehensively considering the feature vectors, obtain the matching relationship between clustering clusters, improve the accuracy and robustness of cross-modal matching analysis, and thus determine the most matching clustering cluster situation of each feature clustering cluster in other modalities. Taking the most matching clustering cluster situation of feature clustering clusters, when dealing with information involving multiple modalities, it can avoid the mismatch phenomenon of clustering in different modalities, enable each cluster in each modality to find the most suitable match in the multi-modal space, and thus achieve a more refined clustering relationship result.
[0030] Further, the method for obtaining the matching degree includes:
[0031] For any feature clustering cluster in the cross-target modality, calculate the dot product of the cross-attention weight of this feature vector in the cross-target modality and the cross-attention weights of each feature vector in this feature clustering cluster in this modality as the similarity weight between this feature vector and each feature vector in this feature clustering cluster;
[0032] Calculate the product of the cross-attention score between this feature vector and each feature vector in this feature clustering cluster and the similarity weight to obtain the cross-similarity between this feature vector and each feature vector in this feature clustering cluster;
[0033] Combining the cross-similarity between the feature vector and all the feature vectors in the feature clustering cluster, the matching degree between the feature vector and the feature clustering cluster is obtained.
[0034] Further, the method for obtaining the cluster matching index includes:
[0035] For any feature clustering cluster in the cross-target modality, taking the feature clustering cluster where the feature vector is located as the analysis clustering cluster;
[0036] Calculating the mean of the matching degrees between all the feature vectors in the analysis clustering cluster and this feature clustering cluster, and taking it as the cluster matching index between the analysis clustering cluster and this feature clustering cluster.
[0037] Further, the method for determining the most matching clustering cluster of each feature clustering cluster in this modality according to the cluster matching index between each feature clustering cluster in this modality and the feature clustering clusters in the cross-target modality includes:
[0038] Successively taking each feature clustering cluster in this modality as the target clustering cluster;
[0039] When the cluster matching index between the target clustering cluster and the feature clustering clusters in the cross-target modality is the largest, taking the corresponding feature clustering cluster in the cross-target modality as the most matching clustering cluster of the target clustering cluster in the cross-target modality.
[0040] Further, the multi-modal fusion alignment according to the most matching clustering clusters between the feature clustering clusters of different modalities includes:
[0041] For any modality, performing data cross-matching between the most matching clustering clusters of each feature clustering cluster in this modality in each other modality to obtain the cross-matching result between this modality and other modalities; performing data alignment and fusion according to the cross-matching result.
[0042] In the above solution, by finding the most matching clustering clusters for cross-modal cross-matching, restricting the accuracy range of cross-matching, reducing the interference of redundant information or noise, reducing the impact of over-crossing caused by mis-matching, and improving the accuracy of data alignment and fusion.
[0043] According to another aspect of the embodiments of the present invention, there is provided an intelligent museum virtual interaction system for fusing multi-modal data to implement the above-mentioned intelligent museum virtual interaction method for fusing multi-modal data. The system includes:
[0044] A data acquisition module, configured to acquire the feature vectors in each modality data and form a feature set, where the modalities include images, texts, and audios;
[0045] The unimodal division module is used to analyze the correlation degree between the feature vectors in the feature set of each modality and the local other feature vectors based on the self-attention mechanism, and obtain the local correlation index of each feature vector; cluster the feature vectors in each modality according to the local correlation index to obtain feature clustering clusters;
[0046] The multimodal matching module is used to analyze the relevant matching degree of the feature vectors between each feature clustering cluster in each modality and each feature clustering cluster in other modalities based on the cross-attention mechanism, and determine the most matching clustering cluster of each feature clustering cluster in each modality in other modalities;
[0047] The multimodal fusion module is used to perform multimodal fusion alignment according to the most matching clustering clusters of the feature clustering clusters between different modalities.
[0048] According to another aspect of the embodiments of the present invention, a computer device is provided.
[0049] A computer device, the computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of a smart museum virtual interaction method for fusing multimodal data as described above are implemented.
[0050] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided.
[0051] A computer-readable storage medium, the computer-readable storage medium stores a computer program, the computer program can be executed by at least one processor, and when the computer program code runs on the processor of the computer, the computer executes the steps of a smart museum virtual interaction method for fusing multimodal data as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 A flowchart of a smart museum virtual interaction method for fusing multimodal data provided by an embodiment of the present invention;
[0054] Figure 2 A flowchart of a method for obtaining the most matching clustering cluster provided by an embodiment of the present invention;
[0055] Figure 3The structural diagram of a smart museum virtual interaction system that integrates multi-modal data provided by an embodiment of the present invention;
[0056] Figure 4 The structural schematic diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0057] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the drawings and preferred embodiments to specifically describe the specific implementation manners, structures, features and effects of a smart museum virtual interaction method and system that integrates multi-modal data according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0059] The following specifically describes the specific solutions of a smart museum virtual interaction method and system that integrates multi-modal data provided by the present invention with reference to the drawings.
[0060] The feature extraction of multi-modal data refers to extracting feature vectors that can represent its information from data of different modalities. These feature vectors can then be used for model training and prediction. The feature extraction method depends on the specific modality and task. For example: Visual modality: Use a convolutional neural network (CNN) to extract image features. Text modality: Use word embeddings or pre-trained language models such as BERT to extract text features. Audio modality: Extract acoustic features such as Mel-frequency cepstral coefficients (MFCC).
[0061] In multi-modal learning, the key to feature extraction lies in how to effectively represent and fuse information from different modalities so that the model can capture the complex relationships and interactions between different modalities. Through the self-attention mechanism and cross-modal attention mechanism, the model can better understand and utilize these features, improving the performance of multi-modal tasks.
[0062] Multi-modal data fusion alignment refers to integrating data from different modalities, such as vision, text, and audio, at the feature level in multi-modal learning so that the model can simultaneously understand and utilize information from these different sources. The purpose of alignment is to solve the correspondence problems of different modality data in time and semantics, which is the basis of multi-modal fusion.
[0063] Methods for multi-modal data fusion alignment include: Early Fusion: Merge data of different modalities at the feature extraction stage and then input it into the model for processing. Late Fusion: Extract features and learn from data of different modalities separately, and then perform fusion at the decision-making layer or a higher layer.
[0064] The Cross-Attention Mechanism is a very important mechanism in multi-modal learning. It focuses on exploring the relationships between different modalities and is usually used in scenarios that require modeling associations between two different sequences, such as multi-modal learning (images and text), machine translation (encoder and decoder), generation tasks, etc.
[0065] The self-attention mechanism, also known as the internal attention mechanism, is an attention mechanism that associates different positions of a single sequence to calculate the representation of the same sequence. The core idea of the self-attention mechanism is to enable the model to consider the information of all positions in the sequence when processing each element in the sequence, thereby capturing long-range dependencies within the sequence. This mechanism does not depend on the absolute positions of the elements in the sequence but realizes it by calculating the relative relationships between the elements.
[0066] The main steps are to generate queries (Q), keys (K), and values (V) for each element in a sequence. The model calculates the similarity or matching degree between each query and all keys to obtain an attention score matrix, and normalizes the attention scores through the softmax function to obtain the attention weights for each element.
[0067] Example 1:
[0068] The embodiment of the present invention provides a smart museum virtual interaction method for fusing multi-modal data. Please refer to Figure 1 , which shows a flowchart of a smart museum virtual interaction method for fusing multi-modal data provided by an embodiment of the present invention. The method includes the following steps:
[0069] S1: Obtain the feature vectors in each modality data and form a feature set. The modalities include images, text, and audio.
[0070] In the virtual interaction system of a smart museum, the acquisition of multi-modal data is the basis for building an immersive experience. The embodiment of the present invention mainly involves data of three modalities: images, text, and audio, aiming to achieve a richer and more realistic virtual interaction experience by integrating these data sources.
[0071] Multi-modal data feature extraction is a prerequisite and key step for achieving multi-modal alignment and enhancement. Machine learning and deep learning methods can be used to extract features from three types of modal data, namely images, text, and audio, respectively, to form a feature set to represent global and local semantic features for subsequent analysis.
[0072] In the embodiments of the present invention, image feature extraction mainly includes: for image data, a convolutional neural network (CNN) is used for preliminary feature extraction by capturing local features of the image, such as edges, textures, etc. To further enhance the object detection and recognition capabilities, a region proposal network (RPN) is introduced to generate candidate regions, and a feature pyramid network (FPN) is combined to extract multi-scale features. Finally, the local features in the image are transformed into feature vectors to form a feature set of the image modality.
[0073] Text feature extraction mainly includes: for text data, first, word embedding techniques, such as Word2Vec or BERT, are used to convert words into vector representations to capture the semantic relationships between words. Then, a Transformer-based sentence-level encoder is adopted to further extract the semantic features of key sentences and phrases in the text to generate a set of feature vectors of the text.
[0074] Audio feature extraction mainly includes: for audio data, first, the audio signal is converted into a frequency-domain representation by Mel-frequency cepstral coefficients (MFCC) to capture the frequency features of the audio. Subsequently, local features in the audio, such as the speech patterns of specific words or phrases, are further extracted and represented as feature vectors to form a feature set of the audio modality.
[0075] Features are extracted from image, text, and audio data respectively, and the feature representations of each modality are expressed as vector sets, providing a basis for subsequent multi-modal analysis. It should be noted that the techniques used for local feature extraction are all well-known to those skilled in the art and will not be elaborated here.
[0076] S2: Analyze the correlation degree between the feature vectors in the feature set of each modality and other local feature vectors based on the self-attention mechanism to obtain the local correlation index of each feature vector; cluster the feature vectors in each modality according to the local correlation index to obtain feature clusters.
[0077] In a smart museum, the displayed content is usually the exhibit data in the museum and the historical, cultural, and other humanistic information related to the exhibits. Among the multi-modal data of different exhibits, there may be relatively similar components, such as shapes and backgrounds. When performing multi-modal data fusion in a smart museum, it is necessary to align the multi-modal data to provide a better visual and interactive experience.
[0078] Meanwhile, there will be a large information gap between the multi-angle information of the exhibits, resulting in a certain information divide. When using the cross-attention mechanism to fuse the multi-modal data of museum exhibits, it will cause a certain attention cross deviation, affecting the fusion and alignment effects of the multi-modal data.
[0079] Therefore, the multi-modal data of the smart museum can be clustered and segmented. By segmenting the whole of the multi-modal data through technologies such as clustering and self-attention mechanism, the cross-modal attention calculation of the multi-modal data can be restricted within a certain range, improving the accuracy of the cross-modal attention between subsequent multi-modalities, and thus improving the final fusion and alignment accuracy of the multi-modalities.
[0080] The self-attention mechanism can obtain relatively accurate and good results within a small data set. However, when dealing with long data, due to its neglect of the positional relationship in the data, its attention scores on the long time axis are distorted. At the same time, when the smart museum is displaying, when the displayed information changes greatly, the attention scores of the feature data in the single-modal data are usually low. Therefore, clustering can be performed within the single-modal data, and the clustering or data segmentation within the single-modal can be completed by making the attention scores within a single cluster more concentrated.
[0081] The self-attention mechanism can capture the mutual relationship between different feature vectors, effectively evaluate the relationship between the feature vector and other features, identify local potential associations through local analysis, and then optimize the clustering results of the features. Therefore, first analyze the degree of association between the feature vectors in the feature set of each modality and other local feature vectors based on the self-attention mechanism, obtain the local association index of each feature vector, and evaluate each feature vector through the attention association of the local features.
[0082] Preferably, in the embodiment of the present invention, the method for obtaining the local association index includes:
[0083] The self-attention mechanism can obtain the scores of each element in the set, that is, the attention weights, which reflect the degree of attention of the model to other elements in the set when processing the current element. In the virtual interaction system of the smart museum, there is usually one or more main display objects, and these main bodies are manifested as significant elements in the overall or local features of the multi-modal data. Since the self-attention mechanism can automatically assign weights, these main display objects usually show relatively high self-attention scores in the relevant feature data of each modality. Therefore, first calculating the self-attention mechanism for the single-modal data helps to identify and highlight these main display objects.
[0084] First, in the feature set of each modality, the self-attention scores of each feature vector with every other feature vector and the self-attention weights of each feature vector in its own feature set are obtained through the self-attention mechanism. Calculating the self-attention scores through the self-attention mechanism represents the degree of relationship between each feature vector and other vectors, and further obtaining the self-attention weights represents the contribution of each feature vector to the overall feature set. It should be noted that the methods for the self-attention mechanism to obtain the self-attention scores and self-attention weights are well-known technical means to those skilled in the art and will not be elaborated here.
[0085] In the self-attention mechanism, the self-attention weight of each element is a vector because each element calculates the similarity with all other elements. The self-attention weight contains the global attention distribution of an element to the entire set. For example, assuming the length of the set is n, then for an input element, the self-attention weight after Softmax processing is an n-dimensional vector, and each dimension represents the attention allocation of this input element to the n elements in the set.
[0086] Furthermore, for any feature vector in the feature set of each modality, the feature vectors within a preset local range in its own feature set are used as the local feature vectors of this feature vector, and the local analysis range is determined. In the embodiments of the present invention, the preset local range is set to a range with a radius of 20 data lengths to the left and right of the feature vector as the center. When the range limits can be taken on both the left and right of the feature vector, the number of local feature vectors of the feature vector is 40. The specific range value can be adjusted by the implementer according to the specific implementation scenario and is not limited here.
[0087] Furthermore, according to the relative degree between the distribution of the self-attention scores and the self-attention weights of this feature vector and the local feature vectors, the local correlation degree of this feature vector is obtained. By the relative attention situation and the correlation situation, the accuracy of local feature representation is improved. In the embodiments of the present invention, the method for obtaining the local correlation degree includes:
[0088] First, each local feature vector is sequentially used as the target vector, and the projection of the self-attention weight of the target vector on the self-attention weight of this feature vector is calculated to obtain the relative attention degree of the target vector. The projection maps the self-attention weight distribution of the target vector onto the current feature vector, and calculates the attention degree of the target vector relative to the current feature vector.
[0089] Furthermore, the product of the self-attention score of this feature vector and the target vector and the relative attention degree is used as the correlation degree between this feature vector and the target vector. The correlation relationship is reflected through the self-attention score, and weighted by the relative attention degree, which can more accurately measure the relationship between features and avoid the inaccuracy of simple self-attention score measurement.
[0090] Finally, the average of the correlations between this feature vector and all local feature vectors is used as the local correlation of this feature vector, synthesizing the association between this feature vector and local feature vectors. As an example, the expression for local correlation is:
[0091] ; where represents the local correlation of the -th feature vector, represents the total number of local feature vectors of the -th feature vector, represents the self-attention score between the -th feature vector and the -th local feature vector, represents the relative attention degree between the -th local feature vectors, represents the correlation between the -th feature vector and the -th local feature vector.
[0092] Finally, combining the local correlation of this feature vector and the self-attention weight, the local association index of this feature vector is obtained. By combining the correlation of local analysis with the attention situation of this feature vector in global analysis, the local association characteristics of this feature vector are comprehensively characterized. In the embodiments of the present invention, the product of the local correlation of this feature vector and the self-attention weight is used as the local association index of this feature vector.
[0093] Furthermore, the feature set in a single modality can be clustered and segmented based on local relevance. When the main content changes, the local relevance will change significantly. Therefore, different local feature parts are segmented by clustering for subsequent analysis and recognition. Preferably, in the embodiments of the present invention, the method for obtaining feature clustering clusters includes:
[0094] For any modality, the set composed of the local correlations corresponding to the feature vectors in this modality is used as the association set. By analyzing the distribution change of local correlations, the change situation between local features can be reflected. The elbow method is used to determine the number of clusters for the association set of this modality. It should be noted that the elbow method is a commonly used method for determining the number of clusters, which is a well-known technical means for those skilled in the art and will not be elaborated here.
[0095] Furthermore, the feature vectors are clustered according to the local correlation degree of the feature vectors of the modality, and the number of clusters obtained is used as the feature clusters of the modality. In the embodiments of the present invention, the K-means clustering algorithm is used for clustering. When clustering, the seed points are evenly distributed in the data, and the local correlation degree is used as the distance metric to obtain the clusters. The clustering algorithm is a well-known technical means for those skilled in the art and will not be limited herein.
[0096] So far, each single modality has been clustered and divided, and the subsequent analysis range is limited by the clusters.
[0097] S3: Based on the cross-attention mechanism, analyze the correlation matching degree of the feature vectors between each feature cluster in each modality and each feature cluster in other modalities, and determine the most matching cluster of each feature cluster in each modality in other modalities.
[0098] Through the cross-attention mechanism, similar features and semantic relationships in different modalities can be accurately aligned. The feature clusters of each modality can be matched according to their correlation with the clusters in other modalities, thereby ensuring semantic consistency in clustering. By determining the most relevant matching clusters of the feature clusters of each modality in other modalities, the system can avoid over-fusing irrelevant information, thereby reducing the risk of overfitting. Therefore, through the participation of cross-attention in the matching of clusters, the multi-modal information fusion becomes more efficient and compact.
[0099] Preferably, in the embodiments of the present invention, for the method of obtaining the most matching cluster, please refer to Figure 2 , which shows a flowchart of a method for obtaining the most matching cluster provided by an embodiment of the present invention. The method includes the following steps:
[0100] S301: Through the cross-attention mechanism, obtain the cross-attention scores between each feature vector in each modality and each feature vector in other modalities, and the cross-attention weights of each feature vector in the feature set of other modalities; for any one modality, each other modality is sequentially used as the cross-target modality.
[0101] The cross-attention mechanism is an extension based on the self-attention mechanism. It allows the model to directly calculate the attention degree of each feature vector in one modality to all feature vectors in another modality when processing multi-modal data. In this way, the model can understand the mutual correlation between different modalities, rather than just the internal information of a single modality.
[0102] The cross-attention score reflects the correlation degree between feature vectors, and the cross-attention weight represents the relative importance of a feature vector in another modality in the cross-modal feature space. Analyze each other modality of this modality in turn to perform the matching of feature clusters between modalities.
[0103] S302: For any eigenvector in this modality, according to the cross-attention scores and cross-attention weights between this eigenvector and the eigenvectors in each feature clustering cluster in the cross-target modality, and the cross-attention weights of the eigenvectors in each feature clustering cluster in the cross-target modality in this modality, obtain the matching degree between this eigenvector and each feature clustering cluster in the cross-target modality.
[0104] Combining the attention situation between the eigenvector and the feature clustering cluster for which the matching degree is analyzed among the eigenvectors, and the global attention situation in the mutual feature sets, evaluate the matching degree between the eigenvector and the feature clustering cluster, and characterize the overall relationship between the eigenvector and the feature clustering cluster.
[0105] Preferably, in the embodiments of the present invention, the method for obtaining the matching degree includes:
[0106] First, for any feature clustering cluster in the cross-target modality, calculate the dot product of the cross-attention weight of this eigenvector in the cross-target modality and the cross-attention weights of each eigenvector in this feature clustering cluster in this modality, as the similarity weight between this eigenvector and each eigenvector in this feature clustering cluster. The dot product result of the cross-attention weights reflects the similarity between the two eigenvectors at the modality level, and characterizes the potential attention similarity degree between these eigenvectors.
[0107] Furthermore, calculate the product of the cross-attention score between this eigenvector and each eigenvector in this feature clustering cluster and the similarity weight, to obtain the cross-similarity between this eigenvector and each eigenvector in this feature clustering cluster. The cross-attention score reflects the overall similarity degree between the eigenvector and all vectors in the cross-target modality. Combining the similarity weight to adjust the similarity situation between the eigenvectors can more finely reflect the similarity relationship.
[0108] Finally, combining the cross-similarities between this eigenvector and all eigenvectors in this feature clustering cluster, obtain the matching degree between this eigenvector and this feature clustering cluster. In the embodiments of the present invention, take the mean value of the cross-similarities between this eigenvector and all eigenvectors in this feature clustering cluster as the matching degree between this eigenvector and this feature clustering cluster, which comprehensively reflects the overall correspondence situation between the eigenvector and the feature clustering cluster.
[0109] S303: Combining the matching degrees between all eigenvectors in the feature clustering cluster where this eigenvector is located and each feature clustering cluster in the cross-target modality, obtain the cluster matching index between the feature clustering cluster where this eigenvector is located and each feature clustering cluster in the cross-target modality; according to the cluster matching indices between each feature clustering cluster in this modality and the feature clustering clusters in the cross-target modality, determine the most matching clustering cluster for each feature clustering cluster in this modality.
[0110] By combining the relationship between the matching degree of a single eigenvector in a clustering cluster and other feature clustering clusters, the matching relationship between clustering clusters is obtained. Preferably, in the embodiments of the present invention, the method for obtaining the clustering matching index includes:
[0111] For any feature clustering cluster in the cross-target modality, use the feature clustering cluster where the eigenvector is located as the analysis clustering cluster. Calculate the mean value of the matching degrees of all eigenvectors in the analysis clustering cluster with this feature clustering cluster as the clustering matching index between the analysis clustering cluster and this feature clustering cluster. Obtain the matching between clustering clusters by analyzing the matching situation between all eigenvectors in the analysis clustering cluster and this feature clustering cluster.
[0112] Furthermore, the feature clustering cluster with the best matching situation in other modalities for each feature clustering cluster can be screened out through the matching situation between clustering clusters. According to the clustering matching index between each feature clustering cluster in this modality and the feature clustering clusters in the cross-target modality, determine the best matching clustering cluster for each feature clustering cluster in this modality, including:
[0113] Successively use each feature clustering cluster in this modality as the target clustering cluster. When the clustering matching index between the target clustering cluster and the feature clustering clusters in the cross-target modality is the largest, use the corresponding feature clustering cluster in the cross-target modality as the best matching clustering cluster of the target clustering cluster in the cross-target modality. Determine the feature clustering cluster with the best matching situation in other modalities through the maximum clustering matching index of the target clustering cluster.
[0114] So far, the relevant matching at the local fine-grained level is completed, and the matching relationship of the divided clusters between different modalities is obtained.
[0115] S4: Perform multi-modal fusion alignment according to the best matching clustering clusters of the feature clustering clusters between different modalities.
[0116] By cross-matching the feature clustering clusters in one modality with the best matching clustering clusters in other modalities, the similarity and potential associations between different modalities can be captured more precisely. The alignment of each clustering cluster in different modalities is optimized not only based on local features but also through global matching, thereby improving the accuracy of cross-modal matching.
[0117] In the embodiments of the present invention, for any modality, perform data cross-matching between the best matching clustering clusters of each feature clustering cluster in this modality in each other modality to obtain the cross-matching result between this modality and other modalities, and perform data alignment and fusion according to the cross-matching result.
[0118] Improve the expression ability of multimodal information through more precise alignment and fusion, and enhance the interactive experience of the intelligent museum between different exhibits and visitors. For example, the color and shape features in the image modality can be more precisely combined with the historical background and cultural connotations in the text modality, making the expression of exhibit information more complete and vivid. Or when the visitor mentions the name of an artist or a historical period, the system can more accurately identify the relevant image and text information according to the voice command, quickly switch the exhibit or provide a detailed background explanation.
[0119] In summary, the present invention considers the global alignment difference of the cross-attention mechanism in processing multimodal data. First, by combining the self-attention mechanism with clustering division, the self-attention mechanism is calculated for the feature set in the single modality. Through the attention degree correlation of the local feature vectors, the change of the content features of the identified items is reflected, and clustering division is performed to better capture local details and improve the precise fusion of local features in different modalities. At the same time, there will be an excessive cross-influence of redundant information parts between multimodal data. The cross-attention mechanism is used to further correct and match the feature clustering clusters between different modalities to obtain the most matching clustering cluster to optimize the information alignment between modalities, improve the accurate matching of fine-grained information, and perform multimodal data fusion using the matching relationship between different modalities to reduce inconsistency. The present invention combines the self-attention mechanism with clustering segmentation to improve the accurate recognition of local features, uses the cross-modal attention mechanism to correct and match the clusters of different modalities, reduces the excessive cross situation, and improves the data fusion quality to provide a richer and more realistic virtual interaction experience.
[0120] Embodiment 2:
[0121] An embodiment of the present invention provides an intelligent museum virtual interaction system for fusing multimodal data. Please refer to Figure 3 , which shows the structural diagram of an intelligent museum virtual interaction system for fusing multimodal data provided by an embodiment of the present invention. The system 400 includes: a data acquisition module 401, a single modality division module 402, a multimodal matching module 403, and a multimodal fusion module 404.
[0122] The data acquisition module 401 is used to acquire the feature vectors in each modality data and form a feature set, and the modalities include images, texts, and audios;
[0123] The single modality division module 402 is used to analyze the association degree between the feature vectors in the feature set of each modality based on the self-attention mechanism, and obtain the local association index of each feature vector; cluster the feature vectors in each modality according to the local association index to obtain feature clustering clusters;
[0124] The multimodal matching module 403 is used to analyze the correlation matching degree of feature vectors between each feature clustering cluster in each modality and each feature clustering cluster in other modalities based on the cross-attention mechanism, and determine the most matching clustering cluster of each feature clustering cluster in each modality in other modalities.
[0125] The multimodal fusion module 404 is used to perform multimodal fusion alignment according to the most matching clustering clusters of feature clustering clusters between different modalities.
[0126] It should be noted that the system provided in the above embodiment is only illustrated by the division of the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. Since the specific implementation process of a smart museum virtual interaction system for fusing multimodal data in this embodiment is the same as the specific implementation process of a smart museum virtual interaction method for fusing multimodal data described above, it will not be elaborated here.
[0127] Embodiment 3:
[0128] The embodiment of the present invention also provides a computer device. Please refer to Figure 4 , which shows the structural schematic diagram of a computer device provided by an embodiment of the present invention. The computer device 500 includes at least, but is not limited to, a memory 501, a processor 502, and a communication interface 503 that can be communicatively connected to each other through a bus 510.
[0129] The memory 501 may include a large-capacity memory for data or instructions, and at least includes one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a read-only memory, a magnetic disk, or an optical disc, etc. The memory 501 can be inside or outside the computer device. In the embodiment of the present invention, the memory 501 can be a read-only memory for storing computer program instructions.
[0130] The processor 502 may include one or more integrated circuits for controlling the overall operation of the computer device 500. In the embodiment of the present invention, the processor 502 reads and executes the computer program instructions stored in the memory 501 to implement Figure 1 the steps of a smart museum virtual interaction method for fusing multimodal data in Embodiment 1 shown.
[0131] The communication interface 503 is used to establish a communication connection between the computer device 500 and other electronic devices. In the embodiment of the present invention, the communication interface 503 is used to connect the computer device 500 to an external terminal through a network, and establish a data transmission channel and a communication connection between the computer device 500 and the external terminal.
[0132] Example 4:
[0133] The embodiment of the present invention further provides a computer-readable storage medium, such as a flash memory, a hard disk, an optical disc, or a server, etc., on which computer program instructions are stored, and when the program instructions are executed by a processor, corresponding functions are implemented. The computer-readable storage medium in the embodiment of the present invention is used to store computer program instructions for a virtual interaction method of an intelligent museum that fuses multi-modal data, and the steps of the above-mentioned virtual interaction method of an intelligent museum that fuses multi-modal data are executed by a processor.
[0134] It should be noted that: the above sequence of embodiments of the present invention is only for description and does not represent the advantages or disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0135] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A virtual interaction method for a smart museum that integrates multi-modal data, characterized in that, The method includes: Obtaining feature vectors in each modality data and forming a feature set, where the modalities include images, text, and audio; Analyzing the degree of association between the feature vectors in the feature set of each modality and other local feature vectors based on the self-attention mechanism to obtain the local association index of each feature vector; clustering the feature vectors in each modality according to the local association index to obtain feature clustering clusters; Analyzing the degree of relevant matching between the feature vectors of each feature clustering cluster in each modality and those of each feature clustering cluster in other modalities based on the cross-attention mechanism to determine the most matching clustering cluster of each feature clustering cluster in each modality in other modalities; Performing multi-modal fusion alignment according to the most matching clustering clusters of feature clustering clusters between different modalities; The method for obtaining the most matching clustering cluster includes: Obtaining the cross-attention scores between each feature vector in each modality and each feature vector in other modalities, and the cross-attention weights of each feature vector in the feature set of other modalities through the cross-attention mechanism; for any one modality, sequentially taking each other modality as the cross-target modality; For any feature vector in this modality, according to the cross-attention scores and cross-attention weights between this feature vector and the feature vectors in each feature clustering cluster in the cross-target modality, and the cross-attention weights of the feature vectors in each feature clustering cluster in the cross-target modality in this modality, obtain the matching degree between this feature vector and each feature clustering cluster in the cross-target modality; Combining the matching degrees between all the feature vectors in the feature clustering cluster where this feature vector is located and each feature clustering cluster in the cross-target modality to obtain the clustering matching index between the feature clustering cluster where this feature vector is located and each feature clustering cluster in the cross-target modality; determining the most matching clustering cluster of each feature clustering cluster in this modality according to the clustering matching index between each feature clustering cluster in this modality and the feature clustering clusters in the cross-target modality.
2. The method for virtual interaction in a smart museum integrating multi-modal data according to claim 1, wherein, The method for obtaining the local association index includes: In the feature set of each modality, obtaining the self-attention scores between each feature vector and each other feature vector, and the self-attention weights of each feature vector in the feature set where it is located through the self-attention mechanism; For any feature vector in the feature set of each modality, taking the feature vectors within a preset local range in the feature set where this feature vector is located as the local feature vectors of this feature vector; according to the relative degree between the distribution of the self-attention scores of this feature vector and the local feature vectors and the self-attention weights, obtain the local correlation degree of this feature vector; Combining the local correlation degree of this feature vector and the self-attention weights to obtain the local association index of this feature vector.
3. The intelligent museum virtual interaction method for fusing multi-modal data according to claim 2, characterized in that, The method for obtaining the local correlation degree includes: Sequentially taking each local feature vector as the target vector, calculating the projection of the self-attention weight of the target vector on the self-attention weight of this feature vector to obtain the relative attention degree of the target vector; Taking the product of the self-attention score between this feature vector and the target vector and the relative attention degree as the correlation degree between this feature vector and the target vector. The average of the correlation degrees between the feature vector and all local feature vectors is used as the local correlation degree of the feature vector.
4. The intelligent museum virtual interaction method for fusing multi-modal data according to claim 1, characterized in that, The method for obtaining the feature clustering clusters includes: For any modality, the set composed of the local correlation degrees corresponding to the feature vectors in this modality is used as the correlation set; the elbow method is used to determine the number of clusters for the correlation set of this modality. The feature vectors are clustered according to the magnitudes of the local correlation degrees of the feature vectors in this modality, and the number of clusters obtained is used as the feature clustering clusters of this modality.
5. The intelligent museum virtual interaction method for fusing multi-modal data according to claim 1, wherein The method for obtaining the matching degree includes: For any feature clustering cluster in the cross-target modality, calculate the dot product of the cross-attention weight of the feature vector in the cross-target modality and the cross-attention weight of each feature vector in this feature clustering cluster in this modality as the similarity weight between the feature vector and each feature vector in this feature clustering cluster. Calculate the product of the cross-attention score between the feature vector and each feature vector in this feature clustering cluster and the similarity weight to obtain the cross-similarity between the feature vector and each feature vector in this feature clustering cluster. Combine the cross-similarities between the feature vector and all feature vectors in this feature clustering cluster to obtain the matching degree between the feature vector and this feature clustering cluster.
6. The intelligent museum virtual interaction method for fusing multi-modal data according to claim 1, characterized in that, The method for obtaining the cluster matching index includes: For any feature clustering cluster in the cross-target modality, use the feature clustering cluster where the feature vector is located as the analysis clustering cluster. Calculate the average of the matching degrees between all feature vectors in the analysis clustering cluster and this feature clustering cluster as the cluster matching index between the analysis clustering cluster and this feature clustering cluster.
7. The intelligent museum virtual interaction method for fusing multi-modal data according to claim 1, characterized in that Determining the most matching clustering cluster of each feature clustering cluster in this modality according to the cluster matching index between each feature clustering cluster in this modality and the feature clustering clusters in the cross-target modality includes: Successively use each feature clustering cluster in this modality as the target clustering cluster. When the cluster matching index between the target clustering cluster and the feature clustering clusters in the cross-target modality is the largest, use the corresponding feature clustering cluster in the cross-target modality as the most matching clustering cluster of the target clustering cluster in the cross-target modality.
8. The intelligent museum virtual interaction method for fusing multi-modal data according to claim 1, wherein Performing multi-modal fusion alignment according to the most matching clustering clusters between feature clustering clusters of different modalities includes: For any modality, perform data cross-matching between the most matching clustering clusters of each feature clustering cluster in this modality in each other modality to obtain the cross-matching result between this modality and other modalities; perform data alignment and fusion according to the cross-matching result.
9. A smart museum virtual interaction system integrating multi-modal data, which is used to implement the method for a smart museum virtual interaction integrating multi-modal data according to any one of claims 1-8, characterized in that, The system includes: A data acquisition module for acquiring the feature vectors in the data of each modality and forming a feature set, where the modalities include images, texts, and audios. A single-modal partitioning module for analyzing the correlation degree between the feature vectors in the feature set of each modality and local other feature vectors based on the self-attention mechanism to obtain the local correlation index of each feature vector; clustering the feature vectors in each modality according to the local correlation index to obtain feature clustering clusters. A multi-modal matching module for analyzing the relevant matching degree of the feature vectors between each feature clustering cluster in each modality and each feature clustering cluster in other modalities based on the cross-attention mechanism to determine the most matching clustering cluster of each feature clustering cluster in each modality in other modalities. A multimodal fusion module for performing multimodal fusion alignment based on the most matching cluster among the feature clusters between different modalities.
Citation Information
Patent Citations
Method and device for generating information
CN111783808A
Urban underground space environment monitoring method and system based on Internet of Things
CN118898733A
Multi-modal fusion method based on cross attention gating unit
CN119046863A