Data interaction display method and system for digital exhibition hall

By using a visual language attribute recognition network and an interactive emotion recognition network, combined with particle swarm optimization methods, the problem of inaccurate linkage between exhibit semantics and audience preferences in digital exhibition halls was solved, adaptive adjustment of exhibit displays and personalized recommendations were achieved, and the user experience was improved.

CN120670607AInactive Publication Date: 2025-09-19NANJING YOUQIUBIYIN INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510785354.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing digital exhibition hall systems find it difficult to achieve accurate linkage between exhibit semantics and audience preferences, resulting in unfocused recommendations and poor user experience.

Method used

By obtaining the text information and image data of cultural relics, using the visual language attribute recognition network and interactive emotion recognition network, the semantic features of cultural relics and the audience's voice semantics, image behavior and emotional feedback are extracted, and the semantic tendency index, behavioral adsorption index and emotional response index are generated. The user interaction preference index is calculated by combining the particle swarm optimization method, and the exhibition content is adjusted to achieve precise matching.

Benefits of technology

It improves the ability to accurately match the semantics of cultural relics with user interests, enhances user immersion and information acquisition efficiency, and realizes adaptive switching and personalized recommendations of exhibit displays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670607A_ABST
    Figure CN120670607A_ABST
Patent Text Reader

Abstract

The invention discloses a data interaction display method and system for a digital exhibition hall, and relates to the technical field of data interaction display. The data interaction display method for the digital exhibition hall comprises the steps of obtaining text information of each cultural relic in a to-be-displayed cultural relic restoration exhibition hall and cultural relic display image data of each angle, inputting the text information and the cultural relic display image data into a pre-trained cultural relic attribute recognition model to extract an attribute set, and constructing a cultural relic exhibition model; inputting the attribute set into a cultural relic restoration display model, calculating a display recommendation index, generating a preference response evaluation set in combination with browsing video stream data and a user perception model, and calculating a user interaction preference index; and finally fusing the display recommendation index and the user interaction preference index to obtain an exhibition item response index, the cultural relics of the cultural relic restoration display model are displayed and adjusted through the exhibition item response index, so that the accurate matching capability between the semantic meaning of the cultural relics and the interest of the user is effectively improved, the immersion of the user is enhanced, and the information acquisition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data interactive display, and in particular to a data interactive display method and system for a digital exhibition hall. Background Art

[0002] With the advancement of the digitalization of cultural heritage, cultural relics restoration exhibition halls, as an important carrier for spreading the concept of cultural relics restoration and displaying the results of cultural relics restoration, have gradually introduced digital exhibition hall systems to achieve more efficient, interactive and immersive display methods. Existing digital exhibition halls mostly rely on static graphic display boards or preset three-dimensional models to display exhibits. The exhibition content is usually organized in linear sorting or category division, which makes it difficult to dynamically adjust according to the visiting preferences and real-time interactive behaviors of different audiences.

[0003] In the current display method, the display information of cultural relics mainly relies on manual annotation and static description, and lacks the ability to automatically extract the semantic attributes of cultural relics from multi-angle images and text information, resulting in coarse granularity of display content and a single form of interaction. In addition, the audience's interactive behaviors in the exhibition hall, such as voice expression, behavioral trajectory, gaze state and emotional feedback, have not yet been effectively perceived and analyzed. The display system cannot respond to the interests of different audiences, thus limiting the exhibition hall's intelligence level and personalized recommendation capabilities.

[0004] Existing technology, such as the patent application with publication number CN119781891A, discloses a data interactive display method and system for a digital exhibition hall. The method includes: collecting historical display interactive behavior data from multiple display devices in the exhibition hall, extracting user browsing behavior preference characteristics to determine multiple browsing behavior preference patterns; extracting browsing time period preference characteristics and determining time period distribution patterns to generate a first interactive display strategy; collecting historical display environment data from multiple display devices and performing content environment association analysis to construct a content environment analysis model and generate a second interactive display strategy; fusing and generating a target interactive display strategy and performing interactive display optimization; obtaining real-time display interactive behavior data from display devices and performing user behavior collaborative impact analysis to generate multiple interactive collaborative optimization strategies for interactive display collaborative optimization. The present invention achieves dynamic and personalized display content optimization.

[0005] Based on the above solution, it is found that the limitations of the existing technology include at least the following problems. The existing technology lacks a fine-grained and dynamic linkage mechanism between the semantic structure of cultural relics and the real preferences of the audience, resulting in a lack of accuracy and responsiveness in the display content, which in turn limits the intelligent recommendation and interactive guidance capabilities of the digital exhibition hall. For example, when a user stops to observe multiple exhibits in an exhibition area with the theme of the cultural relic restoration process, the existing technology only records their browsing path and time, but it is difficult to identify the user's interest due to exquisite patterns or coordinated colors, and it is also difficult to capture the cognitive burden caused by obscure structure or severe damage. The lack of this preference information directly leads to the loss of focus of subsequent display sorting and recommendation results, making it difficult to prioritize the exhibits that truly resonate with the user's emotion or cognitive interest, thereby affecting the user's immersive experience and information acquisition efficiency. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention provides a data interactive display method and system for a digital exhibition hall, which solves the problem that the existing technology is difficult to achieve accurate linkage between exhibit semantics and audience preferences, resulting in out-of-focus recommendations and poor experience.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a data interactive display method for a digital exhibition hall, comprising the following steps: obtaining text information of each cultural relic to be displayed in a cultural relic restoration exhibition hall and cultural relic display image data at each angle, and inputting them into a pre-trained cultural relic attribute recognition model for recognition analysis to obtain an attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and constructing a cultural relic exhibition model; inputting the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model, and performing feature analysis to obtain a display recommendation index for each cultural relic in the cultural relic restoration display model; obtaining browsing video stream data of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and performing comprehensive analysis in combination with a user perception model to obtain a preference response evaluation set for each cultural relic to be displayed in the cultural relic restoration exhibition hall; inputting the preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model for data analysis to obtain a user interaction preference index for each cultural relic in the cultural relic restoration display model, and performing comprehensive analysis in combination with the display recommendation index to obtain an exhibition item response index for each cultural relic in the cultural relic restoration display model; and performing display adjustment on each cultural relic in the cultural relic restoration display model based on the exhibition item response index.

[0008] Furthermore, the cultural relic display image data is specifically the pixel value of each pixel point in the cultural relic display image, the cultural relic attribute set includes a component set and a display indicator set, the cultural relic attribute recognition model is specifically a visual language attribute recognition network, and the visual language attribute recognition network includes an input layer, a feature encoding layer, a multi-modal fusion layer, and an output layer.

[0009] Furthermore, the specific steps for obtaining the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: in the input layer of the visual language attribute recognition network, the text information of each cultural relic to be displayed in the cultural relic restoration exhibition hall and the cultural relic display image data at each angle are received and preprocessed; in the feature encoding layer of the visual language attribute recognition network, the preprocessed text information of each cultural relic to be displayed in the cultural relic restoration exhibition hall and the cultural relic display image data at each angle are respectively subjected to feature extraction processing to obtain the text feature vector and image feature vector set of each cultural relic to be displayed in the cultural relic restoration exhibition hall; in the multi-mode fusion layer of the visual language attribute recognition network, the text feature vector and image feature vector set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are fused to obtain the joint semantic representation vector of each cultural relic to be displayed in the cultural relic restoration exhibition hall; in the output layer of the visual language attribute recognition network, the joint semantic representation vector of each cultural relic to be displayed in the cultural relic restoration exhibition hall is parsed to obtain the component component set and display index set of each cultural relic to be displayed in the cultural relic restoration exhibition hall; the display index set includes restoration intervention index, visual coherence index, cognitive load index, display medium dependence index, color attractiveness index, and structural hierarchy index.

[0010] Furthermore, the specific steps for obtaining the display recommendation index of each cultural relic in the cultural relic restoration and display model are as follows: read the display index set in the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and perform comprehensive analysis on each of them to obtain the display evaluation index set of each cultural relic to be displayed in the cultural relic restoration exhibition hall, including the visible load index and the exhibition experience index; and perform comprehensive analysis on the display evaluation index set of each cultural relic to be displayed in the cultural relic restoration exhibition hall to obtain the display recommendation index of each cultural relic in the cultural relic restoration exhibition hall to be displayed, and input it into the cultural relic restoration display model for embedding processing to obtain the display recommendation index of each cultural relic in the cultural relic restoration display model.

[0011] Furthermore, the specific formula for calculating the display recommendation index of a cultural relic to be displayed in the cultural relic restoration exhibition hall is as follows: Among them, ZsT is the display recommendation index of a certain cultural relic in the cultural relic restoration exhibition hall to be displayed, TyZ is the exhibition experience index of a certain cultural relic in the cultural relic restoration exhibition hall to be displayed, η1 is the experience adjustment coefficient stored in the database, XgF is the explicit load index of a certain cultural relic in the cultural relic restoration exhibition hall to be displayed, η2 is the explicit adjustment coefficient stored in the database, and η3 is the interaction adjustment coefficient stored in the database.

[0012] Furthermore, the browsing video stream data is specifically voice information and several frames of browsing image data, the browsing image data is specifically the browsing pixel value and browsing two-dimensional coordinates of each browsing pixel point in the browsing image, and the user perception model is specifically an interactive emotion recognition network, which includes a multi-modal input layer, a user recognition layer, a modal encoding layer, a modal fusion layer, and an emotion feature decoding output layer.

[0013] Furthermore, the specific steps of obtaining the preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: in the multi-modal input layer of the interactive emotion recognition network, the browsing video stream data of each cultural relic to be displayed in the cultural relic restoration exhibition hall is received and preprocessed; in the user recognition layer of the interactive emotion recognition network, the preprocessed browsing video stream data of each cultural relic to be displayed in the cultural relic restoration exhibition hall is recognized and processed to obtain the interactive modal data of each audience within a set range for each cultural relic to be displayed in the cultural relic restoration exhibition hall; in the modal coding layer of the interactive emotion recognition network, the interactive modal data of each audience within a set range for each cultural relic to be displayed in the cultural relic restoration exhibition hall is feature coded and obtained. The speech emotion feature vector and image interaction feature vector of each audience member within the set range of the cultural relics to be displayed in the cultural relics restoration exhibition hall are obtained; in the modal fusion layer of the interactive emotion recognition network, the speech emotion feature vector and image interaction feature vector of each audience member within the set range of each cultural relic to be displayed in the cultural relics restoration exhibition hall are cross-modally fused to obtain the interactive emotion expression vector of each audience member within the set range of each cultural relic to be displayed in the cultural relics restoration exhibition hall; in the feature decoding output layer of the interactive emotion recognition network, the interactive emotion expression vector of each audience member within the set range of each cultural relic to be displayed in the cultural relics restoration exhibition hall is feature decoded to obtain the semantic tendency index, behavioral adsorption index and emotional response index of each cultural relic to be displayed in the cultural relics restoration exhibition hall, that is, the preference response evaluation set.

[0014] Furthermore, the specific steps for obtaining a preference response evaluation set for each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: based on the particle swarm optimization algorithm, a comprehensive analysis is performed on the semantic tendency index, behavioral adsorption index, and emotional response index of each cultural relic to be displayed in the cultural relic restoration exhibition hall to obtain a user interaction preference index for each cultural relic to be displayed in the cultural relic restoration exhibition hall; the user interaction preference index of each cultural relic in the cultural relic restoration exhibition hall to be displayed is input into the cultural relic restoration display model for embedding processing to obtain a user interaction preference index for each cultural relic in the cultural relic restoration display model.

[0015] Furthermore, the specific formula for calculating the response index of a certain cultural relic in the cultural relic restoration and display model is as follows: Among them, ZxY is the exhibition item response index of a certain cultural relic in the cultural relic restoration and display model, ZsT is the display recommendation index of a certain cultural relic in the cultural relic restoration and display model, μ1 is the recommendation adjustment coefficient stored in the database, HyJ is the user interaction preference index of a certain cultural relic in the cultural relic restoration and display model, μ2 is the interaction preference adjustment coefficient stored in the database, and μ3 is the deviation adjustment coefficient stored in the database.

[0016] A data interactive display system for a digital exhibition hall comprises: a model construction module for acquiring text information of each cultural relic to be displayed in a cultural relic restoration exhibition hall and cultural relic display image data at each angle, and inputting the data into a pre-trained cultural relic attribute recognition model for recognition analysis to obtain an attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and constructing a cultural relic exhibition model; a feature analysis module for inputting the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model, performing feature analysis, and obtaining a display recommendation index for each cultural relic in the cultural relic restoration display model; a user perception analysis module for acquiring browsing video stream data of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and performing comprehensive analysis in combination with the user perception model to obtain a preference response evaluation set for each cultural relic to be displayed in the cultural relic restoration exhibition hall; a comprehensive interaction analysis module for inputting the preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model for data analysis to obtain a user interaction preference index for each cultural relic in the cultural relic restoration display model, and performing comprehensive analysis in combination with the display recommendation index to obtain an exhibition item response index for each cultural relic in the cultural relic restoration display model; and a display feedback module for adjusting the display of each cultural relic in the cultural relic restoration display model based on the exhibition item response index.

[0017] The present invention has the following beneficial effects:

[0018] (1) The data interactive display method for digital exhibition halls extracts the structural semantic features and display index set of each cultural relic through the visual language attribute recognition network, and combines the interactive emotion recognition network to identify the audience's voice semantics, image behavior and emotional feedback, and generates the cultural relic's semantic tendency index, behavior adsorption index and emotional response index respectively. Then, the particle swarm optimization method is used to calculate the user interaction preference index, and then the user interaction preference index is integrated with the display recommendation index to output the final exhibition item response index. The display order of cultural relics in the exhibition system is adjusted accordingly, thereby realizing active response to the audience's focus and adaptive switching of local display content, thereby effectively improving the precise matching ability between cultural relic semantics and user interests, and avoiding the disconnection between display content and user expectations, thereby enhancing user immersion and improving information acquisition efficiency.

[0019] (2) The data interactive display method for digital exhibition halls extracts joint features of the graphic and text information of cultural relics based on the visual language attribute recognition network, uses the position attention mechanism and the graphic and text alignment strategy to identify the structural relationship of the components of cultural relics, and further divides the main structure and the auxiliary structure on the basis of generating the component set, constructing a cultural relic component organization tree with clear semantics and clear hierarchy, and each component is bound to the image attribute and function label at the same time, so as to realize the multi-semantic nested expression for color, texture, material, symbol and other dimensions, and then combines the user browsing behavior to perform personalized structural navigation and focus enhancement, thereby improving the presentation granularity and semantic comprehensibility of the exhibits.

[0020] (3) The data interactive display method used in digital exhibition halls integrates three types of modal information: voice content, emotional tone, and image behavior, to construct an interactive emotion recognition network, which can fully restore the audience's real emotional state and attention behavior in front of specific cultural relics. The network uses the voice modality to extract the semantic tendency index, and also uses the image modality to identify the gaze area and behavior adsorption index, and combines facial expressions and voice energy perception to analyze the emotional response index, thereby achieving multi-dimensional capture of audience preferences, interests, and emotions, and then effectively identifying fine-grained emotions and physical response relationships such as confusion due to complex patterns or emotional resonance due to bright colors, thereby providing high-precision data support for user preference modeling and emotional resonance-driven display optimization.

[0021] (4) The data interactive display system for digital exhibition halls realizes closed-loop information flow integration from semantic analysis, user perception to exhibit feedback based on a modular structure, thereby forming a dynamic linkage mechanism between graphic data, emotional feedback and display strategy. The system takes the model construction module as the starting point to build a cultural relic exhibition model oriented to graphic semantics, and then obtains a quantifiable display recommendation index through the feature analysis module; the user perception analysis module accesses the video stream and extracts emotional behavior features, which are integrated by the comprehensive interactive analysis module to generate an exhibit response index, and finally the display feedback module performs display updates based on preferences and recommendations, thereby strengthening the task closed-loop capability and response timeliness of the display process, and providing a systematic and logically rigorous interactive linkage execution framework for cultural relic display.

[0022] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The present invention is a flow chart of a data interactive display method for a digital exhibition hall.

[0024] Figure 2The present invention is a flowchart of the specific steps of obtaining the attribute set of each cultural relic to be displayed in a cultural relic restoration exhibition hall in a data interactive display method for a digital exhibition hall.

[0025] Figure 3 The present invention is a block diagram of a data interactive display system for a digital exhibition hall. DETAILED DESCRIPTION

[0026] See also Figure 1 , an embodiment of the present invention provides a technical solution: a data interactive display method for a digital exhibition hall, comprising the following steps: obtaining text information of each cultural relic in the cultural relic restoration exhibition hall to be displayed and cultural relic display image data at each angle, and inputting them into a pre-trained cultural relic attribute recognition model for recognition analysis to obtain an attribute set of each cultural relic in the cultural relic restoration exhibition hall to be displayed, and constructing a cultural relic exhibition model; inputting the attribute set of each cultural relic in the cultural relic restoration exhibition hall to be displayed into the cultural relic restoration display model, and performing feature analysis to obtain a display recommendation index for each cultural relic in the cultural relic restoration display model; obtaining browsing video stream data of each cultural relic in the cultural relic restoration exhibition hall to be displayed, and performing comprehensive analysis in combination with a user perception model to obtain the cultural relic to be displayed. a preference response evaluation set for each cultural relic in the cultural relic restoration exhibition hall; inputting the preference response evaluation set for each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model for data analysis, obtaining the user interaction preference index of each cultural relic in the cultural relic restoration display model, and performing comprehensive analysis in combination with the display recommendation index to obtain the exhibition item response index of each cultural relic in the cultural relic restoration display model; adjusting the display of each cultural relic in the cultural relic restoration display model based on the exhibition item response index, specifically: arranging the exhibition item response index of each cultural relic in the cultural relic restoration display model in descending order, generating a cultural relic display browsing table, and displaying based on the sequence of the corresponding cultural relics in the cultural relic display browsing table, such as giving priority to displaying the cultural relics in the first sequence in the cultural relic display browsing table.

[0027] The specific formula for calculating the response index of a cultural relic in the cultural relic restoration and display model is as follows: Among them, ZxY is the exhibition item response index of a certain cultural relic in the cultural relic restoration and display model, ZsT is the display recommendation index of a certain cultural relic in the cultural relic restoration and display model, μ1 is the recommendation adjustment coefficient stored in the database, HyJ is the user interaction preference index of a certain cultural relic in the cultural relic restoration and display model, μ2 is the interaction preference adjustment coefficient stored in the database, and μ3 is the deviation adjustment coefficient stored in the database.

[0028] It should be explained that μ1, μ2, and μ3 can be obtained through the following steps: based on historical data, determine the initial impact weights of each variable (display recommendation index, user interaction preference index) on the exhibit response index through statistical regression analysis, and then use the sensitivity analysis method to adjust the value range of the coefficient to evaluate the stability and applicability of these parameters to the formula output. Next, further fit the weights through model optimization (such as machine learning algorithms or multi-objective optimization) to ensure that the formula can accurately reflect the actual exhibition status of cultural relics.

[0029] Among them, the specific steps of constructing the cultural relics exhibition model are as follows: according to the information of each component part contained in the component set, each component part is logically divided according to functional attributes and structural relationships to form a structural hierarchical model containing a main structure and an auxiliary structure. The main structure is used to express the core components of the cultural relics (such as the drum head, drum belly, drum bottom, etc. of the bronze drum), and the auxiliary structure is used to express the extended components (such as buttons, decorative bands, etc.), forming a clearly structured cultural relics component organization tree. Secondly, the image attributes (including main color, texture type, pattern style) and semantic labels (including functional role, material category, cultural symbol, etc.) corresponding to each component part are attached to the corresponding cultural relics to construct a multi-dimensional attribute expression at the part level, and generate a cultural relics exhibition model for realizing the refined display and semantic visualization of cultural relics in the digital exhibition hall, which serves as the core semantic model to drive the interactive display of the cultural relics in the digital exhibition hall.

[0030] The cultural relics display image data is specifically the pixel value of each pixel in the cultural relics display image. The cultural relics attribute set includes a component set and a display indicator set. The cultural relics attribute recognition model is specifically a visual language attribute recognition network (a fusion of SwinTransformer and BERT). The visual language attribute recognition network includes an input layer, a feature encoding layer, a multi-modal fusion layer, and an output layer.

[0031] Among them, the component set is a set of structured component part information.

[0032] The input layer is used to receive and preprocess the text information and multi-angle image data of cultural relics to generate a standardized feature input format.

[0033] The feature encoding layer is used to extract the semantic features of text information and the spatial visual features of image data respectively, and output text feature vectors and image feature vector sets.

[0034] The multimodal fusion layer semantically aligns and fuses the text feature vector with the image feature vector set to generate a joint semantic representation vector.

[0035] The output layer is used to parse the joint semantic representation vector and extract the component set and display indicator set.

[0036] Specifically, if Figure 2As shown, the specific steps of obtaining the attribute set of each cultural relic in the cultural relic restoration exhibition hall to be displayed are as follows: in the input layer of the visual language attribute recognition network, the text information of each cultural relic in the cultural relic restoration exhibition hall to be displayed and the cultural relic display image data at each angle are received, and preprocessed (i.e., the image data is size normalized and channel normalized, and the text information is word segmentation encoded and sequence filled to generate a standardized input format that can be processed by the feature extraction module);

[0037] In the feature encoding layer of the visual language attribute recognition network, the pre-processed text information of each cultural relic to be displayed in the cultural relic restoration exhibition hall and the cultural relic display image data at each angle are respectively subjected to feature extraction processing (that is, the text information of each cultural relic is input into the language encoding sub-network stored in the database, and the language encoding sub-network adopts the pre-trained BERT structure, and encodes the input text based on the Transformer mechanism, including word segmentation, position encoding and multi-layer attention calculation of the input cultural relic description text, so as to construct the grammatical dependency, contextual semantic association and keyword distribution characteristics in the cultural relic description, and finally extract a fixed-length text semantic feature vector as the semantic representation of the cultural relic in the language modality, and at the same time, the cultural relic display images of multiple angles corresponding to the cultural relic are respectively input into the visual encoding sub-network stored in the database, and the visual encoding sub-network adopts the Swin Transformer structure, and through the hierarchical local window self-attention mechanism, the spatial hierarchical information, edge texture features, geometric configuration features, etc. of the image can be extracted by Hierarchical Window-based Self-Attention, and combined with Patch Embedding operations and multi-scale feature aggregation mechanisms achieve comprehensive expression of images at the detail and structure levels. Each image ultimately outputs a fixed-length image feature vector that reflects the deep visual semantics of the image content at that perspective. After merging the image feature vectors corresponding to all perspective images, a multi-angle image feature vector set of the cultural relic is formed. This yields a text feature vector and image feature vector set for each cultural relic to be displayed in the cultural relic restoration exhibition hall.

[0038] In the multimodal fusion layer of the visual language attribute recognition network, the text feature vector and image feature vector set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are fused (that is, the image feature vector set is aggregated, that is, the image feature vectors of each perspective are merged based on maximum pooling, specifically, the maximum response value on each dimension is selected to form a fusion vector with the strongest feature response, and a visual fusion vector is obtained. Next, the text feature vector is used as the query vector, and the visual fusion vector is used as the key and value, and input into the cross-modal attention module, that is, the Cross-Modal AttentionModule. In this module, the similarity score between the query vector and the key is calculated, and the attention weights of different feature dimensions in the image modality are dynamically adjusted, so that the final fusion representation is more in line with the visual area and structural information concerned by the text semantics. In this process, a multi-head attention mechanism such as Multi-Head can be used. Attention is used to improve the alignment of different semantic channels between images and text, and residual connections and layer normalization are used to ensure stable information transmission. Ultimately, a joint semantic representation vector is output, which combines the semantic information of the cultural relic in the text modality with the spatial structure characteristics of the image modality. This vector is used to obtain the joint semantic representation vector of each cultural relic to be displayed in the cultural relic restoration exhibition hall.

[0039] In the output layer of the visual language attribute recognition network, the joint semantic representation vector of each cultural relic to be displayed in the cultural relic restoration exhibition hall is parsed (for the component set, the joint semantic representation vector is input into the composition structure analysis module, which partitions and decodes the visual space dimension in the fusion representation based on the position attention mechanism and feature aggregation mechanism, extracts several potential component area representations, and combines the structural description semantics contained in the text for feature alignment. Through image-text matching and regional attention allocation, it is determined which significant components the cultural relic consists of, and for each component, its image attribute vector and semantic description vector are further extracted, including visual attributes such as main color, texture type, pattern style, and functional role, such as Supporting parts, sound-generating parts, cultural symbols, such as the sun, dragon patterns, Bagua, etc., material categories and other semantic labels are finally output, that is, a set of structured component part information sets. For the display indicator set, the sub-feature vector related to restoration is extracted in the joint semantic representation vector based on the attention mechanism. The sub-vector mainly contains the semantic information of the abnormal area of ​​surface texture in the image modality, and the semantic channels activated by keywords such as restoration, splicing, and recasting in the text. Then, the sub-vector is subjected to feature projection and nonlinear decoding operations to extract the intermediate features of potential restoration area, restoration description strength, semantic confidence and other dimensions. Finally, the restoration intervention index of the cultural relic is obtained by combining the normalization functions, which reflects its original preservation degree. The attention mechanism is used to extract sub-feature vectors related to the structural hierarchy from the joint semantic representation vector. The sub-vector mainly responds to the semantic dimensions of structural features such as spatial nesting, component separation, and axial symmetry in the image modality. At the same time, it combines the activation response of structural description keywords such as multi-layer, loop, segmentation, mosaic, upper and lower parts that appear in the text modality, and performs spatial relationship modeling and structural complexity perception projection to extract intermediate structural features such as the number of component levels, symmetry score, and segmentation logic. Finally, after nonlinear normalization mapping, a structural hierarchy index reflecting the structural complexity of the cultural relic is output to assist in interactive design and hierarchical display decisions during display. The sub-vector related to the overall composition coordination of the image is extracted from the joint semantic representation vector. , this sub-vector integrates the potential feature expressions of the image modality in terms of composition structure, element arrangement order, style continuity and other dimensions, and responds to the activation of semantic keyword channels such as symmetrical layout, integrated molding, unified style, and seamless decoration in the text description. Then, spatial consistency modeling and style continuity projection processing are performed to extract intermediate representation features such as rhythm balance, composition tightness, and decorative continuity in the semantic image structure. Finally, the visual coherence index is output through a combination of normalized functions. The color feature-guided gating mechanism is used to extract the color-related sub-vector from the joint semantic representation vector. This sub-vector integrates the main color channel response in the image modality, such as HSV space projection, color distribution balance, etc., as well as the color perception vocabulary in the text modality.For example, the semantic embedding of gilt, dark green, and iron black is activated, and then the main color clustering score and user preference semantic similarity are calculated. The relationship between the main color and the aesthetic tendency in the cultural context is further combined, and finally the matching attraction score of the cultural relic color is output, which is recorded as the color attraction index. Based on the semantic abstraction perception mechanism, the cognitive load related sub-vector is extracted from the joint semantic representation vector. This sub-vector is reflected in the image modality as an implicit representation of perceptual burden such as structural complexity, texture density, and pattern irregularity, and in the text modality as a semantic projection response of language complexity indicators such as term density, semantic jumping, and context nonlinearity. Then, language abstraction modeling and visual complexity fusion analysis are performed to obtain the audience's understanding under visual-language fusion. This prediction value ultimately outputs the cognitive load index of the cultural relic. The semantic path attention mechanism extracts a sub-vector related to media adaptation from the joint semantic representation vector. This sub-vector reflects the expression requirements for spatial dimensions such as three-dimensional structure, local details, and hidden areas in the image modality, as well as the activation of semantic features related to descriptions such as rotational observation, inner cavity, decomposed structure, and cross-sectional display in the text modality. Spatial visibility analysis and expression requirement mapping are then performed, ultimately outputting the display medium dependence index. This yields a set of component components and a set of display indicators for each cultural relic to be displayed in the cultural relic restoration exhibition hall. The display indicator set includes the restoration intervention index, visual coherence index, cognitive load index, display medium dependence index, color appeal index, and structural hierarchy index.

[0040] Among them, the restoration intervention index is the degree to which the cultural relics have undergone artificial restoration.

[0041] The visual coherence index refers to the overall visual coordination and style unity of the cultural relics.

[0042] The cognitive load index is the information processing cost required for the audience to understand the content of the cultural relic display.

[0043] The display media dependence index is the degree to which cultural relics rely on enhanced display media (such as AR and 3D decomposition) during the display process.

[0044] The color attraction index is the attractiveness of the cultural relic color to the user's visual preference.

[0045] The structural hierarchy index is the depth and complexity of the hierarchical nesting of the cultural relics structure.

[0046] The pre-training process of the visual language attribute recognition network is as follows: construct a joint image and text annotation dataset containing multi-category cultural relics samples, including multi-angle image sequences and text description information of each cultural relic, and is equipped with annotation fields such as component labels, image attribute labels (such as color, texture, pattern) and semantic labels (such as material, function, cultural meaning), and divide the dataset into attribute training set and attribute verification set.

[0047] The parameters of the visual language attribute recognition network are initialized, where the image encoding module (using the SwinTransformer structure) and the text encoding module (based on pre-trained BERT) are loaded with weights pre-trained on ImageNet and Chinese corpus (such as Chinese Wikipedia) respectively to improve the model's modeling foundation for visual structure and language semantics.

[0048] During the training phase, multi-task training is performed using image-text alignment tasks and attribute parsing tasks. The number of training rounds is set (e.g., 120 rounds), and the following steps are performed in each round of training: first, the input image is channel-normalized and patch-divided, and the input text is segmented and position-encoded. Subsequently, feature vectors are extracted through the visual encoder and language encoder respectively, and input into the multimodal fusion module for semantic alignment, and a joint semantic representation vector is output; then, the matching loss (e.g., contrast loss of image-text similarity), component decoding loss (e.g., cross entropy), and display indicator regression loss (e.g., mean square error) are calculated respectively. The three losses are combined by weight to form a total loss function, and the AdamW optimizer is used for parameter update. At the same time, strategies such as learning rate warm-up and decay, Dropout, and residual regularization are combined to control the training dynamics and improve the robustness of training.

[0049] After each round of training, the text-image matching accuracy, attribute recognition accuracy, and MAE values ​​of various display indicators are calculated on the validation set, and loss and performance curves are plotted to dynamically evaluate the model convergence and generalization performance. If the validation indicators do not improve for several consecutive rounds, the Early Stopping mechanism is triggered and training is terminated early.

[0050] Finally, the converged model parameters are exported into a deployment format for subsequent loading and calling in the digital exhibition hall system to support the structured analysis of cultural relics graphic information and the extraction of display indicators.

[0051] In this implementation, the dual feature encoding of text and image is combined to break through the limitations of a single modality. For example, text can explain structural semantics, and images can capture visual details. The fusion of the two can more accurately analyze the physical characteristics and cultural connotations of cultural relics, thereby avoiding attribute misjudgment caused by one-sided data. Secondly, through the cross-modal attention module and multi-head attention mechanism, the alignment and enhancement of image and text features are achieved. For example, in the calculation of the restoration intervention index, the splicing traces of the text description and the texture abnormality areas in the image can be automatically focused, thereby improving the interpretability and accuracy of the indicator generation. Finally, the extracted display indicators are objectively weighted through the entropy weight method, and abstract concepts such as the restoration status of cultural relics and display needs are converted into quantifiable explicit load index and exhibition experience index, providing a data-driven basis for curatorial priority sorting and media resource allocation, and reducing the risk of subjective experience decision-making.

[0052] Specifically, the specific steps for obtaining the display recommendation index of each cultural relic in the cultural relic restoration and display model are as follows: read the display index set in the attribute set of each cultural relic in the cultural relic restoration exhibition hall to be displayed, and conduct comprehensive analysis respectively to obtain the display evaluation index set of each cultural relic in the cultural relic restoration exhibition hall to be displayed, including the explicit load index (used to evaluate the technical cost, deconstruction difficulty and media dependence required for the cultural relic to be perceived and expressed completely, clearly and structuredly in the digital exhibition hall) and the exhibition experience index; and conduct comprehensive analysis on the display evaluation index set of each cultural relic in the cultural relic restoration exhibition hall to be displayed to obtain the display recommendation index of each cultural relic in the cultural relic restoration exhibition hall to be displayed. And input it into the cultural relics restoration and display model for embedding processing (that is, the display recommendation index of the cultural relic is taken as a structured attribute field and written into the data entity of the corresponding cultural relic in the cultural relics restoration and display model. Specifically, in the cultural relics restoration and display model, a unique identification number is assigned to each cultural relic to be displayed, and the one-to-one correspondence between the recommendation index and the target cultural relic entity is determined by index matching with the cultural relic number field in the display recommendation index output structure. Then, in the data structure of the cultural relics restoration and display model, an attribute field is extended to define the display recommendation index for each cultural relic object, and it is a numerical type that can be represented by a floating point), and the display recommendation index of each cultural relic in the cultural relics restoration and display model is obtained.

[0053] Among them, the specific steps for obtaining the display evaluation index set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: based on the entropy weight method, the restoration intervention index, structural hierarchy index, and display media dependence index of each cultural relic to be displayed in the cultural relic restoration exhibition hall are weighted (and in the weighted processing process, the weight coefficients corresponding to the restoration intervention index, structural hierarchy index, and display media dependence index are obtained based on the entropy weight method, which is as follows: for all the cultural relics to be displayed in the cultural relic restoration exhibition hall, their corresponding restoration intervention index, structural hierarchy index, and display media dependence index are collected respectively, and the three indicators have been normalized to the same The numerical interval constitutes an indicator matrix. Each row in the matrix represents a cultural relic sample, and each column represents a certain type of indicator value. Secondly, the information entropy value of each indicator is calculated. For each column of indicators, the proportion of each cultural relic under the indicator is first calculated, that is, the normalized score of the cultural relic under the indicator is divided by the sum of the normalized scores of all cultural relics in the column. At the same time, a very small constant is introduced to avoid the denominator being zero. Subsequently, the proportion value is multiplied by its own natural logarithm, and then the sum of such product results of all cultural relics in the column is multiplied by a negative normalization coefficient, which is the inverse of the total number of samples under the natural logarithm. Finally, The result is the information entropy value of the indicator. The larger the information entropy value, the smaller the discrimination of the indicator in the sample; the smaller the entropy value, the greater the difference between the indicator in different cultural relics. Again, according to the obtained information entropy value, the entropy weight coefficients corresponding to the three indicators are calculated. The specific method is to first subtract one from the information entropy value of each indicator to obtain its corresponding difference value, and then add up all the difference values, and divide the difference of each indicator by the sum to obtain the normalized weight value. Finally, three weight coefficients are obtained, which correspond to the restoration intervention index, the structural hierarchy index and the display media dependence index, reflecting these three indicators. The proportion of its importance in the overall visible structural evaluation) is used to obtain the visible structural load index of each cultural relic in the cultural relics restoration exhibition hall to be displayed; the visual coherence index, cognitive load index, and color attractiveness index of each cultural relic in the cultural relics restoration exhibition hall to be displayed are weighted based on the entropy weight method (and in the weighted processing process, the weight coefficients corresponding to the visual coherence index, cognitive load index, and color attractiveness index are obtained based on the entropy weight method, and its acquisition logic is consistent with the weight coefficients corresponding to the restoration intervention index, structural hierarchy index, and display media dependence index), and the visible structural load index of each cultural relic in the cultural relics restoration exhibition hall to be displayed is obtained.

[0054] The specific formula for calculating the display recommendation index of a cultural relic to be displayed in a cultural relic restoration exhibition hall is as follows: Among them, ZsT is the display recommendation index of a certain cultural relic in the cultural relic restoration exhibition hall to be displayed, TyZ is the exhibition experience index of a certain cultural relic in the cultural relic restoration exhibition hall to be displayed, η1 is the experience adjustment coefficient stored in the database, XgF is the explicit load index of a certain cultural relic in the cultural relic restoration exhibition hall to be displayed, η2 is the explicit adjustment coefficient stored in the database, and η3 is the interaction adjustment coefficient stored in the database.

[0055] What needs to be explained is that the specific expression of the tanh function is: Here, e is a natural constant and can be 2.71 in this embodiment, with a domain of (-∞, +∞) and a range of (-1, +1).

[0056] η1, η2, and η3 can be obtained through the following steps: using historical data, combined with the exhibition experience index and the visible load index, to conduct statistical regression analysis, quantify the specific impact of each factor on the display recommendation index, and thus fit the initial weight value. Secondly, using the sensitivity analysis method, adjust the value range of each coefficient, observe its impact on the display recommendation evaluation results, and ensure the stability and rationality of the model.

[0057] The specific implementation example of calculating the display recommendation index of a cultural relic to be displayed in the cultural relic restoration exhibition hall is as follows. The following data is available: including the exhibition experience index and visible load index of a cultural relic to be displayed in the cultural relic restoration exhibition hall, as shown in Table 1:

[0058] Table 1 Example of the evaluation index set data for the display of cultural relics sequence in the cultural relics restoration exhibition hall

[0059] Exhibition Experience Index Explicit load index Cultural Relic 1 0.858 0.624 Relics 2 0.924 0.537 Relic 3 0.762 0.754 Relics 4 0.813 0.658 Relic 5 0.649 0.716

[0060] The experience adjustment coefficient η1 stored in the database is approximately: 0.427;

[0061] The explicit structural adjustment coefficient η2 stored in the database is approximately: 0.267;

[0062] The interaction adjustment coefficient η3 stored in the database is approximately: 1.172;

[0063] Substituting the data in Table 1 and the above adjustment coefficient into the specific formula for calculating the display recommendation index of a cultural relic to be displayed in the cultural relic restoration exhibition hall, we obtain:

[0064] The display recommendation index of the first cultural relic to be displayed in the cultural relic restoration exhibition hall = (exp(0.427×0.858) / (1+ln(1+0.624 0.267 )))×(1+tanh(1.172×0.858×0.624))≈1.382;

[0065] The display recommendation index of the second cultural relic to be displayed in the cultural relic restoration exhibition hall = (exp(0.427×0.924) / (1+ln(1+0.537 0.267 )))×(1+tanh(1.172×0.924×0.537))≈1.401;

[0066] The display recommendation index of the third cultural relic to be displayed in the cultural relic restoration exhibition hall = (exp(0.427×0.762) / (1+ln(1+0.754 0.267 )))×(1+tanh(1.172×0.762×0.754))≈1.324;

[0067] The fourth cultural relic display recommendation index of the cultural relic restoration exhibition hall is (exp(0.427×0.813) / (1+ln(1+0.658 0.267 )))×(1+tanh(1.172×0.813×0.658))≈1.342;

[0068] The fifth cultural relic to be displayed in the cultural relic restoration exhibition hall is recommended for display index = (exp(0.427×0.649) / (1+ln(1+0.716 0.267 )))×(1+tanh(1.172×0.649×0.716))≈1.196.

[0069] In this implementation plan, by constructing a display evaluation index set and using the entropy weight method to weight and fuse multiple indicators, a display recommendation index reflecting the characteristics of cultural relic display is generated, thereby achieving quantitative expression and precise control of multi-dimensional factors on display decisions. The explicit load index is based on the degree of restoration intervention, structural complexity, and dependence on display media, thereby comprehensively measuring the technical load and structural analysis difficulty required for the display of cultural relics. The exhibition experience index integrates user perception dimensions such as color appeal, visual coherence, and cognitive load, and can accurately reflect the audience's aesthetic adaptability and understanding difficulty of the exhibits. By embedding these two types of indices into the cultural relic data model, traceable and adjustable structured recommendation parameters are formed, thereby achieving the system's reasonable sorting and personalized recommendation of the display priorities of different exhibits. The weight parameters are extracted through historical data regression and sensitivity analysis, thereby improving the model's adaptability and prediction accuracy, and ensuring the interpretability and stability of the display recommendation mechanism, thereby significantly enhancing the intelligent and scientific level of content scheduling in digital exhibition halls.

[0070] Specifically, the browsing video stream data is specifically voice information and several frames of browsing image data. The browsing image data is specifically the browsing pixel value and browsing two-dimensional coordinates of each browsing pixel point in the browsing image. The user perception model is specifically an interactive emotion recognition network. The interactive emotion recognition network includes a multi-modal input layer, a user recognition layer, a modal encoding layer, a modal fusion layer, and an emotion feature decoding output layer.

[0071] Among them, the multimodal input layer is used to receive voice information and browse image frame data during the display process, and perform standardized preprocessing on the voice and image modalities respectively, providing a unified format input basis for subsequent modality classification and feature extraction.

[0072] The user identification layer is used to identify individual viewers and complete the identity binding of voice and image.

[0073] The modality encoding layer is used to extract semantic, behavioral, and emotional features from speech and images.

[0074] The modal fusion layer is used to fuse multimodal features and generate interactive emotion expression vectors.

[0075] The feature decoding output layer is used to decode and generate semantic tendency, behavioral adsorption and emotional response index from the expression vector.

[0076] The specific steps of obtaining the preference response evaluation set of each cultural relic in the exhibition hall for restoration of cultural relics to be displayed are as follows: in the multi-modal input layer of the interactive emotion recognition network, the browsing video stream data of each cultural relic in the exhibition hall for restoration of cultural relics to be displayed is received and preprocessed; in the user recognition layer of the interactive emotion recognition network, the browsing video stream data of each cultural relic in the exhibition hall for restoration of cultural relics to be displayed after preprocessing is recognized and processed (face detection and human contour recognition operations are performed on each frame image in the browsing video stream, and a multi-target tracking algorithm, such as a trajectory matching method based on deep features, is used to maintain the track of the audience in consecutive frames, and a unique identity code is assigned to each audience member). The number is used to identify each audience individual who appears within the set display time range, and based on the synchronously collected voice signal data, the method of combining sound source direction estimation with voice facial state detection is used to judge the user face in the voice state in the current video frame, and the corresponding voice segment is bound to the user number to achieve accurate correspondence between the voice modal data and the user entity in the image frame. Then, for each audience number, the voice segment and image sequence in the corresponding time period are extracted respectively, and the two are encapsulated into a unified audience interaction modal data, including two parts of voice modal data and image modal data), and each cultural relic to be displayed in the cultural relics restoration exhibition hall is obtained. The interactive modal data of each audience member within the set range; in the modal coding layer of the interactive emotion recognition network, the interactive modal data of each audience member within the set range of each cultural relic to be displayed in the cultural relic restoration exhibition hall is subjected to feature coding processing (the voice modal information in the interactive modal data is input into the voice coding sub-network stored in the database. The voice coding sub-network adopts a one-dimensional convolutional network and a bidirectional recurrent network structure, namely Bi-LSTM, to extract the rhythm frequency, pitch change, and emotional intonation features of the voice; and combines the keyword semantic embedding information of the voice content, such as like, complex, good-looking, etc. to extract the high-dimensional semantic vector and output the audience's emotion expression in the cultural relic. The speech emotion feature vector in front of the audience is obtained; each frame of display browsing image data in the interactive modal data is input into the image coding sub-network stored in the database. The image coding sub-network is based on a convolutional neural network to extract the emotional expression features of the audience's face area, the attention distribution features of the gaze area, and the behavioral direction features of the head or posture. The local attention model is performed on the user's facial area in each frame of the image, and an interactive behavior response map is generated in combination with the pixel coordinate sequence. The image interaction feature vector of the audience in front of the cultural relic is output, and the speech emotion feature vector and image interaction feature vector of each audience member within the set range of each cultural relic to be displayed in the cultural relic restoration exhibition hall are obtained;In the modal fusion layer of the interactive emotion recognition network, the voice emotion feature vector and image interaction feature vector of each audience member within the set range of each cultural relic to be displayed in the cultural relics restoration exhibition hall are cross-modally fused (the voice feature vector is used as the query vector and the image feature vector is used as the key-value pair to dynamically capture the visual interaction area and behavioral feature dimension that the voice expression focuses on, and through the multi-head attention mechanism, multi-layer residual connection and layer normalization operation, the semantic alignment and information complementarity between the two modalities are achieved, and finally a fused multi-modal joint semantic representation vector is output), and the interactive emotion expression vector of each audience member within the set range of each cultural relic to be displayed in the cultural relics restoration exhibition hall is obtained; in the feature decoding output layer of the interactive emotion recognition network, the interactive emotion expression vector of each audience member within the set range of each cultural relic to be displayed in the cultural relics restoration exhibition hall is feature decoded (for the semantic tendency index, based on the preset semantic channel attention mechanism, the high response channels related to factors such as voice content, keyword density, and grammatical structure are identified in the interactive emotion expression vector, and the semantic sub-feature vector is extracted and input into the semantic preference mapping network, and the semantic tendency is extracted through the full connection layer and normalization function, and the output is for each The semantic tendency score of each viewer within the set range of the cultural relic is averaged to obtain a semantic tendency index for each cultural relic, reflecting the intensity of attention and semantic match of the user's voice expression towards the cultural relic. For the behavioral adsorption index, a behavioral preference feature screening function is constructed to identify feature dimensions related to behavioral feedback such as gaze hot spots, dwell time, and operation frequency in the interactive emotional expression vector. The behavioral adsorption sub-feature vectors are extracted and weighted to obtain the behavioral adsorption score of each viewer within the set range for each cultural relic. The average is then processed to obtain the behavioral adsorption index of each cultural relic, which measures the user's interactive concentration and physical attention during the display of the cultural relic. For the emotional response index, feature sub-channels related to the intensity of expression changes, frequency of intonation fluctuations, and emotional energy in the interactive emotional expression vector are extracted and weighted to obtain the emotional response score of each viewer within the set range. The average is then processed to obtain the emotional response index of each cultural relic, which reflects the intensity of emotional resonance and subjective preference response of the user in front of the cultural relic. The semantic tendency index, behavioral adsorption index, and emotional response index of each cultural relic in the cultural relic restoration exhibition hall to be displayed are obtained, namely the preference response evaluation set.

[0077] The pre-training process of the interactive emotion recognition network is as follows:

[0078] A dataset of annotated emotional interaction videos is obtained. The dataset contains audience interaction samples in different scenarios. Each sample consists of synchronously collected speech audio, facial image sequences and their corresponding multi-dimensional annotations. The annotations include: speech content labels (positive / negative / neutral), facial emotion classification (such as joy, surprise, indifference), gaze area heat maps, action behavior classification (such as attention, removal, operation), etc. The annotated emotional interaction video dataset is divided into an emotion training set and an emotion verification set.

[0079] Initialize the interactive emotion recognition network, adopt Kaiming or Xavier initialization strategy for the trainable weight parameters in each module of the model, and load some parameters (such as speech encoder and image expression recognition module) pre-trained on common emotion recognition datasets (such as RAVDESS and AffectNet) to enhance the model's feature extraction ability and convergence speed in multimodal emotion modeling.

[0080] Multiple rounds of training are carried out based on the interactive training set, and the number of training rounds is set (such as 100 rounds). In each round of training, the following steps are performed in sequence: first, forward propagation is performed, the speech and image modalities in the sample are input into the network, and three types of preference prediction values ​​are output: semantic tendency index, behavioral adsorption index, and emotional response index; then the multi-task loss function is calculated, including: semantic consistency loss: based on the mean square error (MSE) between the semantic label and the predicted value; behavioral focus loss: based on the IoU and KL divergence between the predicted gaze heat map and the real behavior area; emotional response loss: the cross entropy of the predicted emotion and the real label is used, and the three losses are weighted and summed according to the set weights to form a total loss. Backpropagation is performed and the model parameters are updated. The optimizer uses AdamW, combined with weight decay, learning rate warm-up and multi-stage decay strategy to control the training dynamics, and gradient clipping is applied to prevent gradient explosion and improve training stability.

[0081] After each round of training, the model performance is evaluated based on the interactive validation set, forward reasoning is performed on the validation samples, and three types of preference prediction values ​​are output. The corresponding evaluation indicators are calculated respectively, such as semantic accuracy, fixation area IoU, emotion recognition F1-score, and mean error between the three types of indices (MAE). At the same time, the training and validation loss curves and indicator change graphs are plotted to monitor the convergence status and generalization performance of the model. If the validation loss no longer decreases or the indicator tends to be stable for several consecutive rounds, the Early Stopping mechanism is triggered to terminate the training process early.

[0082] After training, the final converged model parameter file is saved and exported to a deployable format for subsequent loading and use in preference recognition and emotion-driven recommendation systems in exhibition interaction scenarios.

[0083] In this implementation, a multimodal user perception mechanism based on an interactive emotion recognition network is constructed to achieve accurate modeling of audience preferences and emotional responses during the exhibition process. At the same time, voice information and browsing image frame data are collected, and the multimodal input layer and user recognition layer are used to identity-bind and standardize the voice and image information, thereby establishing a unified foundation for the subsequent extraction of emotion and preference features. In the modal encoding layer, the user's emotions and interaction status are comprehensively extracted by combining voice rhythm, intonation, keyword semantics, and facial expressions, gaze areas, and behavioral actions in the image. A unified interactive emotion expression vector is constructed through modal fusion and attention mechanisms. This vector is further parsed into three quantitative indicators: semantic tendency, behavioral adsorption, and emotional response in the feature decoding layer, forming a preference response evaluation set that can be used for subsequent display recommendations and content adjustments. This significantly improves the system's ability to understand the user's real emotional reactions and interactive interests, which in turn helps to realize the regulatory mechanism of the linkage between exhibits and audience emotions, and enhances the human-computer interaction intelligence and personalized adaptability of the display.

[0084] Specifically, the specific steps of obtaining the preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: based on the particle swarm optimization algorithm, the semantic tendency index, behavioral adsorption index and emotional response index of each cultural relic to be displayed in the cultural relic restoration exhibition hall are comprehensively analyzed (that is, weighted processing, and the weight coefficients corresponding to the semantic tendency index, behavioral adsorption index and emotional response index are obtained through the particle swarm optimization algorithm, which specifically sets the fusion objective function as a weighted linear combination function of the semantic tendency index, behavioral adsorption index and emotional response index, that is, the user interaction preference index is equal to the sum of the three preference indices multiplied by the corresponding weight coefficients, where the sum of each coefficient is 1, and the value range is limited to the interval of 0-1, then, constructs the search space of the particle swarm optimization algorithm, regards each set of weighted coefficients as the position vector of a particle, initializes the positions and velocities of multiple particles, and sets the individual historical optimal position and the global optimal position for each particle, and defines the fitness function as the current fusion result and the fitness function. The error function between the actual preference labels of the audience in the sample, for example, the minimum mean square error, maximum accuracy or maximum AUC can be used as the optimization target. During the iterative process, the position and velocity of the particles are updated according to the standard formula of particle swarm optimization, and the updated coefficient vector is normalized to ensure that the sum of all weights is 1; the fitness of all particles is re-evaluated in each round of iteration, and the global optimal solution is updated until the preset maximum number of iterations is reached, or when the fitness function tends to be stable, the optimization process is terminated early. Finally, the optimal weighted coefficient after convergence is taken as the weight configuration of this round, and it is applied to the three indexes corresponding to each cultural relic) to obtain the user interaction preference index of each cultural relic in the cultural relic restoration exhibition hall to be displayed; the user interaction preference index of each cultural relic in the cultural relic restoration exhibition hall to be displayed is input into the cultural relic restoration display model for embedding processing (consistent with the embedding processing logic of the display recommendation index) to obtain the user interaction preference index of each cultural relic in the cultural relic restoration display model.

[0085] In this implementation, the particle swarm optimization algorithm is introduced to perform weighted fusion analysis of the semantic tendency index, behavioral adsorption index, and emotional response index, so as to accurately calculate the user interaction preference index of each cultural relic. The particle swarm optimization can dynamically adjust the importance ratio of each factor according to the historical preference samples, so that the fusion result is more in line with the actual audience response. By constructing the search space and fitness function, and then combining multiple rounds of iterative optimization and normalization operations, the fusion weight is ensured to have global optimality and stability, thereby improving the accuracy of the interaction preference index, enhancing the system's adaptability to the behavioral patterns of different audience groups, and helping to achieve a more emotionally connected cultural relic display sorting and recommendation strategy.

[0086] See also Figure 3The embodiment of the present invention provides a technical solution: a data interactive display system for a digital exhibition hall, comprising: a model construction module, for obtaining text information of each cultural relic to be displayed in the cultural relic restoration exhibition hall and cultural relic display image data at each angle, and inputting the information into a pre-trained cultural relic attribute recognition model for recognition analysis, obtaining the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and constructing a cultural relic exhibition model; a feature analysis module, for inputting the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model, and performing feature analysis, obtaining the display recommendation index of each cultural relic in the cultural relic restoration display model; a user perception analysis module, for ... The method is used to obtain browsing video stream data of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and conduct a comprehensive analysis in combination with the user perception model to obtain a preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall; the comprehensive interaction analysis module is used to input the preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model for data analysis to obtain the user interaction preference index of each cultural relic in the cultural relic restoration display model, and conduct a comprehensive analysis in combination with the display recommendation index to obtain the exhibition item response index of each cultural relic in the cultural relic restoration display model; the display feedback module is used to adjust the display of each cultural relic in the cultural relic restoration display model based on the exhibition item response index.

[0087] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0088] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A data interactive display method for a digital exhibition hall, characterized in that: The following steps are involved: Obtain text information of each cultural relic to be displayed in the cultural relic restoration exhibition hall and image data of cultural relic display at each angle, and input them into a pre-trained cultural relic attribute recognition model for recognition analysis, obtain the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and construct a cultural relic exhibition model; Input the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model, and perform feature analysis to obtain the display recommendation index of each cultural relic in the cultural relic restoration display model; Obtain browsing video stream data of each cultural relic in the cultural relic restoration exhibition hall to be displayed, and conduct comprehensive analysis in combination with the user perception model to obtain a preference response evaluation set for each cultural relic in the cultural relic restoration exhibition hall to be displayed; Input the preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model for data analysis to obtain the user interaction preference index of each cultural relic in the cultural relic restoration display model, and perform comprehensive analysis in combination with the display recommendation index to obtain the exhibition item response index of each cultural relic in the cultural relic restoration display model; Based on the exhibit response index, the display of each cultural relic in the cultural relic restoration display model is adjusted.

2. The data interactive display method for a digital exhibition hall according to claim 1, characterized in that: The cultural relic display image data is specifically the pixel value of each pixel point in the cultural relic display image, the cultural relic attribute set includes a component set and a display indicator set, and the cultural relic attribute recognition model is specifically a visual language attribute recognition network, which includes an input layer, a feature encoding layer, a multi-modal fusion layer, and an output layer.

3. The data interactive display method for a digital exhibition hall according to claim 2, characterized in that: The specific steps for obtaining the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: In the input layer of the visual language attribute recognition network, the text information of each cultural relic to be displayed in the cultural relic restoration exhibition hall and the image data of the cultural relic display at each angle are received and preprocessed; In the feature encoding layer of the visual language attribute recognition network, feature extraction processing is performed on the pre-processed text information of each cultural relic in the cultural relic restoration exhibition hall and the cultural relic display image data at each angle, thereby obtaining a text feature vector and an image feature vector set for each cultural relic in the cultural relic restoration exhibition hall. In the multimodal fusion layer of the visual language attribute recognition network, the text feature vector and image feature vector set of each cultural relic to be displayed in the cultural relic restoration exhibition hall are fused to obtain a joint semantic representation vector for each cultural relic to be displayed in the cultural relic restoration exhibition hall. In the output layer of the visual language attribute recognition network, the joint semantic representation vector of each cultural relic to be displayed in the cultural relic restoration exhibition hall is parsed to obtain the component set and display index set of each cultural relic to be displayed in the cultural relic restoration exhibition hall; The display index set includes a repair intervention index, a visual coherence index, a cognitive load index, a display medium dependence index, a color attractiveness index, and a structural hierarchy index.

4. The data interactive display method for a digital exhibition hall according to claim 3, characterized in that: The specific steps for obtaining the display recommendation index of each cultural relic in the cultural relic restoration and display model are as follows: Read the display index set in the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and perform comprehensive analysis on each of them to obtain a display evaluation index set for each cultural relic to be displayed in the cultural relic restoration exhibition hall, including a visible load index and an exhibition experience index; A comprehensive analysis is performed on the display evaluation index set of each cultural relic to be displayed in the cultural relic restoration exhibition hall to obtain the display recommendation index of each cultural relic to be displayed in the cultural relic restoration exhibition hall, and the index is input into the cultural relic restoration display model for embedding processing to obtain the display recommendation index of each cultural relic in the cultural relic restoration display model.

5. The data interactive display method for a digital exhibition hall according to claim 4, characterized in that: The specific formula for calculating the display recommendation index of a cultural relic to be displayed in a cultural relic restoration exhibition hall is as follows: Among them, ZsT, TyZ, and XgF are the display recommendation index, exhibition experience index, and explicit structure load index of a certain cultural relic to be displayed in the cultural relic restoration exhibition hall, respectively. η1, η2, and η3 are the experience adjustment coefficient, explicit structure adjustment coefficient, and interaction adjustment coefficient stored in the database, respectively.

6. The data interactive display method for a digital exhibition hall according to claim 1, characterized in that: The browsing video stream data is specifically voice information and several frames of browsing image data. The browsing image data is specifically the browsing pixel value and browsing two-dimensional coordinates of each browsing pixel point in the browsing image. The user perception model is specifically an interactive emotion recognition network. The interactive emotion recognition network includes a multi-modal input layer, a user recognition layer, a modal encoding layer, a modal fusion layer, and an emotion feature decoding output layer.

7. The data interactive display method for a digital exhibition hall according to claim 6, characterized in that: The specific steps for obtaining the preference response evaluation set for each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: In the multimodal input layer of the interactive emotion recognition network, the browsing video stream data of each cultural relic to be displayed in the cultural relic restoration exhibition hall is received and preprocessed; In the user recognition layer of the interactive emotion recognition network, the pre-processed browsing video stream data of each cultural relic in the exhibition hall to be displayed is recognized and processed to obtain the interaction modality data of each viewer within a set range for each cultural relic in the exhibition hall to be displayed; In the modal coding layer of the interactive emotion recognition network, feature coding is performed on the interactive modal data of each viewer within a set range of each cultural relic to be displayed in the cultural relic restoration exhibition hall, thereby obtaining a voice emotion feature vector and an image interaction feature vector for each viewer within a set range of each cultural relic to be displayed in the cultural relic restoration exhibition hall; In the modal fusion layer of the interactive emotion recognition network, the voice emotion feature vectors and image interaction feature vectors of each viewer within the set range of each cultural relic to be displayed in the cultural relic restoration exhibition hall are cross-modally fused to obtain the interactive emotion expression vectors of each viewer within the set range of each cultural relic to be displayed in the cultural relic restoration exhibition hall; In the feature decoding output layer of the interactive emotion recognition network, the interactive emotion expression vector of each audience member within the set range of each cultural relic to be displayed in the cultural relic restoration exhibition hall is feature decoded to obtain the semantic tendency index, behavioral adsorption index, and emotional response index of each cultural relic to be displayed in the cultural relic restoration exhibition hall, that is, the preference response evaluation set.

8. The data interactive display method for a digital exhibition hall according to claim 7, characterized in that: The specific steps for obtaining the preference response evaluation set for each cultural relic to be displayed in the cultural relic restoration exhibition hall are as follows: Based on the particle swarm optimization algorithm, a comprehensive analysis is conducted on the semantic tendency index, behavioral adsorption index, and emotional response index of each cultural relic in the cultural relic restoration exhibition hall to be displayed, and the user interaction preference index of each cultural relic in the cultural relic restoration exhibition hall to be displayed is obtained; The user interaction preference index of each cultural relic to be displayed in the cultural relic restoration exhibition hall is input into the cultural relic restoration display model for embedding processing to obtain the user interaction preference index of each cultural relic in the cultural relic restoration display model.

9. The data interactive display method for a digital exhibition hall according to claim 1, characterized in that: The specific formula for calculating the response index of a cultural relic in the cultural relic restoration and display model is as follows: Among them, ZxY, ZsT, and HyJ are the exhibition item response index, display recommendation index, and user interaction preference index of a certain cultural relic in the cultural relic restoration and display model, respectively; μ1, μ2, and μ3 are the recommendation adjustment coefficient, interaction preference adjustment coefficient, and deviation adjustment coefficient stored in the database, respectively.

10. A data interactive display system for a digital exhibition hall, applying the data interactive display method for a digital exhibition hall according to any one of claims 1 to 9, characterized in that: include: A model building module is used to obtain text information of each cultural relic to be displayed in the cultural relic restoration exhibition hall and image data of cultural relic display at each angle, and input them into a pre-trained cultural relic attribute recognition model for recognition analysis, thereby obtaining an attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall and building a cultural relic exhibition model; A feature analysis module is used to input the attribute set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model, and perform feature analysis to obtain a display recommendation index for each cultural relic in the cultural relic restoration display model; The user perception analysis module is used to obtain the browsing video stream data of each cultural relic in the cultural relic restoration exhibition hall to be displayed, and conduct a comprehensive analysis in combination with the user perception model to obtain a preference response evaluation set for each cultural relic in the cultural relic restoration exhibition hall to be displayed; A comprehensive interaction analysis module is used to input the preference response evaluation set of each cultural relic to be displayed in the cultural relic restoration exhibition hall into the cultural relic restoration display model for data analysis to obtain the user interaction preference index of each cultural relic in the cultural relic restoration display model, and to perform comprehensive analysis in combination with the display recommendation index to obtain the exhibition item response index of each cultural relic in the cultural relic restoration display model; The display feedback module is used to adjust the display of each cultural relic in the cultural relic restoration display model based on the exhibit response index.

Citation Information

Patent Citations

  • A data interactive display method and system for a digital exhibition hall

    CN119781891A