Digital culture space full-process intelligent collaborative display system based on AI
Through the AI-driven digital cultural space display system, combined with three-dimensional point cloud data and historical video stream data, the display location of museum exhibits is optimized, the problem of insufficient user behavior perception is solved, the display effect of exhibits is matched with audience attention, and the intelligence and stability of the display system is improved.
Patent Information
- Application Number
- CN202510728076.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to systematically collect and analyze users' interaction behavior before the exhibit, resulting in the display effect of the exhibit and the audience's attention, the configuration of display resources is imbalanced, and the lack of intelligent collaborative display of interaction, feedback and optimization.
Through the AI-based digital cultural space full-process intelligent collaborative display system, a museum structure display model is constructed using three-dimensional point cloud data and text information, and the user behavior index is extracted in combination with historical video stream timing data, initial regulation and optimization of display locations are carried out, and a posture-guided behavior recognition model and component attribute recognition network are introduced, a user behavior response index and local semantic coordination index are generated to optimize the display location to match user preferences.
It significantly improves the degree of matching between display information and user preferences, enhances the learning ability and real-time allocation ability of the display system, maintains semantic consistency and structural coherence between display logic, avoids the chaos caused by large-scale rearrangement, and builds an intelligent regulatory mechanism of behavior-driven and structural constraints.
Smart Images

Figure CN120495592A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent collaborative display technology, and specifically to an AI-based full-process intelligent collaborative display system for digital cultural spaces. Background Art
[0002] With the continuous development of artificial intelligence, three-dimensional reconstruction, virtual reality and big data technologies, the digitalization process of cultural space has gradually evolved from static visualization to semantic intelligence, forming a new technical paradigm for the full-process intelligent collaborative display of digital cultural space. In the digital cultural space scenario, museums are digital twin objects of core cultural entities. The structure, function and display effect of their exhibits become the key media for building an intelligent linkage between semantic space and user experience. In particular, the display of cultural relics in museums, which are composed of a large number of complex components, needs not only to restore the geometric shape, but also to reveal the internal structural semantics and historical connections, and at the same time, to dynamically adjust the display based on the behavioral characteristics of users during the exhibition.
[0003] The limitations of existing technologies include at least the following problems: it is difficult for existing technologies to systematically collect and analyze users' interactive behaviors in front of each exhibit, making it difficult to truly capture the popularity and information communication effects of exhibits during actual exhibitions. Due to the difficulty in perceiving user preferences, existing technologies often place exhibits that attract strong attention from some audiences at the visual edge or in low-priority positions, while mistakenly placing some low-participation exhibits in core areas, resulting in an imbalance in the allocation of exhibition resources. This makes it difficult to dynamically optimize exhibition strategies based on past exhibition behaviors, and there is a lack of intelligent collaborative display with interaction, feedback, and optimization. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention provides an AI-based full-process intelligent collaborative display system for digital cultural spaces, which solves the problem that the display process of the existing technology lacks user behavior perception drive, resulting in a disconnect between the display effect of exhibits and audience attention.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: an AI-based full-process intelligent collaborative display system for digital cultural spaces, comprising: a data acquisition and construction module, for acquiring three-dimensional point cloud data and text information of several museum collections to be displayed, and inputting the data into a pre-trained structure recognition model for comprehensive analysis to obtain a component attribute set for each museum collection to be displayed, and constructing a museum structure display model; a behavior acquisition and analysis module, for acquiring historical video stream time series data of the museum to be displayed, and inputting the data into a pre-trained behavior recognition model for comprehensive analysis to obtain a user behavior index set for each museum collection to be displayed; a comprehensive response analysis module, for inputting the user behavior index set for each museum collection to be displayed into the museum structure display model for comprehensive analysis to obtain a user behavior response index for each collection in the museum structure display model; an initial display control module, for performing initial display position control for each collection in the museum structure display model based on the user behavior response index; and an optimized display control module, for performing comprehensive analysis on the component attribute set of each collection in the museum structure display model after the initial display position control and each collection within a set range, to obtain a local semantic coordination index for each collection in the museum structure display model after the initial display position control, and to perform optimized display position control.
[0006] The present invention has the following beneficial effects: (1) This AI-based digital cultural space full-process intelligent collaborative display system introduces a posture-guided behavior recognition model. Based on the time series data of historical exhibition video streams, it extracts the user's behavioral semantic feature set in front of each collection and constructs a user behavior response index based on it, thereby achieving a structured quantitative characterization of user attention. Based on the user behavior response index, the display position is initially adjusted, so that the collections with high response indexes are matched with priority visual display positions, which significantly improves the matching degree between display information and user preferences, and enhances the exhibition system's ability to learn historical user behavior and real-time deployment capabilities. At the same time, it effectively establishes a dynamic display feedback mechanism driven by user behavior.
[0007] (2) This AI-based digital cultural space full-process intelligent collaborative display system uses three-dimensional point cloud data and text information, and extracts the component attribute set of each collection based on the component attribute recognition network, and then generates a unified component vector representation in the semantic embedding coding module. Based on the initial display position regulation, the component vectors of each collection and the surrounding collections are analyzed for semantic similarity, and combined with the spatial distance value, a local semantic coordination index is constructed, which is then used to guide the secondary optimization regulation of the display position, thereby ensuring that the display layout maintains semantic consistency and structural coherence of the display logic while taking into account user preferences, and significantly enhances the spatial coordination and cultural context coherence of the overall display model.
[0008] (3) This AI-based digital cultural space full-process intelligent collaborative display system constructs an optimization control mechanism based on the comprehensive display efficiency objective function, thereby dynamically evaluating the comprehensive efficiency of the current display configuration of each collection. For collections with low semantic coordination, the system performs a limited number of replacement simulations within the preset physical neighborhood. If the efficiency is significantly improved after the replacement, the optimal allocation is performed. If there is no improvement, the original position is retained, thereby avoiding the display chaos caused by large-scale rearrangement, and improving the overall layout performance through controllable disturbance optimization, taking into account stability and adaptability, and then effectively constructing an intelligent control mechanism with two-way feedback of behavior drive and structural constraints.
[0009] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a block diagram of the AI-based full-process intelligent collaborative display system for digital cultural space in the present invention.
[0011] Figure 2 This is a flowchart of the specific steps for obtaining the user behavior response index of each collection in the museum structure display model in the AI-based digital cultural space full-process intelligent collaborative display system of the present invention.
[0012] Figure 3 This is a flowchart of the specific steps for obtaining the local semantic coordination index of each collection in the museum structure display model after the initial display position is adjusted in the AI-based digital cultural space full-process intelligent collaborative display system of the present invention. DETAILED DESCRIPTION
[0013] See also Figure 1An embodiment of the present invention provides a technical solution: an AI-based full-process intelligent collaborative display system for digital cultural spaces, comprising: a data acquisition and construction module, configured to acquire three-dimensional point cloud data and text information of a plurality of museum collections to be displayed, and input the data into a pre-trained structure recognition model for comprehensive analysis to obtain a component attribute set for each museum collection to be displayed, and to construct a museum (three-dimensional semantic) structure display model; a behavior acquisition and analysis module, configured to acquire historical video stream time series data of the museum to be displayed, and input the data into a pre-trained behavior recognition model for comprehensive analysis to obtain a user behavior index set for each museum collection to be displayed; a comprehensive response analysis module, configured to input the user behavior index set for each museum collection to be displayed into the museum structure display model for comprehensive analysis to obtain a user behavior response index for each collection in the museum structure display model; an initial display control module, configured to perform initial display position control for each collection in the museum structure display model based on the user behavior response index; and an optimized display control module, configured to perform comprehensive analysis on the component attribute set of each collection in the museum structure display model after the initial display position control and each collection within a set range to obtain a local semantic coordination index for each collection in the museum structure display model after the initial display position control, and to perform optimized display position control.
[0014] The specific steps of constructing the museum structure display model are as follows: based on the component identifier, each component is numbered and identified as an independent node in the display model to ensure the uniqueness and indexability of each component in the overall display structure. Secondly, the minimum circumscribed bounding box coordinates, component center of gravity coordinates and voxel point data in the spatial position set are used to restore the geometric shape and actual spatial position of each component in three-dimensional space, and all components are spatially assembled to form the basic geometric framework of the collection in the museum scene. Thirdly, semantic labels are given to each component based on the component type, material category, functional use and other information in the semantic attribute set to realize component-level functional labeling and semantic hierarchical division in the model, so that the model not only has It not only has a geometric structure but also has semantic interpretability. Then, according to the color coding, transparency level and surface reflection level in the rendering parameter set, the visual performance of each component in the display model is controlled to achieve realistic restoration of materials and layered highlight rendering, thereby improving the visual readability and immersion of the display model. Finally, based on the connection target identification, connection method, connection position and connection direction information contained in the semantic association set, the structural connection relationship diagram between components is automatically established, and the logical assembly and structural linkage between components are realized in three-dimensional space. At the same time, through structural group mapping, semantic structural units such as cover components and central support groups are identified to achieve semantic combination display between components, thereby completing the construction of the museum's three-dimensional semantic structure display model.
[0015] The specific steps for optimizing the display position regulation are as follows: the local semantic coordination index of each collection of the museum structure display model after the initial display position regulation is judged and analyzed with the preset local semantic coordination index threshold; if the local semantic coordination index of each collection of the museum structure display model after the initial display position regulation is higher than the preset local semantic coordination index threshold, the collection will not be optimized for display position regulation; if the local semantic coordination index of each collection of the museum structure display model after the initial display position regulation is lower than or equal to the preset local semantic coordination index threshold, the collection will be optimized for display position regulation, which is specifically as follows: for each collection in the museum structure display model, extract its user behavior response index, local semantic coordination index, and the preset display priority weight value corresponding to its current display position, so as to construct a comprehensive display effectiveness objective function, which is calculated in sequence with each collection as a unit as follows: multiply the user behavior response index of each optimized collection by the first regulation factor, multiply the local semantic coordination index by the second regulation factor, then add the two and multiply by the priority weight value of its display position, and finally, the display effectiveness of all collections is optimized. The values are accumulated to obtain the total display effectiveness value for the current display state, which is recorded as the original objective function value. Then, all items with a local semantic coordination index below the preset local semantic coordination index threshold are screened and considered as candidate optimization objects. For each optimized item (i.e., an item below the preset local semantic coordination index threshold), attempts are made to swap the display position with other items in its physically adjacent display locations (for example, limiting the display position to positions within a set radius). For each set of swap attempts, the global display effectiveness objective function value after the swap is recalculated. If the objective function value after the swap increases by more than a preset significance threshold (for example, an increase of 0.02 from the original value), the swap is determined to have an optimization benefit and the display position swap operation is executed. Conversely, if the improvement is not significant, the swap attempt is canceled and the optimized item is marked as non-exchangeable to avoid repeated processing in subsequent optimizations. The maximum number or proportion of items that can be manipulated in each optimization round is limited (for example, no more than 10% of the total number of items). While meeting the manipulation limit, items with high user behavior response indexes but poor semantic coordination are prioritized for optimization adjustment.
[0016] Specifically, the three-dimensional point cloud data includes the voxel value and three-dimensional coordinates of each voxel point. The component attribute set includes the component identifier, spatial position set, semantic attribute set, rendering parameter set, and semantic association set of each component. The structure recognition model is specifically a component attribute recognition network (Point-BERT and BERT fusion architecture). The multimodal component attribute recognition network includes an input layer, a multimodal guided encoding layer, a semantic fusion recognition layer, a semantic attribute enhancement layer, and an attribute set output layer.
[0017] Among them, the component identifier is a unique identifier, that is, a coding identifier of each component.
[0018] The spatial position set includes the coordinates of the minimum circumscribed bounding box, the three-dimensional coordinates of the component's center of gravity, and each voxel point that constitutes the component.
[0019] The semantic attribute set includes component type (such as lid, body, base), material category (such as copper, ceramic, wood), and functional use (such as holding water, sealing, and display).
[0020] The rendering parameter set includes color encoding, such as RGB [185, 140, 80] corresponding to copper color, as well as transparency parameters (such as 0.95 for slightly transparent) and surface reflection level.
[0021] The semantic association set includes the connection target identifier (the unique identifier of the target component to which the component is connected, used to clarify the structural dependence or connection object between components), the connection method (the connection method between the component and the target component), the connection position (the specific spatial position of the component involved in the connection), and the connection direction (the connection direction between the component and the target component).
[0022] The input layer is used to receive the three-dimensional point cloud data and corresponding text information of each collection to be displayed in the museum, and preprocess the multimodal input data in a unified format.
[0023] The multimodal guided encoding layer is used to extract geometric features and semantic cues, fuse point cloud and text information, and generate semantically enhanced feature representations.
[0024] The semantic fusion recognition layer is used to identify the component-level regional divisions in each collection based on the fusion semantic feature set, and to construct a candidate set of semantic labels for each component to preliminarily express its semantic attributes. The semantic attribute enhancement layer is used to verify and optimize candidate tags, confirm the main attributes of each component, and improve semantic consistency.
[0025] The attribute set output layer is used to encapsulate component number, spatial position, semantic attributes and rendering parameters, and output standardized component attribute sets.
[0026] The specific steps of obtaining the component attribute set of each collection to be displayed in the museum are as follows: in the input layer of the component attribute recognition network, the three-dimensional point cloud data and text information of each collection to be displayed in the museum are received and preprocessed; in the multimodal guided encoding layer of the component attribute recognition network, the preprocessed three-dimensional point cloud data and text information of each collection to be displayed in the museum are semantically guided encoding (the three-dimensional point cloud data of each collection is input into the Point-BERT-based point cloud encoder stored in the data to perform point-by-point geometric feature embedding on the point cloud, such as extracting the spatial coordinates, normal vector, local curvature and other geometric description information of the point, and combining the local neighborhood extraction and global aggregation mechanism to obtain the high-dimensional structural semantic representation of each point, forming a three-dimensional structural embedding feature matrix for each collection; secondly, the text information corresponding to the collection, including the name, purpose, construction description, etc. stored in the input data, is embedded in the point cloud encoder. The BERT model is used to perform context-aware word vector encoding, extract key semantic units such as component names, attribute terms, and functional terms, and generate a text semantic feature representation matrix for each collection. Then, a set of keyword vectors is extracted from the text semantic features as semantic hint embedding to guide semantic understanding in the point cloud encoding process. The correlation distribution weight between text semantics and local features of the point cloud is calculated through the attention mechanism. The structural feature matrix and the hint vector set are then guidedly fused using a bidirectional cross-attention mechanism to generate a point-level fused feature representation with text semantic enhancement capabilities. The fused semantic feature set of each collection to be displayed in the museum is obtained (including but not limited to the geometric feature dimensions of each point, such as spatial coordinates, normal vectors, curvature values, etc., as well as the semantic correlation embedding vector between the point and the component keyword, which is used to support the input representation of component-level recognition and attribute label decoding).In the semantic fusion recognition layer of the component attribute recognition network, the fused semantic feature set of each collection of the museum to be displayed is subjected to semantic-driven component recognition processing (for the embedding vector of each point in the fused semantic feature set, a feature similarity measurement method such as Euclidean distance or Gaussian kernel function is used to calculate the feature distance between point pairs, and a preliminary feature adjacency graph is constructed based on the distance threshold to identify point sets with similar local structural features. Subsequently, a semantic attention mechanism is introduced based on the feature adjacency graph, and the semantic response weight between each point pair and the text semantic prompt vector is used as the edge weighting factor to generate a semantic weighted feature map). The feature map not only retains the local geometric structure information of the point cloud, but also introduces a semantic cue-driven point pair similarity adjustment mechanism, thereby improving the sensitivity of component region division to semantic differences. Then, based on the semantic weighted feature map, a semantic-aware clustering algorithm, such as the semantic-guided K-means or MeanShift algorithm, is used to perform component-level region division on the point cloud data, obtaining multiple component instance regions with geometric continuity and semantic consistency. For example, although the lid and the base may be similar in geometry, they can be effectively distinguished in the clustering process through the attention guidance of semantic cue keywords. After the component region division is completed, the correlation between the fusion feature center of each component instance region and each semantic keyword vector is calculated, and the correlation scores are sorted according to the correlation scores to generate a semantic label candidate list for each component. For example, for a component region, its semantic label candidates may include lid, rivet connection, copper, etc.), and a component semantic candidate set for each collection to be displayed in the museum is obtained (including but not limited to the spatial range of all component regions, region identification numbers and their corresponding semantic label candidate sets, etc.). In the semantic attribute enhancement layer of the component attribute recognition network, the component semantic candidate set for each collection to be displayed in the museum is attribute-calibrated (multi-dimensional similarity comparison is performed on each keyword vector in the semantic label candidate list of each component with the fusion feature center of the component region, such as cosine similarity or Transformer attention score, and the semantic consistency score of each candidate label is calculated. Then, low-confidence labels that deviate significantly from the fusion feature semantic direction are removed to form a converged candidate label set. Based on the predefined semantic logic rule map, for example, rivet connection cannot appear at the same time as pottery and lid cannot be the bottom component, conflict detection is then performed on the label combination of each component.If there are logically contradictory label combinations, those with lower semantic scores or lower frequencies are preferentially eliminated to ensure that the output label set is logically consistent in terms of function, structure, material, etc., and then the main label of each component's attribute domain, such as type, purpose, connection method, material, etc., is confirmed. In the same attribute domain, if there are still multiple high-confidence labels, a voting mechanism is introduced, that is, a unique label is determined based on the context label coupling degree, historical training distribution or domain prior. For example, in the material domain, copper is confirmed as the main label from copper and pottery. Then, for each component label, the context consistency is enhanced in combination with the attribute distribution of its spatially adjacent components. For example, if the adjacent components are all bronze systems, the score of the copper label in the current component can be improved, while the score of the pottery label can be further lowered, forming an attribute adjustment with coherent semantic style and coordinated exhibit configuration. Finally, a component semantic attribute set is formed to obtain the component semantic attribute set of each collection to be displayed in the museum (including but not limited to the spatial range identification of each component, component number, attribute domain main label, such as type, material, purpose, connection method, etc. In the attribute set output layer of the component attribute recognition network, the semantic attribute set of each component of the museum's collection to be displayed is encapsulated (i.e., each component is assigned a unique component identifier, such as CPT_001, and the component's spatial extent, such as the minimum bounding box, centroid coordinates, and 3D point set, is embedded in the attribute structure to bind attribute semantics with spatial location. Semantic attribute labels are uniformly named and populated according to predefined fields, such as type, material, usage, and connectMode, ensuring that the output data structure can be directly parsed by the visualization module or knowledge graph system. Visual rendering parameters, such as material textures, color coding, and transparency levels, are automatically matched based on the attribute content, such as copper or ceramic. The above numbers, spatial information, standard fields, and visualization parameters are combined into a complete component attribute structure, forming the final output component attribute set). The component identifier, spatial location set, semantic attribute set, rendering parameter set, and semantic association set are obtained for each component of each museum's collection to be displayed.
[0027] The pre-training process of the component attribute recognition network is as follows: Obtain annotated 3D collection point cloud and semantic attribute datasets, including 3D laser scanning point cloud data of several real exhibits and their corresponding expert-annotated text information. The data of each exhibit includes: point cloud voxel coordinates, geometric features (such as normal vectors, local curvature), spatial partition labels (such as component divisions such as lids, bodies, and bases). The text data includes semantic texts such as exhibit name, usage description, structural feature terms, and a set of attribute labels for each component (type, material, usage, connection method, etc.). The dataset is divided into a component training set and a component verification set.
[0028] The component attribute recognition network is initialized, and a multimodal input structure of the network is constructed, including a three-dimensional point cloud encoding module based on Point-BERT, a text semantic encoding module based on BERT, a multimodal cross-guided fusion module, a semantic-aware component recognition module, and an attribute enhancement decoding module. During initialization, the key hyperparameters of each layer are set, such as the embedding dimension of Point-BERT, the K-neighborhood sampling scale, the context window length of the BERT model, the number of semantic fusion attention heads, the semantic consistency alignment weight coefficient, etc. The network weights are initialized using the Kaiming method.
[0029] Training is performed based on the component training set, and the number of training rounds is set (e.g. 80 rounds). The following stages are executed in sequence during each round of training: In the forward propagation stage, the three-dimensional point cloud data and text information of each collection are input into the network, and the structured semantic attributes of each component are output through each layer of the network in turn. A loss function is constructed, and a multi-branch task loss is designed for each attribute domain, including the IoU loss for component area division accuracy, the cross entropy loss for attribute classification, and the rule conflict penalty loss for attribute logical consistency. Weights are set according to the relevance between the attribute domain and the task to form an overall weighted multi-task loss function. Then, in the backpropagation and parameter optimization stage, the AdamW optimizer is used to update the network parameters, and the gradient clipping, batch normalization and learning rate cosine annealing strategies are combined to enhance the stability and global convergence ability of the training process.
[0030] After each round of training, the component validation set is used to verify the network, evaluate the accuracy of point cloud semantic segmentation (such as mIoU, PA), attribute classification accuracy (such as F1-score, Accuracy), and the semantic completeness and logical consistency of the final component attribute set, and draw loss curves and indicator change trend charts to evaluate model convergence and generalization performance. If there is no improvement for multiple consecutive cycles, the Early Stopping mechanism is triggered to terminate training.
[0031] After the training is completed, save the final component attribute recognition model and all its network weight parameter files.
[0032] In this implementation, by deeply fusing three-dimensional point cloud data with semantic text information, the point cloud, which originally only had geometric information, is given semantic guidance capabilities during the recognition process, breaking through the limitations of traditional point cloud recognition in component function discrimination. Secondly, in the multimodal guided encoding layer, a bidirectional cross-attention mechanism is used to inject the semantic hint vector extracted by BERT into the spatial feature expression process of Point-BERT, so that the embedded features of each point not only consider the local geometry, but also pay attention to its correlation with the component vocabulary at the cultural semantic level, so that isomorphic morphological components can be successfully distinguished in the semantic dimension. In addition, in the semantic fusion recognition and attribute enhancement stage, by constructing a semantic weighted feature map, introducing a semantic logic rule map and a context consistency enhancement mechanism, the accuracy of component division and the logical consistency of attribute annotation are effectively improved. Finally, the generated component attribute set has standard field definition and semantic logic completeness, which greatly enhances the overall intelligence, structure and usability of the museum's digital collection modeling.
[0033] Specifically, the historical video stream time series data is the historical video stream data of each historical exhibition, and the historical video stream data includes the pixel value and two-dimensional coordinates of each pixel point in several frames of historical exhibition images. The user behavior index set includes the immersion response index, the gaze focus stability index, the stay behavior structure index, and the observation path deviation index. The behavior recognition model is specifically a posture-guided behavior network, that is, a behavior recognition network that combines time-series video behavior modeling (SlowFast) with cross-entity graph modeling (ST-GCN). The posture-guided behavior network includes a video input layer, a dual-stream time series feature extraction layer, a posture extraction construction layer, a graph convolution behavior recognition layer, and a behavior output layer.
[0034] The video input layer is used to receive video stream data of each historical exhibition.
[0035] The dual-stream temporal feature extraction layer is used for character detection. It extracts the appearance and motion information (Slow branch) and fast micro-motion features (Fast branch) of each character in the video, and performs fusion encoding to form multi-scale temporal features that reflect behavioral dynamics.
[0036] The posture extraction construction layer is used to extract the skeleton key points of each person in each frame based on the fused image features, and construct a complete temporal posture trajectory to represent their entire body movement pattern.
[0037] The graph convolution behavior recognition layer is used to model the user's posture trajectory in front of the exhibit as a spatiotemporal graph structure, and extract the semantic features of their behavior at different times through graph convolution to generate a structured representation that can distinguish behavior categories.
[0038] The behavior output layer is used to perform multi-branch prediction and calculate the immersion response index, gaze focus stability index, stay behavior structure index and observation path deviation index respectively.
[0039] The specific steps of obtaining the user behavior index set for each collection of the museum to be displayed are as follows: in the video input layer of the posture-guided behavior network, the historical video stream time series data of the museum to be displayed (i.e., the pixel value and two-dimensional coordinate of each pixel point in each frame of the historical exhibition image of each historical exhibition) are received and preprocessed; in the dual-stream time series feature extraction layer of the posture-guided behavior network, the historical video stream time series data of the museum to be displayed are detected and encoded (multi-target person detection is performed on each frame of the historical exhibition image of each exhibition, and all user entities that appear are identified using the person detection model stored in the database, such as YOLOv8, and a cross-frame target tracking algorithm, such as DeepSort, is used). The same user is numbered and marked to form a continuous user trajectory sequence. Then, the image trajectory area corresponding to each user in each historical exhibition is cropped into an independent video subsequence, and the time sequence length, image size and light intensity are unified to eliminate the perceptual deviation caused by different exhibition sites and locations. The standardized image sequence input of each user is obtained. Then, the video subsequence of each user is input into the fast branch and slow branch of the SlowFast network structure stored in the database. The fast branch processes at a high frame rate to capture fast action features, such as turning and hand operations, and the slow branch processes at a low frame rate to extract slow behavior patterns such as staying and looking. Finally, a multi-scale temporal feature alignment mechanism is used, such as Temporal convolution fusion or attention weighted fusion performs synchronous alignment, semantic complementarity and feature weighting on the features of the two temporal channels in the temporal dimension to integrate micro-action information, such as short-term hand clicks and macro-behavior patterns, such as long-term gaze and stay, and outputs a fused temporal behavior feature sequence with a complete temporal semantic structure), and obtains a fused temporal behavior feature sequence set of each user in front of each collection for each historical exhibition of the museum to be displayed; in the posture extraction construction layer of the posture guidance behavior network, the posture trajectory extraction processing is performed on the fused temporal behavior feature sequence set of each user in front of each collection for each historical exhibition of the museum to be displayed (for each frame input image in the fused temporal behavior feature). , use the human posture estimation model stored in the database, such as HRNet or OpenPose, to extract the three-dimensional key point coordinates of the human skeleton, including the head, eyes, shoulders, elbows, hands, torso, knees, etc., then concatenate the key points of each frame in chronological order to construct a continuous posture trajectory data sequence of each user in front of a specific collection, form a posture time series feature vector set, and perform time series modeling on the posture trajectory to extract the behavioral state parameters exhibited by the user during the exhibition, including but not limited to the gaze angle change rate, movement amplitude change value, posture holding time, posture stability coefficient, etc.), and obtain the posture time series feature set of each user in front of each collection in each historical exhibition of the museum to be displayed;In the graph convolution behavior recognition layer of the posture-guided behavior network, the posture time series feature set of each user in front of each collection of each exhibition in the history of the museum is processed for graph behavior recognition (a static skeleton graph is constructed for the posture key points of each frame, with the key points as nodes and the connections between body parts as edges, and the same nodes in adjacent frames are introduced into the time connection edges to construct a unified spatiotemporal skeleton graph structure sequence, and then the constructed spatiotemporal skeleton graph is input into the space-time graph convolution network stored in the database, such as ST-GCN, and then the relative motion characteristics of the key parts are captured in the spatial dimension, such as the change in the distance between the hand and the head, and the continuous action change trend is modeled in the time dimension, such as the behavior chain of approaching, looking, turning, etc., and the multi-scale convolution layer and sliding time window mechanism are used to provide We take microscopic features, such as small-scale rapid movements, and macroscopic features, such as long-term slow movements, and fuse the features at each scale through residual connections to form a hierarchical sequence of behavioral semantic embedding vectors. We then introduce a key point weighted attention mechanism to weightedly encode the importance of different nodes, such as hand, head, and eye movements. We also introduce a time segment attention module to enhance the timing sensitivity of key action segments, such as gaze and turning, thereby obtaining the final structured behavioral semantic feature vector representation of each user in front of each collection. We then obtain a set of behavioral semantic features for each user in front of each collection in each historical exhibition of the museum to be displayed (including but not limited to posture stability vectors, movement amplitude change vectors, observation direction change vectors, gaze area movement trajectory embeddings, and local action frequency embeddings).In the behavior output layer of the gesture-guided behavior network, the behavior semantic feature set of each user in front of each collection for each exhibition of the museum's history is subjected to branch prediction processing (for the immersive response index, the action amplitude change vector is used as the core feature to judge the intensity of the interactive behavior. For each user's continuous observation period in front of a certain collection, the frame-by-frame motion amplitude changes of their key body parts are first counted to form a time series vector. If the amplitude change is significantly higher than the background noise threshold in multiple consecutive time periods, such as more than twice the average motion amplitude of the static state, it is regarded as a valid interactive action. Each time a significant interaction is identified, it is recorded as a step in the behavior chain. At the same time, combined with the local action frequency embedding feature, the analysis is carried out. Whether the user has performed multiple stages of small-scale interactive behaviors in front of the collection, such as short-distance hand movements, frequent perspective adjustments or slight displacement operations. If within a short time window, the frequency of the action is observed to reach the set threshold every 3 seconds, such as at least 3 times, then the time window is determined to be a micro-interaction stage, and then the length of the user's behavior chain in front of the collection is constructed. Subsequently, the behavior chains formed by all users in front of the collection are analyzed. If the user has 3 or more consecutive effective interaction steps and the total duration exceeds the set threshold, such as 10 seconds, then the user is considered to have a deep interaction in front of the collection. Then, the proportion of users with such deep interaction is counted as an indicator of the attractiveness of the collection. At the same time, each historical exhibition will also be analyzed. The average behavior chain length of all users who browsed the collection was calculated to evaluate the depth of interaction. Finally, the proportion of deep interaction users and the average behavior chain length were weighted and fused to generate a standardized value, namely the immersion response index. For the gaze focus stability index, the gaze area movement trajectory embedding was extracted, that is, the gaze path trajectory of each user's eyes or head direction in front of the collection, such as the change in gaze point position per second, and the distribution range of their gaze points in the collection display area was calculated, such as the ratio of focusing on a small area to scanning a large area. If a user's gaze point is concentrated on the key areas of the collection for a long time, such as the structural center and the inscription, and the offset angle is small and the jump frequency is low, then the user is considered to have stable focus. The proportion of such users is counted and combined with their The average gaze offset fluctuation value is used to form the focus stability evaluation result, and finally a normalized gaze focus stability index is output; for the dwell behavior structure index, the posture stability vector and the movement amplitude change vector are read. If the user first appears in a stable static state in front of the artifact, such as standing for more than 2 seconds, followed by continuous movement changes, and then enters a static leaving state, it is marked as a structured dwell. Then, a structured detection is performed on the behavior of all users in front of the artifact, and the proportion of users with the above pattern is counted. Then, the number of stages and duration of each dwell behavior are analyzed to see if they are balanced. For example, if the observation time is much longer than the interaction time, it is unbalanced. Finally, the structural proportion and stage balance are weighted averaged to generate the dwell behavior structure index;For the observation path deviation index, the observation direction change vector is used to construct each user's observation path trajectory in front of the artifact. This is calculated, for example, from head direction and eye gaze point. This trajectory is then smoothed and fitted to identify directional deviation points, such as sudden and sharp head turns, jerking to other areas, and the frequency of path transitions. A smooth path with small deviation angles and high coherence indicates natural observation. Conversely, frequent jumps and large deviations are considered unstable paths. A statistical weighted average of all users' observation trajectory deviations is then taken to assess the effectiveness of the artifact in visual guidance, i.e., the observation path deviation index. All indices in this layer are standardized, with output values ranging from 0 to 1. For each artifact in the museum to be displayed, the immersive response index (measuring the user's interactive immersion in front of the exhibit), the gaze focus stability index (measuring the degree of user attention focused on the exhibit), the dwell behavior structure index (measuring the degree of organization and rhythm of the user's behavior when dwelling in front of the exhibit), and the observation path deviation index (measuring the degree of deviation of the user's gaze path while viewing the exhibit) are obtained.
[0040] The pre-training of the posture-guided behavior network is as follows: A dataset of annotated historical exhibition behavior videos was obtained, which contains time-series video clips of historical exhibitions manually annotated by behavior recognition experts. The data includes the coordinates of the person detection frame in each frame, the positions of key posture points, behavioral action labels (such as staying and observing, interactive operation, and leaving the exhibits), and the user-exhibit interaction identifiers and four behavioral index true value labels associated with each video clip (immersion response index, gaze focus stability index, stay behavior structure index, and observation path deviation index). The dataset of annotated historical exhibition behavior videos was divided into a behavior training set and a behavior verification set.
[0041] Initialize the posture-guided action recognition network, construct a dual-stream temporal feature encoding network (SlowFast structure), a temporal posture extraction module, a graph convolution behavior analysis module, and an exponential prediction branch module. During initialization, key hyperparameters of each layer are set, including the Fast branch frame rate, Slow branch step size, graph structure adjacency method, graph convolution kernel size, depth of each exponential branch structure, dropout ratio, etc. The Kaiming method is used to initialize the weights of each network. At the same time, the action recognition backbone parameters pre-trained on the Kinetics or PoseTrack datasets can be loaded to accelerate convergence and improve stability.
[0042] Training is conducted based on the behavioral training set, with a set number of training rounds (e.g., 100 rounds). The following stages are performed in sequence in each round of training: In the forward propagation stage, each historical exhibition video stream is first input into the network, and then four index prediction values are output through each layer of the network. In the loss function construction stage, a multi-task joint loss function is designed, including the L2 loss of the immersive response index, the trajectory concentration loss (e.g., center deviation variance) of the gaze focus stability index, the behavioral rhythm balance loss of the stay behavior structure index, and the observation path offset regression loss of the observation path offset index. These are weighted and merged according to the task weights. In the backpropagation and parameter optimization stage, backpropagation is performed on the total loss function, and the Adam optimizer is used to update the network weights. The learning rate cosine annealing strategy, gradient clipping, and batch normalization mechanism are combined to ensure the stability of the training process and improve the generalization ability.
[0043] After each round of training, the behavior validation set is used for inference verification to evaluate the accuracy and fitting ability of each index prediction. Indicators such as the prediction error of each index (such as MSE, MAE, R²), user behavior clustering consistency, and trajectory alignment accuracy are output separately. Loss curves and prediction performance trend charts are plotted to monitor training convergence. If the indicators do not improve after multiple consecutive cycles, the Early Stopping mechanism is triggered to terminate the training process.
[0044] After the training is completed, save the final posture-guided behavior recognition model and the corresponding weight parameter file.
[0045] In this implementation, a posture-guided behavior network (PGN) is introduced that integrates SlowFast dual-stream temporal behavior modeling with ST-GCN graph-structured behavior recognition. This allows for accurate characterization of users' dynamic behavior in front of specific exhibits over the entire time period, thereby obtaining multi-dimensional behavioral semantic indices reflecting interaction depth, gaze focus, dwell structure, and observation path deviation. Furthermore, in the behavior extraction layer, the fast and slow branches model micro-movement and slow-motion features, respectively, ensuring the temporal adaptability of behavior recognition. In the posture modeling layer, the spatial structure and temporal evolution of human skeleton trajectories are combined to construct a high-fidelity spatiotemporal graph behavior representation. Finally, in the graph convolutional analysis layer, the importance of key points and the attention paid to key time periods are integrated to improve the ability to discriminate behavioral patterns. Finally, the four output user behavior indices are standardized to balance differences in behavior amplitude and stability, providing a quantitative, controllable, and feedback-based user response evaluation basis for subsequent display control, significantly enhancing the display system's intelligent perception and responsiveness to users' real experiences.
[0046] Specifically, if Figure 2As shown, the specific steps for obtaining the user behavior response index of each collection of the museum structure display model are as follows: based on the museum structure display model, the user behavior index set of each collection to be displayed in the museum is mapped (the unique identifier of each collection, such as ITEM_001, ITEM_002, is extracted from the museum structure display model. The identifier is used as the primary key identifier of the collection entity node in the three-dimensional semantic model, and a behavior data index table is constructed to bind the user behavior index set of each collection of the museum to be displayed to the corresponding collection entity node), and the user behavior index set of each collection of the museum structure display model is obtained; the user behavior index set of each collection of the museum structure display model is comprehensively analyzed to obtain the user behavior response index of each collection of the museum structure display model.
[0047] The specific formula for calculating the user behavior response index of a collection in the museum structure display model is as follows: ;in, The user behavior response index of a collection of the museum structure display model, The immersive response index of a collection for a museum structure display model, is the immersion adjustment coefficient stored in the database, The focus stability index of a collection in a museum structure display model. is the gaze focus adjustment coefficient stored in the database, The dwell behavior structure index of a collection in the museum structure display model. is the stay adjustment coefficient stored in the database, The observation path deviation index of a collection in a museum structure display model, The path deviation adjustment coefficient stored in the database.
[0048] What needs to be explained is that the specific form of the tanh function is: ,in, is a natural constant and can be taken as 2.71 in this embodiment, with a domain of (−∞, +∞) and a range of (−1, +1).
[0049] 、 、 、 It can be obtained through the following steps: using historical data, combined with the immersion response index, gaze focus stability index, stay behavior structure index, and observation path deviation index, statistical regression analysis is performed to quantify the specific impact of each factor on the user behavior response index, thereby fitting the initial weight value; secondly, using the sensitivity analysis method, adjust the value range of each coefficient, observe its impact on the user behavior response evaluation results, ensure the stability and rationality of the model, and based on the characteristics of the collection and the actual situation, correct and optimize the preliminary fitting coefficients, and finally determine the coefficient value applicable to the specific collection.
[0050] The specific implementation example of calculating the user behavior response index of a collection in the museum structure display model is as follows. The existing data includes the immersion response index, gaze focus stability index, dwell behavior structure index, and observation path deviation index of five (randomly selected) collections in the museum structure display model, as shown in Table 1: Table 1 Collection sequence user behavior index set of the museum structure display model Immersion adjustment coefficients stored in the database Approximately: 0.641; Gaze focus adjustment coefficients stored in the database Approximately: 1.265; Stay adjustment coefficients stored in the database Approximately: 0.869; Path deviation adjustment coefficient stored in the database Approximately: 0.427; Substituting the data in Table 1 and the above adjustment coefficients into the calculation of the user behavior response index of a collection in the museum structure display model, we obtain: The user behavior response index of the first collection of the museum structure display model = tanh ((exp (0.731 0.641 ×1.265×0.701)×0.869×ln(1+0.627)) / (1+0.264 0.427 ))≈0.506; The user behavior response index of the second collection of the museum structure display model = tanh ((exp (0.843 0.641 ×1.265×0.764)×0.869×ln(1+0.849)) / (1+0.206 0.427 ))≈0.686; The user behavior response index of the third collection of the museum structure display model = tanh ((exp (0.758 0.641×1.265×0.863)×0.869×ln(1+0.813)) / (1+0.186 0.427 ))≈0.699; The user behavior response index of the fourth collection of the museum structure display model = tanh ((exp (0.813 0.641 ×1.265×0.872)×0.869×ln(1+0.753)) / (1+0.217 0.427 ))≈0.687; The user behavior response index of the fifth collection of the museum structure display model = tanh ((exp (0.657 0.641 ×1.265×0.713)×0.869×ln(1+0.672)) / (1+0.287 0.427 ))≈0.509.
[0051] In this implementation plan, by constructing a mapping and comprehensive analysis mechanism from the user behavior index set to the user behavior response index, a deep binding of user behavior data and the museum structure display model is achieved, and an adjustable and fitable adjustment coefficient system is introduced in the response index calculation, so that the behavioral response evaluation results of each collection have a high degree of personalized expression and global comparison capabilities. Secondly, with the help of the collection's unique identifier, the semantic binding of behavioral data is achieved, thereby avoiding data mismatch and index confusion problems. At the same time, based on multi-dimensional user behavior factors such as immersion, gaze, stay, and observation path, the multi-angle measurement capability of behavioral response results is ensured. Finally, through regression modeling and sensitivity analysis, the adjustment coefficients can be adaptively optimized according to different exhibit characteristics, thereby achieving dynamic fitting of user attention, and then improving the display system's quantitative understanding ability of user feedback data and the accuracy of regulatory response.
[0052] Specifically, the specific steps for initially adjusting the display position of each collection of the museum structure display model based on the user behavior response index are as follows: arranging the user behavior response index of each collection of the museum structure display model in descending order to generate a priority display table; performing initial adjustment on each collection of the museum structure display model in the priority display table based on a preset museum priority visual display position set (that is, mapping each collection of the museum structure display model in the priority display table to each priority position in the preset museum priority visual display position set in sequence, and the museum priority visual display position set has the same number as the collections of the museum structure display model in the priority display table).
[0053] In this implementation scheme, by taking the user behavior response index as the sorting basis, a priority display table is constructed and mapped to the preset priority visual display position set, thereby realizing the initial layout optimization of the exhibit position driven by the audience's real attention. Secondly, a dynamic sorting method driven by behavioral data is adopted, so that exhibits that are truly interactive, stay, gaze and other strong attention of the audience can occupy high-exposure, easy-to-reach, visually focused areas in the museum space first, thereby maximizing the efficiency of conveying exhibit information and the interactive potential. Finally, the one-to-one mapping of the display position set and the number of exhibits ensures the controllability and integrity of the display plan, thereby avoiding resource redundancy and waste of vacancies. The initial control link adopts a combination of descending order and mapping allocation, which makes the calculation efficiency high, the execution logic clear, and has good engineering practicality and scalability.
[0054] Specifically, if Figure 3 As shown, the specific steps for obtaining the local semantic coordination index of each collection of the museum structure display model after the initial display position adjustment are as follows: obtaining the spatial distance value between each collection of the museum structure display model after the initial display position adjustment and each collection within the set range (that is, extracting the three-dimensional coordinates of the center of gravity of the component attribute set of each collection of the museum structure display model after the initial display position adjustment and each collection within the set range, and calculating and analyzing based on the Euclidean distance formula); performing semantic embedding coding processing on the component attribute set of each collection of the museum structure display model after the initial display position adjustment to obtain the component vector of each collection of the museum structure display model after the initial display position adjustment; and performing a comprehensive analysis on the component vector and spatial distance value of each collection of the museum structure display model after the initial display position adjustment and each collection within the set range to obtain the local semantic coordination index of each collection of the museum structure display model after the initial display position adjustment.
[0055] The specific steps of semantic embedding encoding processing are as follows: The core attribute fields of all components in each collection in the museum structure display model after the initial display position adjustment are extracted. These fields include but are not limited to component type, component material, functional purpose, connection method, color coding, transparency, surface reflectivity level, component spatial center of gravity coordinates, and minimum bounding box center coordinates. Discrete semantic information such as component type, material, purpose, and connection method is encoded into word vectors using a semantic embedding model. Continuous values such as color and transparency are converted into comparable numerical vectors through normalization. Spatial information is normalized to relative geometric coordinates based on the center position of the exhibit. All attribute vectors are then concatenated according to a fixed field order to form a component-level semantic attribute embedding vector of unified dimension. Next, the embedding vectors of all components in the same collection are weighted and aggregated, where the weight value can be determined based on factors such as component volume proportion, display importance, or whether it is a primary component. A high-dimensional semantic vector representing the overall attribute semantic characteristics of the collection is aggregated through weighted averaging. This high-dimensional semantic vector is then normalized to a unified scale to obtain the component vector of each collection in the museum structure display model after the initial display position adjustment.
[0056] The specific formula for calculating the local semantic coordination index of each collection in the museum structure display model after the initial display position adjustment is as follows: ;in, The first section of the museum structure display model after the initial display position adjustment The local semantic coordination index of the collection, The first section of the museum structure display model after the initial display position adjustment The component vector of the collection, The first section of the museum structure display model after the initial display position adjustment Collection and the first within the set range The component vector of the collection, is the similarity adjustment coefficient stored in the database, The first section of the museum structure display model after the initial display position adjustment Collection and the first within the set range The spatial distance value of the collection, is the distance adjustment coefficient stored in the database, is the interaction adjustment coefficient stored in the database, 1, 2, 3, ..., , is the number of collections, 1, 2, 3, ..., , The number of collections within the set range.
[0057] What needs to be explained is that 、 、 It can be obtained through the following steps: conduct similarity analysis based on historical data to obtain the component similarity between each collection of the museum structure display model after the initial display position adjustment and each collection within the set range, and combine the spatial distance value to conduct statistical regression analysis to quantify the specific impact of each factor on the local semantic coordination index, so as to fit the initial weight value. Secondly, use the sensitivity analysis method to adjust the value range of each coefficient and observe its impact on the local semantic coordination evaluation results to ensure the stability and rationality of the model. Based on the characteristics of the collection and the actual situation, the preliminary fitting coefficients are corrected and optimized, and finally the coefficient value applicable to the specific collection is determined.
[0058] In this implementation, the local semantic coordination index is quantitatively calculated for each collection after initial display position adjustment, effectively achieving coordinated optimization of the semantic attributes and spatial layout between collections in the display model. Secondly, through embedded semantic coding, the component attributes of each collection are converted into a unified high-dimensional vector representation, which enables structured comparison and integration of complex and diverse attribute information. Combined with spatial distance, a dual reference system of semantics and geometry is established. Then, an adjustment coefficient is introduced to control the influence of semantic similarity and spatial proximity on coordination. The resulting local semantic coordination index can fully reflect the display and semantic consistency of the collection within its local area, thereby improving the coherence of the display logic and the continuity of the display semantics between exhibits, while also enhancing the professionalism of the entire display space and the immersiveness of the viewing experience. Finally, the training and optimization process of the adjustment coefficient is based on historical data and sensitivity analysis, thereby ensuring the objectivity, stability and adaptability of the index evaluation to different exhibits, thereby providing highly reliable data support for subsequent display position optimization.
[0059] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0060] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. The AI-based digital cultural space full-process intelligent collaborative display system is characterized by: include: The data acquisition and construction module is used to obtain the 3D point cloud data and text information of several museum collections to be displayed, and input them into the pre-trained structure recognition model for comprehensive analysis to obtain the component attribute set of each museum collection to be displayed and construct the museum structure display model; The behavior acquisition and analysis module is used to obtain the historical video stream time series data of the museum to be displayed, and input it into the pre-trained behavior recognition model for comprehensive analysis to obtain the user behavior index set for each collection of the museum to be displayed; A comprehensive response analysis module is used to input the user behavior index set of each collection to be displayed in the museum into the museum structure display model for comprehensive analysis to obtain the user behavior response index of each collection in the museum structure display model; An initial display control module is used to control the initial display position of each collection in the museum structure display model based on the user behavior response index; The optimized display control module is used to comprehensively analyze the component attribute set of each collection in the museum structure display model after the initial display position adjustment and each collection within the set range, obtain the local semantic coordination index of each collection in the museum structure display model after the initial display position adjustment, and perform optimized display position adjustment.
2. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 1 is characterized in that: The three-dimensional point cloud data specifically includes the voxel value and three-dimensional coordinates of each voxel point. The component attribute set includes the component identifier, spatial position set, semantic attribute set, rendering parameter set, and semantic association set of each component. The structure recognition model is specifically a component attribute recognition network. The multimodal component attribute recognition network includes an input layer, a multimodal guided coding layer, a semantic fusion recognition layer, a semantic attribute enhancement layer, and an attribute set output layer.
3. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 2 is characterized in that: The specific steps to obtain the component attribute set of each collection to be displayed in the museum are as follows: In the input layer of the component attribute recognition network, the 3D point cloud data and text information of each museum collection to be displayed are received and preprocessed; In the multimodal guided encoding layer of the component attribute recognition network, semantic guided encoding is performed on the pre-processed 3D point cloud data and text information of each museum collection to be displayed, thereby obtaining a fused semantic feature set for each museum collection to be displayed. In the semantic fusion recognition layer of the component attribute recognition network, semantic-driven component recognition processing is performed on the fused semantic feature set of each collection of the museum to be displayed, and a component semantic candidate set of each collection of the museum to be displayed is obtained; In the semantic attribute enhancement layer of the component attribute recognition network, attribute calibration is performed on the component semantic candidate set of each collection of the museum to be displayed, thereby obtaining the component semantic attribute set of each collection of the museum to be displayed; In the attribute set output layer of the component attribute recognition network, the component semantic attribute set of each collection to be displayed in the museum is encapsulated and processed to obtain the component identifier, spatial position set, semantic attribute set, rendering parameter set, and semantic association set of each component of each collection to be displayed in the museum.
4. According to the AI-based full-process intelligent collaborative display system for digital cultural spaces according to claim 1, the historical video stream time series data is specifically the historical video stream data of each historical exhibition, and the historical video stream data includes the pixel value and two-dimensional coordinates of each pixel point in several historical frames of exhibition images. The user behavior index set includes an immersion response index, a gaze focus stability index, a stay behavior structure index, and an observation path deviation index. The behavior recognition model is specifically a posture-guided behavior network, which includes a video input layer, a dual-stream time series feature extraction layer, a posture extraction construction layer, a graph convolution behavior recognition layer, and a behavior output layer.
5. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 4 is characterized in that: The specific steps to obtain the user behavior index set for each collection of the museum to be displayed are as follows: In the video input layer of the posture-guided behavior network, the historical video stream time series data of the museum to be displayed is received and preprocessed; In the dual-stream temporal feature extraction layer of the gesture-guided behavior network, the historical video stream temporal data of the museum to be displayed is detected and encoded to obtain a fused temporal behavior feature sequence set of each user in front of each collection at each historical exhibition of the museum to be displayed; In the posture extraction construction layer of the posture guidance behavior network, posture trajectory extraction processing is performed on the fused temporal behavior feature sequence set of each user in front of each collection for each historical exhibition of the museum to be displayed, thereby obtaining the posture temporal feature set of each user in front of each collection for each historical exhibition of the museum to be displayed; In the graph convolutional behavior recognition layer of the gesture-guided behavior network, graph behavior recognition processing is performed on the temporal feature set of the gestures of each user in front of each collection at each historical exhibition of the museum to be displayed, thereby obtaining the semantic feature set of the behavior of each user in front of each collection at each historical exhibition of the museum to be displayed; In the behavior output layer of the gesture-guided behavior network, branch prediction processing is performed on the behavioral semantic feature set of each user in front of each collection in each historical exhibition of the museum to be displayed, and the immersive response index, gaze focus stability index, stay behavior structure index, and observation path deviation index of each collection of the museum to be displayed are obtained.
6. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 4 is characterized in that: The specific steps to obtain the user behavior response index of each collection in the museum structure display model are as follows: Mapping the user behavior index set of each collection of the museum to be displayed based on the museum structure display model to obtain the user behavior index set of each collection of the museum structure display model; A comprehensive analysis is conducted on the user behavior index set of each collection of the museum structure display model to obtain the user behavior response index of each collection of the museum structure display model.
7. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 4 is characterized in that: The specific formula for calculating the user behavior response index of a collection in the museum structure display model is as follows: ; in, 、 、 、 、 These are the user behavior response index, immersion response index, gaze focus stability index, stay behavior structure index, and observation path deviation index of a collection of the museum structure display model. 、 、 、 They are the immersion adjustment coefficient, gaze focus adjustment coefficient, stay adjustment coefficient, and path deviation adjustment coefficient stored in the database.
8. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 1 is characterized in that: The specific steps for adjusting the initial display position of each collection in the museum structure display model based on the user behavior response index are as follows: Arrange the user behavior response index of each collection in the museum structure display model in descending order to generate a priority display table; Based on the preset museum priority visual display position set, each collection of the museum structure display model in the priority display table is initially regulated.
9. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 1 is characterized in that: The specific steps for obtaining the local semantic coordination index of each collection in the museum structure display model after the initial display position adjustment are as follows: Obtain the spatial distance value between each collection in the museum structure display model after the initial display position adjustment and each collection within the set range; Performing semantic embedding coding on the component attribute set of each collection in the museum structure display model after the initial display position adjustment, to obtain a component vector of each collection in the museum structure display model after the initial display position adjustment; A comprehensive analysis is then conducted on the component vectors and spatial distance values of each collection in the museum structure display model after the initial display position adjustment and each collection within the set range to obtain the local semantic coordination index of each collection in the museum structure display model after the initial display position adjustment.
10. The AI-based digital cultural space full-process intelligent collaborative display system according to claim 9 is characterized in that: The specific formula for calculating the local semantic coordination index of each collection in the museum structure display model after the initial display position adjustment is as follows: ; in, The first section of the museum structure display model after the initial display position adjustment The local semantic coordination index of the collection, The first section of the museum structure display model after the initial display position adjustment The component vector of the collection, The first section of the museum structure display model after the initial display position adjustment Collection and the first The component vector of the collection, is the similarity adjustment coefficient stored in the database, The first section of the museum structure display model after the initial display position adjustment Collection and the first The spatial distance value of the collection, is the distance adjustment coefficient stored in the database, is the interaction adjustment coefficient stored in the database, 1, 2, 3, ..., , is the number of collections, 1, 2, 3, ..., , The number of collections within the set range.
Citation Information
Cited By
AI-based three-dimensional model display method and system
CN120823327A
Dynamic exhibition system for digital museum exhibits
CN121209695A