Interaction method and system applied to painting and calligraphy exhibition
Through multimodal sensing arrays and coordinate-label mapping networks, identify exhibits that users are concerned about, and generate personalized guided paths, solving the problem of lack of real-time perception in the existing technology, and improving the interactivity and immersion of calligraphy and painting exhibitions.
Patent Information
- Application Number
- CN202510766307.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing interactive methods of calligraphy and painting exhibitions lack real-time perception and accurate identification of users' exhibition behavior, resulting in the inability to realize personalized exhibit linkage and dynamic guide, affecting the audience's immersion and interactive experience.
Through a multimodal sensing array, a user behavior data and exhibit environment data are collected in real time, a user exhibition behavior data set is constructed, and a coordinate-label mapping network is established through multi-dimensional image calibration, target exhibit numbers are identified, and a personalized guide path is generated to realize the coordinated display and in-depth interaction of exhibits.
Real-time perception and personalized guide for user interests are realized, the interactive and intelligent level of calligraphy and painting exhibitions are improved, and the immersive experience of the audience is enhanced.
Smart Images

Figure CN120279221A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data interaction, and particularly to an interaction method and system applied to a calligraphy and painting exhibition. Background Art
[0002] During the exhibition process of calligraphy and painting art, the interactive experience has gradually become an important means to enhance the audience's sense of participation and artistic appeal. With the development of digital technology, the traditional calligraphy and painting exhibition mode mainly based on static display is gradually evolving towards informatization and intelligence. The existing interactive methods for calligraphy and painting exhibitions are mostly based on methods such as QR code scanning, electronic tour guides, or touch screens to assist the audience in understanding the information of the exhibits. However, most of these interactive methods are triggered passively, lacking real-time perception and in-depth understanding of the user's exhibition viewing behavior, thus making it difficult to conduct targeted linkage display of exhibits and information push. At the same time, the organization form of the exhibition content is relatively fixed, and the audience is prone to information overload or miss the exhibits they are interested in, restricting the in-depth interaction between the audience and the exhibits and affecting the exhibition viewing interactive experience and exhibition efficiency. Summary of the Invention
[0003] The present invention provides an interaction method and system applied to a calligraphy and painting exhibition, which solves the technical problem in the prior art that due to the lack of real-time perception and accurate recognition of the user's exhibition viewing behavior, personalized exhibit linkage and dynamic navigation cannot be realized, and achieves the technical effect of automatically identifying target exhibits based on the user's interest and generating a personalized navigation path, thereby enhancing the immersion and interactive experience of the exhibition viewing.
[0004] In view of the above problems, on the one hand, the present invention provides an interaction method applied to a calligraphy and painting exhibition. The method includes: performing real-time acquisition on a target user based on the exhibit area, synchronously obtaining user behavior data and exhibit environment data through a multi-modal sensing array, and constructing a user exhibition viewing behavior data set; traversing the exhibit area for multi-dimensional image calibration, where the multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data, and texture feature data on the surface of the exhibit, and establishing a coordinate-label mapping network through spatio-temporal association; traversing the coordinate-label mapping network set according to the user exhibition viewing behavior data set to perform exhibit node matching, activating an exhibit recognition module to identify the target exhibit number; performing a linkage display of calligraphy and painting for the target user according to the target exhibit number, generating display information, starting an exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, generating a personalized calligraphy and painting exhibition navigation path for the target user, and dynamically adjusting the navigation path in real time to adapt to the dynamic behavior of the user.
[0005] Preferably, based on the exhibition area, real-time acquisition of target users is carried out, and user behavior data and exhibition environment data are synchronously obtained through a multi-modal sensing array to construct a user exhibition behavior data set. The method includes: deploying a multi-modal sensing array around the perimeter of the calligraphy and painting exhibits in the exhibition area, and obtaining a user sensing data set by capturing the target users in real time through the multi-modal sensing array; establishing a mapping relationship between the exhibit coordinate system and the global coordinate system of the exhibition hall, and converting the user sensing data into exhibition behavior metadata; performing spatio-temporal alignment processing on the exhibition behavior metadata according to the mapping relationship to generate a three-dimensional behavior vector; performing semantic parsing based on the three-dimensional behavior vector to generate a semantic association graph, and analyzing the user behavior according to the semantic association graph to construct multi-source behavior event data, and adding the multi-source behavior event data to the user exhibition behavior data set.
[0006] Preferably, traverse the exhibition area for multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data and texture feature data on the surface of the exhibit. A coordinate-label mapping network is established through spatio-temporal association. The method includes: synchronously collecting exhibits through the multi-modal sensing array traversing the exhibition area to obtain an exhibit sensing data set, where the exhibit sensing data set includes multi-spectral image data and three-dimensional point cloud data on the surface of the exhibit; mapping the exhibit sensing data set to the global coordinate system of the exhibition hall to generate an exhibit space topology map; performing association mapping on the three-dimensional point cloud data and the multi-spectral image data on the surface of the exhibit according to the exhibit space topology map to generate an association label set; calibrating the calligraphy and painting exhibits based on the association label set to construct the coordinate-label mapping network.
[0007] Preferably, traverse the coordinate-label mapping network set according to the user exhibition behavior data set to activate the exhibit recognition module to identify the target exhibit number. The method includes: performing motion trajectory analysis based on the user exhibition behavior data set to extract spatio-temporal behavior feature vectors, and performing gaze tracking based on the user exhibition behavior data set to draw a gaze focus heat map; extracting multiple exhibit node information based on the association label set; traversing the coordinate-label mapping network to perform matching calculations on the spatio-temporal behavior feature vectors and the exhibit node information to obtain multiple matching confidence levels; making a determination according to the multiple matching confidence levels in combination with the gaze focus heat map, and activating the exhibit recognition module according to the determination result to obtain the target exhibit number.
[0008] Preferably, perform motion trajectory analysis based on the user exhibition behavior data set to extract spatio-temporal behavior feature vectors. The method includes: performing spatial displacement analysis based on the user exhibition behavior data set to obtain the behavior trajectory sequence of the target user; classifying the trajectory according to the behavior trajectory sequence to obtain multiple movement mode categories, and performing behavior mining according to the multiple movement mode categories to obtain spatio-temporal behavior feature vectors.
[0009] Preferably, to extract multiple exhibit node information based on the associated tag set, the method includes: performing feature analysis on the basis of the associated tag set in combination with the three-dimensional point cloud data to obtain point cloud geometric features; performing feature analysis on the basis of the associated tag set in combination with the hyperspectral image data to obtain spectral texture features; performing multimodal feature fusion on the point cloud geometric features and the spectral texture features according to the associated tag set to construct an associated feature vector; and performing hierarchical node parsing according to the associated feature vector to determine the multiple exhibit node information.
[0010] Preferably, to make a determination in combination with the line-of-sight focus heat map according to the multiple matching confidence levels, and activate the exhibit recognition module according to the determination result to obtain a target exhibit number, the method includes: traversing the line-of-sight focus heat map to extract eye movement feature data, and calculating the regional residence duration according to the eye movement feature data; performing exhibit coordinate screening according to the multiple matching confidence levels to generate a candidate exhibit set; and when the regional residence duration is greater than or equal to a preset duration, activating the exhibit recognition module to traverse the candidate exhibit set for number recognition to determine the target exhibit number.
[0011] Preferably, to perform a linked display of calligraphy and painting for the target user according to the target exhibit number to generate display information, start the exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, generate a personalized calligraphy and painting exhibition tour path for the target user, and adjust the tour path in real time to adapt to the user's dynamic behavior, the method includes: calling multimodal sensing data for dynamic synthesis according to the target exhibit number to obtain a multimodal synthesis data set; performing a linked display of calligraphy and painting for the target user according to the multimodal synthesis data set to generate display information, where the display information includes audio display information, brush stroke trajectory display information, and three-dimensional holographic image display information; starting the exhibit interaction module to continuously sense the multi-dimensional interaction data set of the target user, where the multi-dimensional interaction data set includes voice interaction data, motion interaction data, and eye movement interaction data; performing interaction based on the voice interaction data in combination with the audio display information to generate a first interaction perception parameter; performing interaction based on the motion interaction data in combination with the brush stroke trajectory display information to generate a second interaction perception parameter; performing interaction based on the eye movement interaction data in combination with the three-dimensional holographic image display information to generate a third interaction perception parameter; and performing a tour analysis of the calligraphy and painting exhibition for the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to generate a personalized calligraphy and painting exhibition tour path.
[0012] Preferably, based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, a tour guide analysis of a calligraphy and painting exhibition is performed on the target user to generate a personalized calligraphy and painting exhibition tour path. The method includes: performing a dynamic interest analysis on the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to construct a user interest vector; optimizing the path according to the user interest vector in combination with the exhibit coordinate system to construct a movement path; and mapping the movement path to the global coordinate system of the exhibition hall for dynamic guidance marking to generate the personalized calligraphy and painting exhibition tour path.
[0013] On the other hand, the present invention also provides an interaction system applied to a calligraphy and painting exhibition. The system includes: a behavior data acquisition module for performing real-time acquisition on the target user based on the exhibit area, synchronously obtaining user behavior data and exhibit environment data through a multi-modal sensing array, and constructing a user exhibition viewing behavior data set; an image calibration module for traversing the exhibit area for multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data, and texture feature data on the surface of the exhibit, and establishing a coordinate-label mapping network through spatio-temporal association; a target recognition module for traversing the coordinate-label mapping network set according to the user exhibition viewing behavior data set to perform exhibit node matching, activating the exhibit recognition module, and identifying the target exhibit number; and an interactive display module for performing a linked display of calligraphy and painting on the target user according to the target exhibit number to generate display information, starting the exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, generating a personalized calligraphy and painting exhibition tour path for the target user, and dynamically adjusting the tour path in real time to adapt to the dynamic behavior of the user.
[0014] One or more technical solutions provided in the present invention have at least the following beneficial effects: By using a sensing device to perform real-time acquisition on the behavior of the audience at the exhibition site and constructing a user exhibition viewing behavior data set, it provides basic data support for subsequent identification and personalized recommendation. By traversing the exhibit area for multi-dimensional image calibration and constructing a coordinate-label mapping network between the exhibit space coordinates and semantic information, it creates conditions for the accurate matching between user behavior and exhibits. By traversing the coordinate-label mapping network set according to the user exhibition viewing behavior data set to activate the exhibit recognition module, automatically identifying the target exhibit number that the user is currently interested in, and realizing the association conversion from behavior data to exhibit entities. Performing a linked display of calligraphy and painting on the target user according to the target exhibit number to generate display information, starting the exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, and generating a personalized calligraphy and painting exhibition tour path for the target user, enhancing the immersive exhibition viewing experience.
[0015] In summary, the present invention realizes real-time perception of user interests and personalized guided tour services in calligraphy and painting exhibitions by constructing a user behavior dataset, an image calibration coordinate-label mapping network, and an exhibit recognition and personalized linkage display mechanism based on behavior data, significantly improving the interactivity, intelligence level, and user immersive experience of the exhibition.
[0016] The above description is only an overview of the technical solution of the present invention. In order to be able to more clearly understand the technical means of the present invention, it can be implemented in accordance with the content of the specification. And in order to make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following specifically gives the specific implementation manners of the present invention. Brief Description of the Drawings
[0017] Figure 1 It is a schematic flowchart of an interaction method applied to a calligraphy and painting exhibition provided by an embodiment of the present invention.
[0018] Figure 2 It is a schematic flowchart of identifying the target exhibit number by traversing the coordinate-label mapping network set according to the user exhibition behavior dataset in an interaction method applied to a calligraphy and painting exhibition provided by an embodiment of the present invention.
[0019] Figure 3 It is a schematic structural diagram of an interaction system applied to a calligraphy and painting exhibition provided by an embodiment of the present invention.
[0020] Description of the reference numerals: Behavior data acquisition module 10, image calibration module 20, target recognition module 30, interaction display module 40. Detailed Description of the Invention
[0021] The embodiment of the present invention provides an interaction method and system applied to a calligraphy and painting exhibition, solving the technical problem in the prior art that due to the lack of real-time perception and accurate recognition of the user's exhibition behavior, the personalized exhibit linkage and dynamic guided tour cannot be realized, and achieving the technical effect of automatically identifying the target exhibit based on the user's interest and generating a personalized guided tour path, thereby enhancing the immersive feeling and interaction experience of the exhibition.
[0022] Embodiment 1, as Figure 1 shown, the embodiment of the present invention provides an interaction method applied to a calligraphy and painting exhibition, and the method includes: Step S100: Based on the exhibit area, perform real-time acquisition on the target user, and synchronously obtain user behavior data and exhibit environment data through a multi-modal sensing array to construct a user exhibition behavior dataset.
[0023] Specifically, the exhibition area refers to the physical range within the exhibition space where the calligraphy and painting works are located. Within the exhibition area, a multi-modal sensing array is deployed to synchronously collect user behavior data and exhibition environment data, enabling accurate modeling of users' exhibition viewing behaviors. This multi-modal sensing array integrates multiple sensing devices, including but not limited to: RGB cameras (collecting images and human postures), infrared depth cameras (measuring spatial positions), eye trackers (capturing fixation points and gaze directions), motion capture devices (tracking limb movements), and voice sensors (such as microphones). By using the multi-modal sensing array to record the interaction behaviors of each viewer facing each exhibit, an exhibition viewing behavior dataset containing temporal sequence, position information, and eye movement features is constructed, achieving accurate modeling of users' exhibition viewing behaviors and providing a high-reliability input data basis for subsequent exhibit recognition and personalized guided tours.
[0024] Step S200: Traverse the exhibition area for multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data of the exhibit surface, three-dimensional point cloud data, and texture feature data. A coordinate-label mapping network is established through spatio-temporal association.
[0025] Specifically, traversing the exhibition area for multi-dimensional image calibration means using multiple sensing methods to scan, identify, and model each exhibit one by one to obtain the perceptual data of the exhibit in different dimensions. This perceptual data includes but is not limited to multi-spectral image data of the exhibit surface, three-dimensional point cloud data, and texture feature data. The multi-spectral image data of the exhibit surface is obtained by using a multi-spectral camera to image the surface of the exhibit, acquiring the exhibit image information in different bands such as visible light, near-infrared (NIR), and ultraviolet (UV), reflecting the detailed features such as the color, material, and texture of the calligraphy and painting works; the three-dimensional point cloud data is the spatial depth information of the exhibit collected using structured light or lidar technology, describing the geometric contour and spatial position of the exhibit and used to construct its three-dimensional geometric model. The texture feature data is the surface texture features extracted from the above multi-spectral image data, used to describe the gray-scale changes, edge directions, corner distributions, etc. of local image segments.
[0026] To achieve the joint modeling of image data and spatial data, the above multi-source data (multi-spectral image data of the exhibit surface, three-dimensional point cloud data, and texture feature data) are fused through a spatio-temporal synchronization mechanism. In the unified coordinate system of the exhibition hall, a mapping relationship is established between the structural coordinate points of each exhibit (such as key points, edges, textures, etc. on the exhibit surface) and their corresponding semantic labels, forming a structured multi-label spatial semantic network, that is, a coordinate-label mapping network, enabling each position point in space to accurately correspond to a certain exhibit and its specific content for subsequent user behavior positioning and exhibit recognition.
[0027] Step S300: Traverse the coordinate-label mapping network set according to the user exhibition viewing behavior data set to match the exhibit nodes, activate the exhibit recognition module, and identify the target exhibit number.
[0028] Specifically, the exhibit recognition module is a functional module for analyzing user behavior and determining the exhibit number that the user is currently interested in. Input the user exhibition viewing behavior data (line of sight direction, position trajectory) in Step S100 into the coordinate-label mapping network for traversal analysis. According to the line of sight and the staying position, perform confidence matching calculation with the spatial positions of each exhibit in the coordinate-label mapping network, and activate the exhibit recognition module to determine the target exhibit number that the user is currently interested in. Through precise matching calculation, the exhibit that the user is interested in can be identified in real time, providing a basis for subsequent targeted display and interaction.
[0029] Step S400: Perform linked display of calligraphy and paintings for the target user according to the target exhibit number, generate display information, activate the exhibit interaction module, conduct in-depth interaction with the target user in combination with the display information, generate a personalized calligraphy and painting exhibition tour path for the target user, and adjust the tour path in real time to adapt to the dynamic behavior of the user.
[0030] Specifically, the exhibit interaction module is a functional component responsible for collecting user feedback and performing interactive control. The display information includes a series of content information for displaying to the user related to calligraphy and painting exhibits, such as high-definition images of calligraphy and paintings, audio introductions of the author's life, video analysis of painting techniques, three-dimensional holographic images, etc. According to the identified exhibit number, call the multi-modal content (audio introduction, three-dimensional model, brushstroke dynamics, etc.) of the exhibit for linked display on the interactive terminal (such as a large screen, AR glasses), generate display information, and achieve immersive and multi-modal human-computer interaction. At the same time, the interaction module captures the user's voice questions, gesture actions or line of sight change feedback in real time, generates multi-dimensional interaction perception parameters. Based on these parameters, dynamically evaluate the user's interest tendency, intelligently adjust its recommended path, and finally generate a calligraphy and painting exhibition tour route that better meets the user's individual needs, and guide its subsequent exhibition viewing through a wayfinding system (such as projection or navigation marks), greatly improving the sense of participation and satisfaction of the exhibition.
[0031] Furthermore, Step S100 includes: Step S110: Deploy a multi-modal sensing array around the perimeter of the calligraphy and painting exhibits in the exhibit area, and perform real-time capture of the target user through the multi-modal sensing array to obtain a user sensing data set.
[0032] Step S120: Establish a mapping relationship between the exhibit coordinate system and the global coordinate system of the exhibition hall, and convert the user sensing data into exhibition viewing behavior metadata.
[0033] Step S130: Perform spatio-temporal alignment processing on the exhibition behavior metadata according to the mapping relationship to generate a three-dimensional behavior vector.
[0034] Step S140: Perform semantic parsing based on the three-dimensional behavior vector to generate a semantic association graph, conduct user behavior analysis according to the semantic association graph to construct multi-source behavior event data, and add the multi-source behavior event data to the user exhibition behavior dataset.
[0035] Specifically, the deployment of the multi-modal sensing array needs to ensure coverage of the areas directly in front of, on the left and right sides, and at the top of the exhibits to avoid sensing blind spots, and achieve joint sampling of multi-modal devices through a time synchronization protocol. When the target user enters the sensing range, start the synchronous acquisition of the multi-modal sensing array, record the behavior data including their position, orientation, line of sight, fixation time, interaction behavior, etc., and encapsulate it into a user sensing dataset. The user sensing dataset is initially marked according to the time stamp and associated and bound with the user's unique identity (user ID) to provide the original input for subsequent data processing.
[0036] The exhibit coordinate system is a local coordinate system established with each calligraphy and painting exhibit as the origin, and the exhibition hall global coordinate system is a unified three-dimensional reference system constructed according to the overall exhibition hall. Use a panoramic positioning camera, lidar, or panoramic depth camera device to perform three-dimensional scanning of the exhibition space and establish a three-dimensional exhibition hall point cloud map as the basis of the global coordinate system. Among them, the origin is usually set at the center of the exhibition hall entrance, the Z-axis points upward, and the X and Y axes are respectively extended along the main channels of the exhibition hall. At the same time, arrange several positioning reference points (such as QR code targets, laser reflection balls) in the exhibition hall and record their spatial coordinates in the global coordinate system for subsequent spatial alignment and attitude correction of the data. Subsequently, perform laser calibration on the position and attitude of each exhibit in the exhibition hall, and use the homogeneous transformation matrix to align the local coordinate systems of each exhibit to the unified exhibition hall global coordinate system to achieve unified reference for spatial positions. Project or transform the sensing data such as the user's real-time position and line of sight into the global coordinate system and convert it into exhibition behavior metadata in a unified format, including: user ID, time stamp, position coordinates (X u , Y u , Z u ), sensing type, original exhibition feature value, etc., to provide a unified coordinate basis for the alignment and fusion of multi-time period and multi-sensing channel data.
[0037] Based on the global coordinate system, the multimodal data of the same user within adjacent time windows are sorted and aggregated according to timestamps, and data frame alignment is performed using a unified time window ΔT (such as 100 ms). Each piece of metadata of the exhibition viewing behavior is represented as a three-dimensional vector structure: behavior timestamp, position coordinates in the global coordinate system of the exhibition hall, and interaction events. This enables the user's behavior data to be accurately described and analyzed under a unified time axis and spatial framework. Among them, the interaction event refers to the specific type of exhibition viewing interaction operation performed by the user at that moment and at that position, such as gazing at an exhibit, staying in front of an exhibit, and having a touch interaction with an exhibit.
[0038] Semantic label inference is performed on multiple three-dimensional behavior vectors of the target user to identify the time, space, and interaction pattern features in the user's behavior. Through the set behavior rule library and semantic label inference mechanism, typical behaviors are classified into events. For example, if the position change threshold is less than the preset threshold, it is judged as "staying", and if the fixation points are concentrated, it is judged as "gazing". Combinations of various semantic labels form a semantic association graph in the form of nodes, and the edges between different nodes represent the time sequence or semantic transfer logic of the user's behavior. Further, the node information in the semantic graph is extracted and merged to form a structured multi-source behavior event data record. For example, a combination of judging that the gazing + staying time > 5 seconds can be classified as a "dwelling" event; gazing + gestures or speech constitutes an "active interaction" event. Finally, these event data are added to the user's exhibition viewing behavior dataset, providing a rich behavior basis for subsequent personalized content push and behavior pattern analysis.
[0039] Furthermore, step S200 includes: Step S210: Synchronously collect exhibits through the multimodal sensing array traversing the exhibit area to obtain an exhibit sensing dataset, where the exhibit sensing dataset includes multi-spectral image data and three-dimensional point cloud data on the surface of the exhibits.
[0040] Step S220: Map the exhibit sensing dataset to the global coordinate system of the exhibition hall to generate an exhibit space topology map.
[0041] Step S230: Correlate and map the three-dimensional point cloud data and the multi-spectral image data on the surface of the exhibits according to the exhibit space topology map to generate an association label set.
[0042] Step S240: Calibrate the calligraphy and painting exhibits based on the association label set to construct the coordinate-label mapping network.
[0043] Specifically, in addition to the aforementioned sensors for collecting user behavior data and exhibit environment data, the multimodal sensing array further includes multispectral imaging devices (such as VNIR and UV cameras) and 3D point cloud scanning devices (such as lidar and structured light depth cameras). After the exhibits are arranged, relying on the multimodal sensing array, each exhibit is collected: the multispectral imaging devices and 3D point cloud scanning devices are clock-synchronized through a unified time synchronization module, and a trigger signal is uniformly sent to achieve multimodal synchronous collection of the exhibits. During the collection process, the multispectral module obtains a sequence of reflected images of the exhibits in the specified wavelength band, i.e., multispectral image data, in a continuous scanning or shutter acquisition manner, while the point cloud scanning device obtains the complete 3D contour data of the exhibits through a multi-frame synthesis algorithm. Finally, the original data of each modality is packaged to form an exhibit sensing dataset, which is identified and archived by a unique exhibit number.
[0044] During the exhibit data collection process, a positioning camera or a laser ranging device is used to obtain the actual installation position and pose angle of the exhibits in the exhibition hall, calculate the transformation matrix from the local coordinate system of the exhibits to the global coordinate system of the exhibition hall, and then apply this transformation matrix to each point in the exhibit point cloud data for coordinate transformation, and uniformly map it to the global coordinate system to obtain a set of spatial points of all exhibits under a unified spatial reference. By calibrating the camera pose of the multispectral image data, the multispectral image data is mapped to the global coordinate system to obtain the spatial position of the image perception area, and a geometric projection model of the exhibit image is constructed. Cluster analysis or bounding box extraction is performed on the point cloud data of all exhibits in the global coordinate system of the exhibition hall to construct an exhibit spatial boundary body, where each exhibit corresponds to a geometric boundary. Then, according to the spatial proximity relationship (Euclidean distance between the center points), line-of-sight occlusion relationship or theme association information between the exhibits, the edge connection relationship between the exhibits is constructed to form an exhibit spatial topology graph G=(V, E), where: V is the node set, representing all exhibits; E is the edge set, representing the spatial or semantic adjacency relationship between the nodes, and each edge can be attached with attributes: distance, visibility weight, theme label overlap degree, etc. The final exhibit spatial topology graph is stored in a graph database or a structured JSON manner to achieve a standardized expression of the exhibit spatial position and support spatial logic analysis between multiple exhibits and regional reasoning of the visit path.
[0045] Using image processing algorithms, according to the exhibit spatial topology graph, texture descriptors (such as SIFT, SURF, ORB) in the image region corresponding to each point cloud point in the exhibit point cloud data are extracted, and the extracted texture feature vectors are denoted as {S}={s1, s2, ……, s k}, where each element is a texture descriptor of a local area on the surface of the exhibit. At the same time, semantic labels are assigned to each of the above point cloud points based on the exhibit metadata or an artificial annotation system to generate an associated label set of the exhibits, and the structure is as follows: L tag={(x i ,y i ,z i ,S i ,l i )}, where (x i ,y i ,z i ) are the spatial coordinates of each point of the exhibit, S i is the image texture feature data vector, and l i is the semantic label, such as "flower-and-bird painting", "portrait", "calligraphy of the Qing Dynasty", etc.
[0046] After generating the labels for each area of the exhibit, the area coordinates of the exhibit are corresponded to their associated labels one by one to construct a structured coordinate-label mapping network. The mapping network is stored and managed in the form of a graph structure or a key-value index. Each exhibit is a node in the network, and it contains multiple child nodes, corresponding to the local area labels, their spatial coordinates, image features, and semantic descriptions of the exhibit respectively. The constructed coordinate-label mapping network is essentially a semantic structure grid based on the surface of the exhibit, which can not only accurately calibrate the spatial structure of the exhibit, but also attach semantics to each position point. Exemplarily, a calligraphy and painting exhibit is numbered P001; through the analysis of three-dimensional point cloud data and multi-spectral image data, it is identified that it contains: structural area A (coordinate points A1 to A50) - semantic label: "main landscape"; area B (coordinate points B1 to B20) - label: "postscript"; area C (coordinate points C1 to C30) - label: "author's seal"; area D (coordinate points D1 to D100) - label: "human face". The coordinate-label mapping network can support the fast coordinate reverse lookup label function through a tree-like index structure or a spatial hash function, realize the matching between any observation point and the label, and provide efficient matching support for subsequent user behavior recognition and linkage display.
[0047] Further, as Figure 2 shown, step S300 includes: Step S310: Perform motion trajectory analysis based on the user's exhibition viewing behavior dataset, extract spatio-temporal behavior feature vectors, and perform gaze tracking based on the user's exhibition viewing behavior dataset to draw a gaze focus heat map.
[0048] Step S320: Extract multiple exhibit node information based on the associated label set.
[0049] Step S330: Traverse the coordinate-label mapping network to perform matching calculations on the spatio-temporal behavior feature vectors and the exhibit node information, and obtain multiple matching confidence levels.
[0050] Step S340: performing a determination according to the multiple matching confidences combined with the sight focus heat map, activating the exhibit identification module according to the determination result, and obtaining the target exhibit number.
[0051] Specifically, the spatiotemporal behavior feature vector is a feature representation of the user's movement and stay behavior in the exhibition hall space, including dimensional data such as timestamp, location coordinates, and behavior type (such as stay, turn). The sight focus heat map is a visual map generated based on the cumulative intensity of the user's gaze point in the time and space dimensions, reflecting the user's focus area. The user's spatial position and eye movement data recorded in the user's exhibition behavior data set are analyzed, and the trajectory processing algorithm based on Kalman filtering and sliding window smoothing is used to denoise and interpolate the user's continuous position data in the exhibition hall, so as to construct a high-precision user motion trajectory curve, and combine the time axis for time weighting to generate a spatiotemporal behavior feature vector sequence. At the same time, the eye movement sensor data is called to obtain the gaze point, and the algorithm based on gaze clustering and regional voting is used to determine the high-frequency gaze area, and the gaze point is mapped to the global coordinate system of the exhibition hall. The gaze frequency and gaze duration of a certain spatial area in unit time are counted, and the heat map is rendered to construct a user's sight focus heat map, where the color saturation represents the superposition intensity of the gaze frequency and duration.
[0052] Parse the established coordinate-label mapping network, traverse each exhibit main node and its sub-label nodes, and extract the exhibit number and regional coordinate boundary, corresponding texture type and semantic identification information of each exhibit in turn. To improve the matching efficiency, KD-Tree or hash index is used to spatially group and label the exhibit nodes, thereby forming an exhibit node set containing position features, image features and semantic descriptions. At the same time, the display hierarchy relationship and exhibition area attribution information of the exhibits are marked to prepare for subsequent precision control and intra-regional constraint matching.
[0053] The matching confidence is a quantitative indicator representing the degree of matching between the current user behavior data and a certain exhibit. The value range is generally [0, 1], and the higher the value, the more credible the matching. Based on the coordinate-label mapping network, all exhibit nodes are traversed, and for each exhibit node and its sub-region labels, multi-dimensional matching calculations are sequentially performed with the user's spatio-temporal behavior vector. The matching function comprehensively considers three dimensions: the spatial distance metric Ds (such as the inverse function of the Euclidean distance), the time synchronization weight Wt (the weight of more recent behaviors is higher), and the semantic feature similarity Se (such as the consistency score of texture types or fixation area labels). The final matching confidence is calculated through normalized linear combination, such as: Cf = α*(1 / Ds)+β*Wt+γ*Se, where α, β, and γ are empirically set weight coefficients, satisfying α+β+γ = 1, and α, β, γ ∈ [0.2, 0.5]. The spatial distance metric Ds is the Euclidean distance between the user's current position and the center point of the exhibit. The closer the user is to the exhibit, the greater the likelihood of viewing the exhibit, and the higher the matching confidence; the time synchronization weight Wt reflects the time freshness of the behavior data, with a value range of [0, 1], and the calculation formula is defined as Wt = e -λ*Δt , where Δt is the interval between the current time and the behavior recording time, and λ is the decay coefficient. The closer the time, the greater the weight, which can highlight the exhibit being viewed currently; the semantic feature similarity Se is used to measure the matching degree between the user's fixation area and the exhibit label, and can be calculated by the cosine similarity of texture features, reflecting the consistency between the user's fixation area and the semantic label of the exhibit. Similar semantics indicate credible matching. For each candidate exhibit node, the above three matching dimensions are calculated separately and substituted into the formula to obtain a normalized matching confidence Cf ∈ [0, 1]. Sort them in descending order of confidence, and select the top N candidate exhibit nodes as the subsequent determination candidate set, where N is the capacity of the candidate exhibit nodes, which can be set as a positive integer between 3 and 10, and is comprehensively set according to the exhibition hall density and system response time.
[0054] After obtaining multiple matching confidence results, further screen the exhibits in the overlapping area in the candidate set by combining the position of the heat map area and the physical coordinates of the candidate exhibits, and activate the exhibit recognition module. This module traverses the exhibits in the current overlapping area, and comprehensively scores multiple candidate exhibits based on the matching confidence results, the line-of-sight focus heat map, and the spatial overlap rate of each exhibit node. Select the exhibit node with the highest score as the final determination target, identify the number of the exhibit, and output it as the target exhibit number. By fusing two types of complementary information, confidence ranking and line-of-sight focus map, high-precision recognition of exhibits of viewing interest is achieved while ensuring real-time performance, significantly improving the accuracy and personalized recommendation ability of interaction triggering.
[0055] Furthermore, step S310 includes: Step S311: Perform spatial displacement analysis based on the user's exhibition viewing behavior data set to obtain the behavior trajectory sequence of the target user.
[0056] Step S312: Classify the trajectories according to the behavior trajectory sequence to obtain multiple moving mode categories, and perform behavior mining based on the multiple moving mode categories to obtain spatio-temporal behavior feature vectors.
[0057] Specifically, obtain the continuous position point data of the user during the exhibition viewing process from the user exhibition viewing behavior dataset, and sort it in sequence in combination with the time stamp. Use the sliding window filtering algorithm to smooth the position data and eliminate instantaneous abnormal drift values. At the same time, introduce a micro-displacement determination mechanism to label the stay points (such as when the speed is lower than the threshold and stays for more than the preset time). Finally, structure and store the displacement data of the user's complete exhibition viewing process as a trajectory sequence according to the time axis. Each trajectory point includes: time stamp t, three-dimensional coordinates (x, y, z), displacement vector Δd, and stay flag r, forming the user behavior trajectory sequence.
[0058] Collect a large number of historical user behavior trajectory sequences, and perform segmented coding and multi-dimensional feature extraction, including average segment speed, turning frequency, stay point distribution density, maximum stay time, and trajectory closure degree, etc., and construct a sample feature matrix. Use an unsupervised clustering algorithm (such as DBSCAN or K-Means) to perform trajectory pattern clustering on the feature matrix, and divide it into several common moving mode categories, such as "sequential exhibition tour type", "theme jump type", "high stay intensive reading type", etc. Construct a typical behavior template for each type of moving mode as a reference for subsequent classification. Perform segmented coding and multi-dimensional feature extraction on the user behavior trajectory sequence obtained above, extract features such as the average segment speed, turning frequency, stay point distribution density, maximum stay time, and trajectory closure degree of the target user, and construct a feature matrix. Subsequently, traverse the typical behavior template according to the feature matrix, match the moving mode category corresponding to the target user, and perform behavior mining based on the feature emphasis under this category, extract behavior indicators such as stay-transfer ratio, interest center drift rate, average fixation density, etc., and fuse these behavior indicators to generate spatio-temporal behavior feature vectors, which are used as the basic input for subsequent heat map modeling and exhibit matching, significantly improving the accuracy and adaptability of user behavior modeling.
[0059] Furthermore, step S320 includes: Step S321: Perform feature analysis based on the associated tag set in combination with the three-dimensional point cloud data to obtain point cloud geometric features.
[0060] Step S322: Perform feature analysis based on the associated tag set in combination with the multi-spectral image data to obtain spectral texture features.
[0061] Step S323: Perform multi-modal feature fusion on the point cloud geometric features and the spectral texture features according to the associated tag set to construct an associated feature vector.
[0062] Step S324: Perform hierarchical node parsing according to the associated feature vector to determine the multiple exhibit node information.
[0063] Specifically, based on the established associated tag set, call the subset of 3D point cloud data associated with each exhibit tag. Then, use the voxel grid downsampling algorithm to preprocess the original point cloud to compress redundant data and improve processing efficiency. Next, use the point cloud slicing projection and principal direction fitting algorithm (such as PCA principal axis analysis) to extract local structure information, and calculate indicators such as curvature distribution, boundary point density, and normal vector consistency in the point cloud. At the same time, through edge extraction and concavity-convexity analysis, generate a local geometric contour map as the basic geometric representation of the subsequent fusion features. All geometric feature parameters are bound to their tag structures according to the exhibit number and registered as point cloud geometric feature entries in the global exhibit information set. Through the refined analysis of the exhibit point cloud model, the geometric structure of the calligraphy and painting exhibits in the spatial dimension can be accurately identified, effectively supporting subsequent node attribute reasoning and exhibit status recognition, and enhancing the intelligent understanding ability of the exhibition system for the physical object form.
[0064] By retrieving the multi-spectral images bound to the exhibit tags, perform image registration and illumination normalization processing on each band image in turn. Subsequently, use algorithms such as GLCM (Gray-Level Co-Occurrence Matrix) and LBP (Local Binary Pattern) to extract information such as image texture directionality, roughness, and granularity, and combine color histogram analysis to analyze the main components of the surface color and color saturation. In addition, identify the moisture distribution and dye residue of the paper substrate in the near-infrared and ultraviolet bands to further judge the potential material and aging status. Finally, encode all image features as texture feature vectors, bind them to their tag structures according to the exhibit number, and register them as spectral texture feature entries in the global exhibit information set. By combining multi-band image analysis, the material characteristics, surface structure, and potential aging of calligraphy and paintings can be effectively identified, thereby realizing high-dimensional modeling of the exhibit content and status, and providing a perception basis for exhibit hierarchical recognition and personalized recommendation.
[0065] Using the exhibit number as an index, perform modal normalization and dimension unification processing on the extracted point cloud geometric feature vector and spectral texture feature vector. The fusion algorithm adopts a feature stitching and weighted fusion strategy, in which a priority weight factor (such as the status degradation risk score) is introduced for features related to the preservation status. At the same time, use the three-dimensional center coordinates of the exhibit as the spatial anchor point, and combine with the fusion vector to generate an associated feature vector structure: F = (x, y, z, S i , P), where x, y, z are the exhibit point cloud coordinates, S i is the spectral texture feature vector, and P is the preservation status score.
[0066] According to the associated feature vector, the structured deconstruction operation is performed according to the preset exhibit node template. First, the basic attribute layer (such as exhibit name, number, material, size and coordinates) is analyzed, and then the derived attributes (such as surface aging level, light sensitivity, mobile adaptability) are mined. Finally, the associated edges with other exhibits are constructed by matching content-based semantic labels (such as style category, historical period, author, etc.). All node information is injected into the exhibit graph database in the form of graph structure data organization, forming a structured exhibition network with exhibits as nodes and attributes and relationships as edges, that is, a coordinate-label mapping network. Through hierarchical analysis, each exhibit is converted into a knowledge graph node with semantic understanding capabilities, realizing deep modeling of exhibition objects from physical to semantic, effectively supporting personalized navigation, intelligent recommendation and automatic association and interaction of exhibits.
[0067] Further, step S340 includes: Step S341: traverse the gaze focus heat map to extract eye movement feature data, and calculate the area dwell time according to the eye movement feature data.
[0068] Step S342: Screening the coordinates of the exhibits according to the multiple matching confidences to generate a candidate exhibit set.
[0069] Step S343: when the area residence time is greater than or equal to the preset time, the exhibit identification module is activated to traverse the candidate exhibit set for number identification to determine the target exhibit number.
[0070] Specifically, the gaze focus heat map is traversed at the pixel level to locate the high-heat area (i.e., the area with color intensity greater than the set threshold), and the corresponding eye movement event sequence is extracted. The user's residence time in each hot zone is calculated by performing time-domain aggregation on the start and end times of eye movement events (such as gaze point sequences) in these hot zones. In order to eliminate the interference of rapid scanning, a minimum gaze threshold (such as 80ms) is set to filter out non-attention eye movement events to ensure the authenticity and validity of the residence time. All calculation results are bound to the user's heat map index structure to provide a timeliness criterion for subsequent exhibit matching. Through high-precision eye movement residence calculation, accurate identification of the user's real focus area is achieved, ensuring that the triggering conditions for activating the recognition module have sufficient human-computer interaction basis, and improving the reliability of target exhibit recognition and the fit with user intentions.
[0071] The coordinates of the exhibits are screened according to the multiple matching confidences calculated in step S330, and the exhibits with matching confidences higher than a set threshold are included in the candidate exhibit set, while limiting their spatial range to not exceed the viewing cone radius (eg, 2.5 meters) of the current user gaze area.
[0072] Combine the location of the heat map area with the physical coordinates of the candidate exhibits to further screen the exhibits in the overlapping area of the candidate set. After confirming that the user's residence time in a certain heat area is greater than or equal to the preset time, immediately activate the exhibit recognition module. This module traverses the current candidate exhibit set, focuses the line of sight on the central coordinates of the high-heat area in the heat map as the main judgment window, and calculates the spatial overlap rate Rh between it and each exhibit node. Subsequently, a joint score is calculated for the confidence Cf and the heat area overlap rate Rh of each candidate exhibit (e.g., Score = Cf × Rh), and the exhibit node with the highest score is selected as the current target exhibit, and the number of the exhibit is identified and output. If the scores of multiple exhibits are similar, the fuzzy inference mechanism can be entered to further optimize the number determination based on image similarity and the user's historical preferences. To improve the accuracy, a minimum confidence threshold is set (e.g., Cf > 0.65), and those below the threshold are marked as "uncertain" and wait for subsequent data accumulation to supplement the judgment.
[0073] Further, step S400 includes: Step S410: Dynamically synthesize multi-modal sensing data according to the target exhibit number to obtain a multi-modal synthesis data set.
[0074] Step S420: Perform linkage display of calligraphy and painting for the target user according to the multi-modal synthesis data set to generate display information, and the display information includes audio display information, brushstroke trajectory display information, and three-dimensional holographic image display information.
[0075] Step S430: Start the exhibit interaction module to real-time sense the multi-dimensional interaction data set of the target user, and the multi-dimensional interaction data set includes voice interaction data, action interaction data, and eye movement interaction data.
[0076] Step S440: Perform interaction based on the voice interaction data combined with the audio display information to generate a first interaction perception parameter.
[0077] Step S450: Perform interaction based on the action interaction data combined with the brushstroke trajectory display information to generate a second interaction perception parameter.
[0078] Step S460: Perform interaction based on the eye movement interaction data combined with the three-dimensional holographic image display information to generate a third interaction perception parameter.
[0079] Step S470: Perform a guided tour analysis of the calligraphy and painting exhibition for the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, and generate a personalized calligraphy and painting exhibition guided tour path.
[0080] Specifically, after receiving the target exhibit number, find the data corresponding to the exhibit in the existing multi-modal sensing data. Adopt a time series synchronization mechanism and a coordinate normalization method to fuse the data of each modality, and use tensor joint modeling or cross-attention mechanism to enhance the complementarity between modalities. Finally, generate a multi-modal synthesis dataset in a unified format for subsequent display calls.
[0081] According to the multi-modal synthesis dataset, combined with the pre-stored explanatory audio and stroke restoration data related to the exhibit, call the linkage display interface to generate three types of display information: audio display information, stroke trajectory display information, and three-dimensional holographic image display information. The explanatory audio of the exhibit (audio display information) can be output through the audio device; the pen tip landing points, strength, and direction trajectories of the work (stroke trajectory display information) can be replayed on the screen; the physical shape of the exhibit (three-dimensional holographic image display information) can be reconstructed with the help of a three-dimensional holographic projection system. The display information is synchronized and adapted to the user's spatial orientation to ensure the coordinated integration and consistent rhythm of various types of information.
[0082] Activate the exhibit interaction module to perform full-cycle perception on the target user and collect a multi-dimensional interaction dataset including voice interaction data, action interaction data, and eye movement interaction data. The module receives instruction statements such as explanation requests and questions (voice interaction data) through the voice channel (such as a microphone); recognizes the user's gesture actions and click behaviors (action interaction data) through a camera or motion capture device; records the gaze points and saccade frequencies (eye movement interaction data) through an eye tracker. After being paired with time stamps, various types of data are stored in the multi-dimensional interaction dataset, providing input for subsequent personalized recognition and tour path adjustment.
[0083] Then, perform information fusion and response matching on the three types of interaction data respectively to generate corresponding behavior response parameters, namely the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter. These three parameters correspond to the three channels of voice, action, and eye movement respectively. For the voice channel, compare the user's voice interaction data with the audio explanation content through a semantic understanding model to generate the first interaction perception parameter. For the action channel, map the user's gesture trajectory to the stroke animation path and calculate the spatial trajectory coincidence rate and action rhythm matching degree to generate the second interaction perception parameter. For the eye movement channel, calculate the interest weight according to the coincidence time between the gaze point and the holographic projection area and output the third interaction perception parameter. The interaction perception parameters of the three channels comprehensively depict the active attention dimensions of the user to the exhibit, providing clear weight support for the next tour optimization and making the basis for personalized recommendation more objective and multi-dimensional.
[0084] Perform weighted fusion on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to construct a user's instant interest vector, and perform similarity matching with the exhibit nodes. Combining the exhibition hall space topology and the exhibit position distribution, generate the shortest and highly relevant personalized guided tour path through a graph traversal algorithm. The path planning result is automatically updated according to the multi-dimensional interaction weights. For example, users who prefer voice-type content are preferentially recommended areas with detailed explanations, and users who prefer brushstroke imitation are preferentially pushed to interactive calligraphy areas. By integrating the perception parameters and the coordinate-label mapping network, the dynamic coupling between the guided tour content and the user's interests is realized, transforming the exhibition route from static planning to active guidance, and significantly improving the personalized visit experience and the depth of exhibition participation.
[0085] Further, step S470 includes: Step S471: Perform dynamic interest analysis on the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, and construct a user interest vector.
[0086] Step S472: Optimize the path according to the user interest vector in combination with the exhibit coordinate system to construct a moving path.
[0087] Step S473: Map the moving path to the global coordinate system of the exhibition hall for dynamic guidance marking to generate the personalized calligraphy and painting exhibition guided tour path.
[0088] Specifically, the user interest vector is a multi-dimensional vector representing the user's current interest state, including the interest intensity of the user in different exhibits, exhibition areas, exhibit categories, etc. By performing weighted processing on the first, second, and third interaction perception parameters, a user interest vector is generated. The weight of each type of perception parameter is dynamically adjusted according to the user's interaction frequency, interaction duration, interaction depth, etc. For example, if the user repeatedly asks about the background information of a certain exhibit, increase the weight of the voice interaction parameter of this exhibit; if the user stays in front of a certain exhibit for a long time, correspondingly increase the weight of this exhibit in the interest vector. In this way, the real-time interest changes of the user can be accurately captured.
[0089] Use path planning algorithms (such as the A-STAR algorithm, Dijkstra algorithm, etc.) to prioritize the exhibits according to the user interest vector. Then, in combination with the exhibit coordinate system, rearrange the exhibits according to the shortest path principle to generate the user's visit path. During the path planning process, comprehensively consider factors such as the distance between exhibits, the user's interest changes, and the spatio-temporal relationship of exhibits to ensure that the user can visit the exhibition efficiently and orderly. During the user's visit to the exhibition, this path will be adjusted and optimized in real time to adapt to the user's interest changes and enhance the interactivity and sense of participation.
[0090] According to the optimized user movement path, the exhibition hall global coordinate system is used to map the path to the exhibition space and generate dynamic guide marks. The dynamic guide marks are presented in real time through display screens, indicator lights or ground projections in the exhibition hall to guide users to move to the next target exhibit. = Voice, image or map prompts can also be pushed through mobile devices to help users quickly reach the next exhibit. By continuously tracking the user's spatial position, the path is adjusted in real time to ensure that the user is always on the correct exhibit guidance path. By combining the user's optimized path with the exhibition hall's global coordinate system and providing real-time feedback through dynamic guide marks, the accuracy and convenience of personalized tours are greatly improved, and the overall visiting experience is optimized.
[0091] In summary, the interactive method for calligraphy and painting exhibition provided by the embodiment of the present invention has the following beneficial effects: Through sensing devices, the behavior of the audience at the exhibition site is collected in real time to build a user exhibition behavior data set, providing basic data support for subsequent identification and personalized recommendations. By traversing the exhibit area for multi-dimensional image calibration, a coordinate-label mapping network between the exhibit space coordinates and semantic information is constructed to create conditions for accurate matching between user behavior and exhibits. By traversing the coordinate-label mapping network set according to the user exhibition behavior data set to activate the exhibit recognition module, the target exhibit number that the user is currently paying attention to is automatically identified, and the association conversion from behavior data to exhibit entities is realized. According to the target exhibit number, the target user is linked to display calligraphy and painting, generate display information, start the exhibit interaction module to deeply interact with the target user in combination with the display information, generate a personalized calligraphy and painting exhibition guide path for the target user, and enhance the immersive exhibition experience.
[0092] In general, the embodiments of the present invention realize real-time perception of user interests and personalized guided tour services in calligraphy and painting exhibitions by constructing a user behavior data set, a coordinate-label mapping network for image calibration, and an exhibit recognition and personalized linkage display mechanism based on behavior data, thereby significantly improving the interactivity, intelligence level and user immersive experience of the exhibition.
[0093] Embodiment 2, as Figure 3 As shown, based on the same inventive concept as the above-mentioned embodiment 1, the embodiment of the present invention provides an interactive system applied to calligraphy and painting exhibitions, the system comprising: The behavior data collection module 10 is used to collect data on target users in real time based on the exhibition area, and to synchronously obtain user behavior data and exhibition environment data through a multi-modal sensor array to construct a user exhibition behavior data set.
[0094] The image calibration module 20 is used to traverse the exhibition area for multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data of the exhibit surface, three-dimensional point cloud data, and texture feature data, and a coordinate-label mapping network is established through spatio-temporal association.
[0095] The target recognition module 30 is used to traverse the coordinate-label mapping network set according to the user's exhibition behavior data set for exhibit node matching, activate the exhibit recognition module, and identify the target exhibit number.
[0096] The interactive display module 40 is used to perform linkage display of calligraphy and painting for the target user according to the target exhibit number, generate display information, start the exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, generate a personalized calligraphy and painting exhibition tour path for the target user, and adjust the tour path in real time to adapt to the user's dynamic behavior.
[0097] Furthermore, the behavior data acquisition module 10 of the embodiment of the present invention is further used to execute the following steps: Deploy a multi-modal sensing array on the perimeter of the calligraphy and painting exhibits in the exhibition area, and perform real-time capture of the target user through the multi-modal sensing array to obtain a user sensing data set; establish a mapping relationship between the exhibit coordinate system and the global coordinate system of the exhibition hall, and convert the user sensing data into exhibition behavior metadata; perform spatio-temporal alignment processing on the exhibition behavior metadata according to the mapping relationship to generate a three-dimensional behavior vector; perform semantic parsing based on the three-dimensional behavior vector to generate a semantic association graph, perform user behavior analysis according to the semantic association graph to construct multi-source behavior event data, and add the multi-source behavior event data to the user's exhibition behavior data set.
[0098] Furthermore, the image calibration module 20 of the embodiment of the present invention is further used to execute the following steps: Traverse the exhibition area through the multi-modal sensing array to synchronously collect exhibits to obtain an exhibit sensing data set, where the exhibit sensing data set includes multi-spectral image data of the exhibit surface and three-dimensional point cloud data; map the exhibit sensing data set to the global coordinate system of the exhibition hall to generate an exhibit space topology map; perform associated mapping of the three-dimensional point cloud data and the multi-spectral image data of the exhibit surface according to the exhibit space topology map to generate an associated label set; calibrate the calligraphy and painting exhibits based on the associated label set to construct the coordinate-label mapping network.
[0099] Furthermore, the target recognition module 30 of the embodiment of the present invention is further used to execute the following steps: Perform motion trajectory analysis based on the user exhibition behavior dataset, extract spatio-temporal behavior feature vectors, perform gaze tracking based on the user exhibition behavior dataset, and draw a gaze focus heat map; extract multiple exhibit node information based on the associated tag set; traverse the coordinate-tag mapping network to perform matching calculations on the spatio-temporal behavior feature vectors and the exhibit node information, and obtain multiple matching confidence levels; make a determination according to the multiple matching confidence levels in combination with the gaze focus heat map, and activate the exhibit recognition module according to the determination result to obtain the target exhibit number.
[0100] Furthermore, the target recognition module 30 in the embodiment of the present invention is further configured to perform the following steps: Perform spatial displacement analysis based on the user exhibition behavior dataset to obtain the behavior trajectory sequence of the target user; perform trajectory classification according to the behavior trajectory sequence to obtain multiple movement mode categories, and perform behavior mining according to the multiple movement mode categories to obtain spatio-temporal behavior feature vectors.
[0101] Furthermore, the target recognition module 30 in the embodiment of the present invention is further configured to perform the following steps: Perform feature analysis based on the associated tag set in combination with the three-dimensional point cloud data to obtain point cloud geometric features; perform feature analysis based on the associated tag set in combination with the multi-spectral image data to obtain spectral texture features; perform multi-modal feature fusion on the point cloud geometric features and the spectral texture features according to the associated tag set to construct an associated feature vector; perform hierarchical node parsing according to the associated feature vector to determine the multiple exhibit node information.
[0102] Furthermore, the target recognition module 30 in the embodiment of the present invention is further configured to perform the following steps: Traverse the gaze focus heat map to extract eye movement feature data, calculate the regional residence duration according to the eye movement feature data; perform exhibit coordinate screening according to the multiple matching confidence levels to generate a candidate exhibit set; when the regional residence duration is greater than or equal to the preset duration, activate the exhibit recognition module to traverse the candidate exhibit set for number recognition to determine the target exhibit number.
[0103] Furthermore, the interactive display module 40 in the embodiment of the present invention is further configured to perform the following steps: Call multi-modal sensing data according to the target exhibit number for dynamic synthesis to obtain a multi-modal synthesis data set; perform linked display of calligraphy and painting for the target user according to the multi-modal synthesis data set to generate display information, where the display information includes audio display information, brush stroke trajectory display information, and three-dimensional holographic image display information; start the exhibit interaction module to continuously sense the multi-dimensional interaction data set of the target user, where the multi-dimensional interaction data set includes voice interaction data, motion interaction data, and eye movement interaction data; perform interaction based on the voice interaction data combined with the audio display information to generate a first interaction perception parameter; perform interaction based on the motion interaction data combined with the brush stroke trajectory display information to generate a second interaction perception parameter; perform interaction based on the eye movement interaction data combined with the three-dimensional holographic image display information to generate a third interaction perception parameter; perform tour guide analysis of the calligraphy and painting exhibition for the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to generate a personalized calligraphy and painting exhibition tour guide path.
[0104] Further, the interaction display module 40 in the embodiment of the present invention is further configured to perform the following steps: Perform dynamic interest analysis on the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to construct a user interest vector; optimize the path according to the user interest vector combined with the exhibit coordinate system to construct a movement path; map the movement path to the global coordinate system of the exhibition hall for dynamic guidance marking to generate the personalized calligraphy and painting exhibition tour guide path.
[0105] Through the foregoing detailed description of an interaction method applied to a calligraphy and painting exhibition in this specification, those skilled in the art can clearly know an interaction system applied to a calligraphy and painting exhibition in this embodiment. For the system disclosed in Embodiment 2, since it corresponds to the method disclosed in Embodiment 1 and has corresponding functional modules and beneficial effects, the relevant parts can be referred to the description in the method part.
[0106] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An interactive method applied to calligraphy and painting exhibitions, characterized in that, The method includes: Based on the exhibition area, real-time collection of target users is carried out, and user behavior data and exhibition environment data are synchronously obtained through a multimodal sensing array to construct a user exhibition behavior dataset; Traverse the exhibition area for multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data of the exhibition surface, three-dimensional point cloud data, and texture feature data, and a coordinate-label mapping network is established through spatio-temporal association; According to the user exhibition behavior dataset, traverse the coordinate-label mapping network set to match the exhibition nodes, activate the exhibition recognition module, and identify the target exhibition number; Perform linkage display of calligraphy and painting for the target user according to the target exhibition number, generate display information, start the exhibition interaction module, combine the display information to perform in-depth interaction with the target user, generate a personalized calligraphy and painting exhibition tour path for the target user, and adjust the tour path in real time to adapt to the user's dynamic behavior.
2. The interactive method for calligraphy and painting exhibition according to claim 1, characterized in that, Based on the exhibition area, real-time collection of target users is carried out, and user behavior data and exhibition environment data are synchronously obtained through a multimodal sensing array to construct a user exhibition behavior dataset. The method includes: Deploy a multimodal sensing array on the perimeter of the calligraphy and painting exhibitions in the exhibition area, and use the multimodal sensing array to capture target users in real time to obtain a user sensing dataset; Establish a mapping relationship between the exhibition coordinate system and the global coordinate system of the exhibition hall, and convert the user sensing data into exhibition behavior metadata; Perform spatio-temporal alignment processing on the exhibition behavior metadata according to the mapping relationship to generate a three-dimensional behavior vector; Perform semantic analysis based on the three-dimensional behavior vector to generate a semantic association graph, perform user behavior analysis according to the semantic association graph to construct multi-source behavior event data, and add the multi-source behavior event data to the user exhibition behavior dataset.
3. The interactive method for a painting and calligraphy exhibition according to claim 2, wherein, Traverse the exhibition area for multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data of the exhibition surface, three-dimensional point cloud data, and texture feature data, and a coordinate-label mapping network is established through spatio-temporal association. The method includes: Use the multimodal sensing array to traverse the exhibition area to synchronously collect exhibitions to obtain an exhibition sensing dataset, and the exhibition sensing dataset includes multi-spectral image data of the exhibition surface and three-dimensional point cloud data; Map the exhibition sensing dataset to the global coordinate system of the exhibition hall to generate an exhibition space topology map; Perform associated mapping of the three-dimensional point cloud data and the multi-spectral image data of the exhibition surface according to the exhibition space topology map to generate an associated label set; Calibrate the calligraphy and painting exhibitions based on the associated label set to construct the coordinate-label mapping network.
4. The interactive method for a painting and calligraphy exhibition according to claim 3, wherein According to the user exhibition behavior dataset, traverse the coordinate-label mapping network set to activate the exhibition recognition module and identify the target exhibition number. The method includes: Perform motion trajectory analysis based on the user exhibition behavior dataset, extract spatio-temporal behavior feature vectors, perform gaze tracking based on the user exhibition behavior dataset, and draw a gaze focus heat map; Extract multiple exhibition node information based on the associated label set; Traverse the coordinate-label mapping network to perform matching calculations on the spatio-temporal behavior feature vector and the exhibit node information, and obtain multiple matching confidence levels. Make a determination according to the multiple matching confidence levels in combination with the gaze focus heat map, and activate the exhibit recognition module according to the determination result to obtain the target exhibit number.
5. The interactive method for a calligraphy and painting exhibition according to claim 4, wherein Perform motion trajectory analysis based on the user exhibition behavior data set, and extract spatio-temporal behavior feature vectors. The method includes: Perform spatial displacement analysis based on the user exhibition behavior data set to obtain the behavior trajectory sequence of the target user. Perform trajectory classification according to the behavior trajectory sequence to obtain multiple moving mode categories, and perform behavior mining according to the multiple moving mode categories to obtain spatio-temporal behavior feature vectors.
6. The interactive method applied to a calligraphy and painting exhibition according to claim 4, characterized in that Extract multiple exhibit node information based on the associated tag set. The method includes: Perform feature analysis based on the associated tag set in combination with the three-dimensional point cloud data to obtain point cloud geometric features. Perform feature analysis based on the associated tag set in combination with the multi-spectral image data to obtain spectral texture features. Perform multi-modal feature fusion on the point cloud geometric features and the spectral texture features according to the associated tag set to construct an associated feature vector. Perform hierarchical node parsing according to the associated feature vector to determine the multiple exhibit node information.
7. The interactive method for a calligraphy and painting exhibition according to claim 4, characterized in that, Make a determination according to the multiple matching confidence levels in combination with the gaze focus heat map, and activate the exhibit recognition module according to the determination result to obtain the target exhibit number. The method includes: Traverse the gaze focus heat map to extract eye movement feature data, and calculate the regional residence duration according to the eye movement feature data. Perform exhibit coordinate screening according to the multiple matching confidence levels to generate a candidate exhibit set. When the regional residence duration is greater than or equal to the preset duration, activate the exhibit recognition module to traverse the candidate exhibit set for number recognition to determine the target exhibit number.
8. The interactive method for a calligraphy and painting exhibition according to claim 2, characterized in that, Perform linked display of calligraphy and painting for the target user according to the target exhibit number to generate display information, start the exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, generate a personalized calligraphy and painting exhibition tour path for the target user, and adjust the tour path in real time to adapt to the user's dynamic behavior. The method includes: Call multi-modal sensing data for dynamic synthesis according to the target exhibit number to obtain a multi-modal synthesis data set. Perform linked display of calligraphy and painting for the target user according to the multi-modal synthesis data set to generate display information, and the display information includes audio display information, stroke trajectory display information, and three-dimensional holographic image display information. Start the exhibit interaction module to real-time sense the multi-dimensional interaction data set of the target user, and the multi-dimensional interaction data set includes voice interaction data, action interaction data, and eye movement interaction data. Perform interaction based on the voice interaction data in combination with the audio display information to generate a first interaction perception parameter. Perform interaction based on the action interaction data in combination with the stroke trajectory display information to generate a second interaction perception parameter. Perform interaction based on the eye movement interaction data in combination with the three-dimensional holographic image display information to generate a third interaction perception parameter. Based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, conduct a guided tour analysis of the target user for the calligraphy and painting exhibition, and generate a personalized calligraphy and painting exhibition guided tour path.
9. The interactive method applied to a calligraphy and painting exhibition according to claim 8, wherein Based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, conduct a guided tour analysis of the target user for the calligraphy and painting exhibition, and generate a personalized calligraphy and painting exhibition guided tour path. The method includes: Conduct a dynamic interest analysis of the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, and construct a user interest vector; Optimize the path according to the user interest vector in combination with the exhibit coordinate system, and construct a movement path; Map the movement path to the global coordinate system of the exhibition hall for dynamic guidance marking, and generate the personalized calligraphy and painting exhibition guided tour path.
10. An interactive system applied to calligraphy and painting exhibitions, characterized in that, The system is used to execute an interaction method for a calligraphy and painting exhibition according to any one of claims 1-9, including: A behavior data collection module, configured to perform real-time collection on the target user based on the exhibit area, synchronously obtain user behavior data and exhibit environment data through a multi-modal sensing array, and construct a user exhibition viewing behavior data set; An image calibration module, configured to traverse the exhibit area for multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data, and texture feature data on the exhibit surface, and establish a coordinate-label mapping network through spatio-temporal association; A target recognition module, configured to traverse the coordinate-label mapping network set according to the user exhibition viewing behavior data set for exhibit node matching, activate the exhibit recognition module, and identify the target exhibit number; An interactive display module, configured to perform a linked display of calligraphy and painting for the target user according to the target exhibit number, generate display information, start the exhibit interaction module to conduct in-depth interaction with the target user in combination with the display information, generate a personalized calligraphy and painting exhibition guided tour path for the target user, and adjust the guided tour path in real time to adapt to the dynamic behavior of the user.
Citation Information
Patent Citations
Complex road target detection method based on multi-modal fusion aerial view
CN117058646A
Smart museum user management system
CN117456588A
3D exhibition visiting flow prediction method and system based on user behavior analysis
CN118691441A
Intelligent navigation system based on AIGC
CN118864170A
Information visual management system and method based on digital twinning
CN119271899A
Cited By
Intelligent environmental protection propaganda interaction system and method based on VI identification
CN120510007A
Multimedia exhibition hall AI intelligent interaction control system based on Internet of Things
CN120848729A
Guide path display method and device of virtual scene, computer equipment and medium
CN121165945A
Exhibition hall interactive display method and system supporting artificial intelligence recognition
CN121365267A
An exhibition hall interactive display method and system supporting artificial intelligence recognition
CN121365267B