An interactive method and system applied to calligraphy and painting exhibitions

Through the multimodal sensing array and coordinate-label mapping network, the exhibits that users are concerned about are identified and a personalized guided path is generated, which solves the problem of lack of real-time perception and accurate recognition in the existing technology, and improves the interactivity and immersion of calligraphy and painting exhibitions.

CN120279221BActive Publication Date: 2025-08-26THE ARCHITECTURAL DESIGN & RES INST OF ZHEJIANG UNIV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510766307.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-26
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing interactive methods of calligraphy and painting exhibitions lack real-time perception and accurate identification of users' exhibition behavior, resulting in the inability to realize personalized exhibit linkage and dynamic guide, affecting the audience's immersion and interactive experience.

Method used

Through a multimodal sensing array, a user behavior data and exhibit environment data are collected in real time, a user exhibition behavior data set is built, a coordinate-label mapping network is established, the target exhibit number is identified, and a personalized calligraphy and painting exhibition guide path is generated to achieve in-depth interaction.

Benefits of technology

Real-time perception of user interests and personalized guided tour services are realized, improving the interactiveness of the exhibition and user immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279221B_ABST
    Figure CN120279221B_ABST
Patent Text Reader

Abstract

The present invention provides an interactive method and system for calligraphy and painting exhibitions, which relates to the field of data interaction technology. The method and system construct a user viewing behavior dataset based on real-time collection of exhibit areas; traverse the exhibit areas to perform multi-dimensional image calibration and establish a coordinate-label mapping network; activate the exhibit identification module based on the user viewing behavior dataset by traversing the coordinate-label mapping network set, identify the target exhibit number to perform a linked display of calligraphy and painting for the target user, generate display information, start the exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, and generate a personalized calligraphy and painting exhibition guide path. The present invention solves the technical problem that the existing technology cannot realize personalized exhibit linkage and dynamic navigation due to the lack of real-time perception and accurate identification of user viewing behavior, and achieves the technical effect of generating a personalized guide path based on user interests, thereby enhancing the immersiveness and interactive experience of viewing the exhibition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data interaction technology, and in particular to an interaction method and system applied to calligraphy and painting exhibitions. Background Art

[0002] In the exhibition process of calligraphy and painting art, interactive experience has gradually become an important means to enhance audience participation and artistic appeal. With the development of digital technology, the traditional calligraphy and painting exhibition model, which mainly relies on static display, is gradually evolving towards informationization and intelligentization. Existing interactive methods for calligraphy and painting exhibitions are mostly based on QR code scanning, electronic guides or touch screens to assist the audience in understanding the exhibits. However, this type of interaction is mostly passively triggered and lacks real-time perception and in-depth understanding of user viewing behavior, making it difficult to carry out targeted exhibition linkage display and information push. At the same time, the organization of exhibition content is relatively fixed, and the audience is prone to information overload or missing exhibits of interest, which limits the in-depth interaction between the audience and the exhibits, affecting the interactive viewing experience and exhibition efficiency. Summary of the Invention

[0003] The present invention provides an interactive method and system for calligraphy and painting exhibitions, which solves the technical problem that the existing technology cannot realize personalized exhibit linkage and dynamic navigation due to the lack of real-time perception and accurate recognition of user viewing behavior. It achieves the technical effect of automatically identifying target exhibits based on user interests and generating personalized navigation paths, thereby enhancing the immersiveness of exhibition viewing and the interactive experience.

[0004] In view of the above problems, on the one hand, the present invention provides an interactive method applied to calligraphy and painting exhibitions, the method comprising: real-time collection of target users based on an exhibit area, synchronously acquiring user behavior data and exhibit environment data through a multimodal sensor array, and constructing a user exhibition viewing behavior data set; traversing the exhibit area to perform multi-dimensional image calibration, the multi-dimensional image calibration comprising multi-spectral image data, three-dimensional point cloud data, and texture feature data of the exhibit surface, and establishing a coordinate-label mapping network through spatiotemporal association; traversing the coordinate-label mapping network set to perform exhibit node matching according to the user exhibition viewing behavior data set, activating an exhibit recognition module, and identifying a target exhibit number; performing a linkage display of calligraphy and painting for the target user according to the target exhibit number, generating display information, activating an exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, generating a personalized calligraphy and painting exhibition guide path for the target user, and adjusting the guide path in real time to adapt to the user's dynamic behavior.

[0005] Preferably, real-time data collection of target users is performed based on the exhibit area, and user behavior data and exhibit environment data are synchronously acquired through a multimodal sensor array to construct a user exhibition behavior data set. The method includes: deploying a multimodal sensor array at the perimeter of calligraphy and painting exhibits in the exhibit area, and capturing target users in real time through the multimodal sensor array to obtain a user sensor data set; establishing a mapping relationship between the exhibit coordinate system and the exhibition hall global coordinate system, and converting the user sensor data into exhibition behavior metadata; performing spatiotemporal alignment processing on the exhibition behavior metadata according to the mapping relationship to generate a three-dimensional behavior vector; performing semantic analysis based on the three-dimensional behavior vector to generate a semantic association graph, performing user behavior analysis based on the semantic association graph to construct multi-source behavior event data, and adding the multi-source behavior event data to the user exhibition behavior data set.

[0006] Preferably, multi-dimensional image calibration is performed by traversing the exhibition area, and the multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data and texture feature data of the exhibit surface, and a coordinate-label mapping network is established through spatiotemporal association. The method includes: traversing the exhibition area by the multimodal sensor array to synchronously collect exhibits to obtain an exhibit sensor data set, and the exhibit sensor data set includes multi-spectral image data and three-dimensional point cloud data of the exhibit surface; mapping the exhibit sensor data set to the global coordinate system of the exhibition hall to generate an exhibit space topology map; according to the exhibit space topology map, the three-dimensional point cloud data and the multi-spectral image data of the exhibit surface are associated and mapped to generate an associated label set; based on the associated label set, the calligraphy and painting exhibits are calibrated to construct the coordinate-label mapping network.

[0007] Preferably, the coordinate-label mapping network set is traversed according to the user exhibition viewing behavior data set to activate the exhibit recognition module and identify the target exhibit number. The method includes: performing motion trajectory analysis based on the user exhibition viewing behavior data set to extract the spatiotemporal behavior feature vector, performing line of sight tracking based on the user exhibition viewing behavior data set to draw a line of sight focus heat map; extracting multiple exhibit node information based on the associated label set; traversing the coordinate-label mapping network to perform matching calculations on the spatiotemporal behavior feature vector and the exhibit node information to obtain multiple matching confidence levels; making a judgment based on the multiple matching confidence levels combined with the line of sight focus heat map, activating the exhibit recognition module according to the judgment result, and obtaining the target exhibit number.

[0008] Preferably, motion trajectory analysis is performed based on the user exhibition viewing behavior dataset to extract the spatiotemporal behavior feature vector. The method includes: performing spatial displacement analysis based on the user exhibition viewing behavior dataset to obtain the target user's behavior trajectory sequence; performing trajectory classification according to the behavior trajectory sequence to obtain multiple mobility pattern categories; performing behavior mining based on the multiple mobility pattern categories to obtain the spatiotemporal behavior feature vector.

[0009] Preferably, the method for extracting information of multiple exhibit nodes based on the associated tag set includes: performing feature analysis based on the associated tag set in combination with the three-dimensional point cloud data to obtain point cloud geometric features; performing feature analysis based on the associated tag set in combination with the multispectral image data to obtain spectral texture features; performing multimodal feature fusion on the point cloud geometric features and the spectral texture features according to the associated tag set to construct an associated feature vector; and performing hierarchical node parsing according to the associated feature vector to determine the information of the multiple exhibit nodes.

[0010] Preferably, a determination is made according to the multiple matching confidence levels in combination with the gaze focus heat map, and the exhibit identification module is activated according to the determination result to obtain a target exhibit number. The method includes: traversing the gaze focus heat map to extract eye movement feature data, and calculating a region dwell time according to the eye movement feature data; screening exhibit coordinates according to the multiple matching confidence levels to generate a candidate exhibit set; and when the region dwell time is greater than or equal to a preset time, activating the exhibit identification module to traverse the candidate exhibit set for number identification to determine the target exhibit number.

[0011] Preferably, calligraphy and painting are displayed in a linked manner to the target user according to the target exhibit number, display information is generated, the exhibit interaction module is started to deeply interact with the target user in combination with the display information, a personalized calligraphy and painting exhibition guide path for the target user is generated, and the guide path is adjusted in real time to adapt to the user's dynamic behavior. The method includes: calling multimodal sensor data for dynamic synthesis according to the target exhibit number to obtain a multimodal synthesis data set; displaying calligraphy and painting to the target user in a linked manner according to the multimodal synthesis data set to generate display information, wherein the display information includes audio display information, brushstroke trajectory display information, and three-dimensional holographic image display information; starting the exhibit interaction module to perform in-depth interaction with the target user in combination with the display information, generating a personalized calligraphy and painting exhibition guide path for the target user, and adjusting the guide path in real time to adapt to the user's dynamic behavior. The module perceives the target user's multidimensional interaction data set in real time, where the multidimensional interaction data set includes voice interaction data, action interaction data, and eye movement interaction data; interacts based on the voice interaction data in combination with audio display information to generate a first interaction perception parameter; interacts based on the action interaction data in combination with brushstroke trajectory display information to generate a second interaction perception parameter; interacts based on the eye movement interaction data in combination with three-dimensional holographic image display information to generate a third interaction perception parameter; and conducts a guided analysis of the calligraphy and painting exhibition for the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to generate a personalized guided path for the calligraphy and painting exhibition.

[0012] Preferably, a guide analysis of the calligraphy and painting exhibition is performed on the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to generate a personalized calligraphy and painting exhibition guide path. The method includes: performing a dynamic interest analysis on the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to construct a user interest vector; optimizing the path according to the user interest vector in combination with the exhibit coordinate system to construct a moving path; mapping the moving path to the exhibition hall global coordinate system for dynamic guidance marking to generate the personalized calligraphy and painting exhibition guide path.

[0013] On the other hand, the present invention also provides an interactive system for calligraphy and painting exhibitions, the system comprising: a behavior data acquisition module for real-time collection of target users based on the exhibit area, synchronously acquiring user behavior data and exhibit environment data through a multimodal sensor array, and constructing a user exhibition behavior data set; an image calibration module for traversing the exhibit area for multi-dimensional image calibration, the multi-dimensional image calibration comprising multi-spectral image data, three-dimensional point cloud data and texture feature data of the exhibit surface, and establishing a coordinate-label mapping network through spatiotemporal association; a target recognition module for traversing the coordinate-label mapping network set according to the user exhibition behavior data set to match exhibit nodes, activate an exhibit recognition module, and identify a target exhibit number; an interactive display module for performing a linkage display of calligraphy and painting for the target user according to the target exhibit number, generating display information, and activating an exhibit interaction module for in-depth interaction with the target user in combination with the display information, generating a personalized calligraphy and painting exhibition guide path for the target user, and adjusting the guide path in real time to adapt to the user's dynamic behavior.

[0014] One or more technical solutions provided in the present invention have at least the following beneficial effects:

[0015] Through sensing devices, the audience's behavior at the exhibition site is collected in real time to build a user viewing behavior dataset, providing basic data support for subsequent identification and personalized recommendations. By traversing the exhibition area for multi-dimensional image calibration, a coordinate-label mapping network is constructed between the exhibit space coordinates and semantic information, creating conditions for accurate matching between user behavior and exhibits. By traversing the coordinate-label mapping network set based on the user viewing behavior dataset and activating the exhibit recognition module, the target exhibit number currently being focused on by the user is automatically identified, realizing the association conversion from behavior data to exhibit entity. According to the target exhibit number, calligraphy and painting are displayed to the target user in a linked manner, and display information is generated. The exhibit interaction module is activated to conduct in-depth interaction with the target user based on the display information, generating a personalized calligraphy and painting exhibition guide path for the target user, and enhancing the immersive viewing experience.

[0016] In summary, the present invention realizes real-time perception of user interests and personalized guided services in calligraphy and painting exhibitions by constructing a user behavior dataset, a coordinate-label mapping network for image calibration, and an exhibit identification and personalized linkage display mechanism based on behavior data, significantly improving the interactivity, intelligence level and user immersive experience of the exhibition.

[0017] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A flowchart of an interactive method applied to calligraphy and painting exhibitions provided by an embodiment of the present invention.

[0019] Figure 2 A flowchart illustrating a method for traversing a coordinate-label mapping network set based on a user viewing behavior dataset to activate an exhibit identification module and identify a target exhibit number in an interactive method for a calligraphy and painting exhibition provided by an embodiment of the present invention.

[0020] Figure 3 A schematic structural diagram of an interactive system for calligraphy and painting exhibitions provided by an embodiment of the present invention.

[0021] Description of the accompanying drawings: behavior data collection module 10, image calibration module 20, target recognition module 30, interactive display module 40. DETAILED DESCRIPTION

[0022] The embodiments of the present invention provide an interactive method and system for calligraphy and painting exhibitions, which solves the technical problem that the existing technology cannot realize personalized exhibit linkage and dynamic navigation due to the lack of real-time perception and accurate recognition of user viewing behavior. It achieves the technical effect of automatically identifying target exhibits based on user interests and generating personalized navigation paths, thereby enhancing the immersiveness and interactive experience of viewing the exhibition.

[0023] Example 1, as Figure 1 As shown, an embodiment of the present invention provides an interactive method applied to a calligraphy and painting exhibition, the method comprising:

[0024] Step S100: Real-time data collection of target users is performed based on the exhibition area, and user behavior data and exhibition environment data are synchronously acquired through a multimodal sensor array to construct a user exhibition behavior dataset.

[0025] Specifically, the exhibition area refers to the physical area within the exhibition space where the calligraphy and painting works are located. Within the exhibition area, a multimodal sensor array is deployed to simultaneously collect user behavior data and exhibition environment data to achieve accurate modeling of user viewing behavior. This multimodal sensor array integrates multiple sensor devices, including but not limited to: RGB cameras (capturing images and human posture), infrared depth cameras (measuring spatial position), eye trackers (capturing gaze point and gaze direction), motion capture devices (tracking limb movements), and voice sensors (such as microphones). The multimodal sensor array is used to record each visitor's interactive behavior with each exhibit, constructing an exhibition behavior dataset containing timing, position information, and eye movement characteristics, achieving accurate modeling of user viewing behavior, and providing a highly reliable input data foundation for subsequent exhibit identification and personalized guided tours.

[0026] Step S200: traverse the exhibit area to perform multi-dimensional image calibration, wherein the multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data and texture feature data of the exhibit surface, and establishes a coordinate-label mapping network through spatiotemporal association.

[0027] Specifically, multi-dimensional image calibration is performed by traversing the exhibit area, that is, using multiple sensing methods to scan, identify and model the exhibits one by one, and obtain perception data of the exhibits in different dimensions. The perception data includes but is not limited to multispectral image data, three-dimensional point cloud data and texture feature data of the exhibit surface. The multispectral image data of the exhibit surface is obtained by imaging the exhibit surface with a multispectral camera, and the image information of the exhibit in different bands such as visible light, near-infrared (NIR), and ultraviolet (UV) is obtained, reflecting the color, material, texture and other detailed features of the calligraphy and painting works; the three-dimensional point cloud data is the spatial depth information of the exhibit collected using structured light or lidar technology, which describes the geometric outline and spatial position of the exhibit and is used to construct its three-dimensional geometric model. The texture feature data is the surface texture features extracted from the above-mentioned multispectral image data, which is used to describe the grayscale changes, edge direction, corner point distribution, etc. of the local image fragment.

[0028] To achieve joint modeling of image data and spatial data, the above-mentioned multi-source data (multispectral image data of the exhibit surface, three-dimensional point cloud data and texture feature data) are fused through a spatiotemporal synchronization mechanism. In the unified coordinate system of the exhibition hall, a mapping relationship is established between the structural coordinate points of each exhibit (such as key points, edges, textures, etc. on the exhibit surface) and their corresponding semantic labels, forming a structured multi-label spatial semantic network, namely a coordinate-label mapping network, so that each position point in the space can be accurately mapped to a certain exhibit and its specific content, which can be used for subsequent user behavior positioning and exhibit identification.

[0029] Step S300: traversing the coordinate-label mapping network set according to the user exhibition viewing behavior dataset to perform exhibit node matching, activating an exhibit identification module, and identifying a target exhibit number.

[0030] Specifically, the exhibit identification module analyzes user behavior and identifies the exhibit number of the user's current interest. The user's viewing behavior data (gaze direction and location trajectory) from step S100 is input into the coordinate-label mapping network for traversal analysis. A confidence matching calculation is performed based on the gaze and dwell position with the spatial position of each exhibit in the coordinate-label mapping network. The exhibit identification module is then activated to determine the target exhibit number of the user's current interest. This precise matching calculation allows for real-time identification of the user's preferred exhibit, providing a foundation for subsequent targeted display and interaction.

[0031] Step S400: Perform a linked display of calligraphy and painting for the target user according to the target exhibit number, generate display information, start the exhibit interaction module to conduct in-depth interaction with the target user based on the display information, generate a personalized calligraphy and painting exhibition guide path for the target user, and adjust the guide path in real time to adapt to the user's dynamic behavior.

[0032] Specifically, the exhibit interaction module is a functional component responsible for collecting user feedback and conducting interactive control. Display information includes a range of content related to the calligraphy and painting exhibits for display to users, such as high-definition images of the artwork, audio biography of the artist, video explanations of painting techniques, and 3D holographic images. Based on the identified exhibit number, the exhibit's multimodal content (audio description, 3D model, brushstroke dynamics, etc.) is linked and displayed on interactive terminals (such as large screens or AR glasses), generating display information and enabling immersive, multimodal human-computer interaction. Simultaneously, the interaction module captures user voice questions, gestures, and gaze changes in real time to generate multidimensional interaction perception parameters. Based on these parameters, the user's interests are dynamically assessed and recommended routes are intelligently adjusted. Ultimately, a guided tour of the calligraphy and painting exhibition is generated that better suits the user's individual needs. A guidance system (such as projection or navigation signs) guides the user onward through the exhibition, significantly enhancing their sense of participation and satisfaction with the exhibition.

[0033] Furthermore, step S100 includes:

[0034] Step S110: deploying a multimodal sensor array around the calligraphy and painting exhibits in the exhibition area, and capturing the target user in real time through the multimodal sensor array to obtain a user sensor data set.

[0035] Step S120: establishing a mapping relationship between the exhibit coordinate system and the exhibition hall global coordinate system, and converting the user sensor data into exhibition viewing behavior metadata.

[0036] Step S130: performing spatiotemporal alignment processing on the exhibition viewing behavior metadata according to the mapping relationship to generate a three-dimensional behavior vector.

[0037] Step S140: Perform semantic analysis based on the three-dimensional behavior vector to generate a semantic association graph, perform user behavior analysis based on the semantic association graph to construct multi-source behavior event data, and add the multi-source behavior event data to the user exhibition viewing behavior dataset.

[0038] Specifically, the multimodal sensor array deployment must ensure coverage of the areas directly in front of, to the left and right sides, and above the exhibit to avoid blind spots. A time synchronization protocol is also used to enable linked sampling of multimodal devices. When a target user enters the sensing range, the multimodal sensor array initiates synchronous data collection, recording behavioral data including their location, orientation, line of sight, gaze duration, and interactive behavior. This data is then packaged into a user sensor dataset. This user sensor dataset is initially timestamped and associated with the user's unique identity (user ID), providing raw input for subsequent data processing.

[0039] The exhibit coordinate system is a local coordinate system established with each calligraphy and painting exhibit as the origin, and the exhibition hall global coordinate system is a unified three-dimensional reference system constructed based on the entire exhibition hall. A panoramic positioning camera, lidar or panoramic depth camera equipment is used to perform a three-dimensional scan of the exhibition space and establish a three-dimensional exhibition hall point cloud map as the basis of the global coordinate system. The origin is usually set at the center of the exhibition hall entrance, with the Z axis pointing upwards, and the X and Y axes extending along the main aisle of the exhibition hall. At the same time, a number of positioning reference points (such as QR code targets and laser reflective balls) are arranged in the exhibition hall, and their spatial coordinates in the global coordinate system are recorded for subsequent spatial alignment and posture correction of data. Subsequently, the position and posture of each exhibit in the exhibition hall are laser calibrated, and the local coordinate systems of each exhibit are aligned to the unified exhibition hall global coordinate system using a homogeneous transformation matrix to achieve a unified reference for spatial position. The user's real-time position, line of sight and other sensor data are projected or transformed into the global coordinate system and converted into exhibition behavior metadata in a unified format, including: user ID, timestamp, position coordinates (X u , Y u , Z u ), sensor type, original exhibition feature value, etc., providing a unified coordinate basis for the alignment and fusion of multi-time period and multi-sensor channel data.

[0040] Based on the global coordinate system, multimodal data from the same user within adjacent time windows is sorted and aggregated by timestamp. Data frames are aligned using a unified time window ΔT (e.g., 100ms). Each viewing behavior metadata is represented as a three-dimensional vector structure: behavior timestamp, location coordinates in the exhibition hall's global coordinate system, and interaction events. This allows for accurate description and analysis of user behavior data within a unified timeline and spatial framework. An interaction event refers to the specific type of exhibition interaction performed by a user at that moment and location, such as gazing at an exhibit, lingering in front of an exhibit, or interacting with an exhibit through touch.

[0041] Semantic labels are inferred for multiple three-dimensional behavior vectors of the target user, identifying the temporal, spatial, and interaction pattern characteristics of user behavior. Using a predefined behavioral rule library and semantic label inference mechanism, typical behaviors are classified as events. For example, if the position change threshold is less than a preset threshold, it is considered a "dwell," and if the gaze point is concentrated, it is considered a "gaze." Various semantic label combinations form a semantic association graph in the form of nodes. The edges between different nodes represent the temporal sequence of user behaviors or the semantic transfer logic. Node information in the semantic graph is further extracted and merged to form structured multi-source behavioral event data records. For example, a combination of gaze and dwell time greater than 5 seconds is classified as a "dwell" event; gaze plus gesture or voice constitutes an "active interaction" event. Ultimately, this event data is added to the user's exhibition behavior dataset, providing a rich behavioral foundation for subsequent personalized content delivery and behavioral pattern analysis.

[0042] Furthermore, step S200 includes:

[0043] Step S210: synchronously collecting information about the exhibits by traversing the exhibit area through the multimodal sensor array to obtain an exhibit sensor data set, wherein the exhibit sensor data set includes multispectral image data and three-dimensional point cloud data of the exhibit surface.

[0044] Step S220: Mapping the exhibit sensor data set to the exhibition hall global coordinate system to generate an exhibit space topology map.

[0045] Step S230: performing association mapping on the three-dimensional point cloud data and the multispectral image data of the exhibit surface according to the exhibit space topology map to generate an association tag set.

[0046] Step S240: calibrating the calligraphy and painting exhibits based on the associated label set, and constructing the coordinate-label mapping network.

[0047] Specifically, in addition to the aforementioned sensors for collecting user behavior data and exhibit environment data, the multimodal sensor array also includes multispectral imaging devices (such as VNIR and UV cameras) and three-dimensional point cloud scanning devices (such as lidar and structured light depth cameras). After the exhibits are set up, each exhibit is collected using the multimodal sensor array: the multispectral imaging device and the three-dimensional point cloud scanning device are clock-synchronized through a unified time synchronization module, and a unified trigger signal is issued to achieve multimodal synchronous acquisition of the exhibits. During the acquisition process, the multispectral module obtains a sequence of reflective images of the exhibit in a specified band, namely multispectral image data, through continuous scanning or shutter-type acquisition, while the point cloud scanning device obtains complete three-dimensional contour data of the exhibit through a multi-frame synthesis algorithm. Finally, the raw data of each modality is packaged to form an exhibit sensor data set, which is identified and archived with a unique exhibit number.

[0048] During the exhibit data collection process, a positioning camera or laser rangefinder is used to obtain the actual installation position and attitude angle of the exhibit in the exhibition hall. A transformation matrix is ​​calculated from the exhibit's local coordinate system to the exhibition hall's global coordinate system. This transformation matrix is ​​then applied to each point in the exhibit's point cloud data for coordinate transformation and uniform mapping to the global coordinate system, resulting in a spatial point set for all exhibits under a unified spatial reference. The multispectral image data is calibrated for camera pose and mapped to the global coordinate system to obtain the spatial position of the image perception area and construct a geometric projection model of the exhibit image. Cluster analysis or bounding box extraction is performed on the point cloud data of all exhibits in the exhibition hall's global coordinate system to construct the exhibit's spatial boundary volume, where each exhibit corresponds to a geometric boundary. Next, based on the spatial proximity (Euclidean distance between center points), visual obstruction, or thematic associations between exhibits, edge connections between exhibits are constructed, forming an exhibit space topology graph G = (V, E), where V is the node set representing all exhibits; E is the edge set representing the spatial or semantic adjacency between nodes. Each edge can be assigned attributes such as distance, visibility weight, and topic label overlap. The final exhibit space topology graph is stored in a graph database or structured JSON format, achieving standardized representation of exhibit spatial locations and supporting spatial logic analysis across multiple exhibits and regional reasoning about visitor paths.

[0049] Using image processing algorithms, we extract the texture descriptors (such as SIFT, SURF, ORB) in the image area corresponding to each point cloud point in the exhibit point cloud data according to the exhibit space topology map. The extracted texture feature vector is recorded as {S}={s1, s2, ..., s k}, where each element is a texture descriptor of a local area on the surface of the exhibit. At the same time, a semantic label is assigned to each of the above point cloud points based on the exhibit metadata or manual annotation system to generate an associated label set of the exhibit, with the following structure: L tag={(x i ,y i , z i , S i , l i )}, where (x i ,y i , z i ) is the spatial coordinate of each point of the exhibit, S i is the image texture feature data vector, l i These are semantic labels, such as "flower and bird painting", "character portrait", "Qing Dynasty calligraphy", etc.

[0050] After labeling each exhibit region, the coordinates of the exhibit regions are mapped to their associated labels, constructing a structured coordinate-label mapping network. This mapping network is stored and managed using a graph structure or key-value indexing format. Each exhibit is represented as a node in the network, containing multiple child nodes corresponding to the exhibit's local region labels, spatial coordinates, image features, and semantic descriptions. The constructed coordinate-label mapping network is essentially a semantically structured grid based on the exhibit surface. This network not only accurately demarcates the exhibit's spatial structure but also assigns semantic meaning to each location. For example, consider a calligraphy and painting exhibit numbered P001. Through analysis of 3D point cloud data and multispectral image data, it is identified as containing: structural region A (coordinate points A1-A50) with the semantic label "main landscape scene"; region B (coordinate points B1-B20) with the label "inscription"; region C (coordinate points C1-C30) with the label "author's seal"; and region D (coordinate points D1-D100) with the label "face." This coordinate-label mapping network can support fast coordinate reverse label query through a tree index structure or spatial hash function, achieving matching between any observation point and label, and providing efficient matching support for subsequent user behavior identification and linkage display.

[0051] Further, such as Figure 2 As shown, step S300 includes:

[0052] Step S310: performing motion trajectory analysis based on the user viewing behavior dataset, extracting spatiotemporal behavior feature vectors, performing eye tracking based on the user viewing behavior dataset, and drawing an eye focus heat map.

[0053] Step S320: extracting multiple exhibit node information based on the associated tag set.

[0054] Step S330: traversing the coordinate-label mapping network to perform matching calculations on the spatiotemporal behavior feature vector and the exhibit node information to obtain a plurality of matching confidence levels.

[0055] Step S340: performing a determination based on the multiple matching confidence levels in combination with the gaze focus heat map, activating the exhibit recognition module based on the determination result, and obtaining the target exhibit number.

[0056] Specifically, the spatiotemporal behavior feature vector represents the characteristics of a user's movement and dwelling behavior within the exhibition hall. It includes data such as timestamps, location coordinates, and behavior type (e.g., dwell, turn). The gaze focus heatmap is a visualization generated based on the cumulative intensity of user gaze points across time and space, reflecting the areas where users focus their attention. The user's spatial position and eye movement data recorded in the exhibition behavior dataset are analyzed. A trajectory processing algorithm based on Kalman filtering and sliding window smoothing is used to denoise and interpolate the user's continuous position data within the exhibition hall, thereby constructing a high-precision user movement trajectory curve. This is then time-weighted along the time axis to generate a sequence of spatiotemporal behavior feature vectors. Furthermore, gaze sensor data is used to obtain gaze points. An algorithm based on gaze clustering and regional voting is used to identify high-frequency gaze areas and map these gaze points to the exhibition hall's global coordinate system. The gaze frequency and duration of a given spatial area within a unit time are calculated and rendered as a heatmap to construct a user gaze focus heatmap, where color saturation represents the combined intensity of gaze frequency and duration.

[0057] The established coordinate-label mapping network is parsed, and each main exhibit node and its sub-label nodes are traversed. Information such as the exhibit number, regional coordinate boundaries, corresponding texture type, and semantic identifiers are sequentially extracted. To improve matching efficiency, KD-Tree or hash indexing is used to spatially group and label the exhibit nodes, forming a collection of exhibit nodes that includes location features, image features, and semantic descriptions. The exhibits' display hierarchy and exhibition area affiliation are also annotated to facilitate subsequent precision control and constrained matching within the region.

[0058] Match confidence is a quantitative indicator of the degree of match between the current user's behavior data and a specific exhibit. The value range is generally [0, 1], with higher values ​​indicating a more reliable match. Based on a coordinate-label mapping network, all exhibit nodes are traversed. For each exhibit node and its subregion labels, a multi-dimensional matching calculation is performed, one after the other, with the user's spatiotemporal behavior vector. The matching function comprehensively considers three dimensions: a spatial distance metric Ds (e.g., the inverse Euclidean distance function), a temporal synchronization weight Wt (with higher weights for more recent behavior), and a semantic feature similarity Se (e.g., a consistency score for texture type or fixation region labels). The final matching confidence is calculated by normalized linear combination, such as: Cf=α*(1 / Ds)+β*Wt+γ*Se, where α, β, and γ are empirically set weight coefficients, satisfying α+β+γ=1, and α, β, γ∈[0.2, 0.5]. The spatial distance metric Ds is the Euclidean distance between the user's current location and the center of the exhibit. The closer the user is to the exhibit, the greater the possibility of viewing the exhibit, and the higher the matching confidence; the time synchronization weight Wt reflects the temporal freshness of the behavioral data, with a value between [0, 1]. The calculation formula is defined as Wt=e -λ*Δt , where Δt is the interval between the current time and the time the behavior was recorded, and λ is the attenuation coefficient. The closer the time, the greater the weight, which can highlight the currently viewed exhibit. The semantic feature similarity Se is used to measure the degree of match between the user's gaze area and the exhibit label. It can be calculated by cosine similarity using texture features, reflecting the degree of consistency between the user's gaze area and the exhibit's semantic label. Separate semantics indicate a reliable match. For each candidate exhibit node, the three matching dimensions are calculated and substituted into the formula to obtain a normalized matching confidence Cf∈[0,1]. The top N candidate exhibit nodes are sorted from high to low according to confidence, and the top N candidate exhibit nodes are selected as the candidate set for subsequent determination. N is the candidate exhibit node capacity, which can be set to a positive integer between 3 and 10, depending on the exhibition hall density and system response time.

[0059] After obtaining multiple matching confidence scores, the system further screens exhibits in the overlapping areas of the candidate set by combining the heatmap region locations with the physical coordinates of the candidate exhibits. The exhibit recognition module is then activated, which iterates through the exhibits in the current overlapping area. It then comprehensively scores the candidate exhibits based on the matching confidence scores, the gaze focus heatmap, and the spatial overlap ratio of each exhibit node. The exhibit node with the highest score is selected as the final target, and its number is identified and output as the target exhibit number. By integrating the complementary information of confidence ranking and gaze focus maps, the system achieves high-precision identification of exhibits of interest while ensuring real-time performance, significantly improving the accuracy of interactive triggering and the ability to provide personalized recommendations.

[0060] Furthermore, step S310 includes:

[0061] Step S311: performing spatial displacement analysis based on the user exhibition viewing behavior dataset to obtain a behavior trajectory sequence of the target user.

[0062] Step S312: performing trajectory classification according to the behavior trajectory sequence to obtain multiple mobility pattern categories, performing behavior mining based on the multiple mobility pattern categories to obtain a spatiotemporal behavior feature vector.

[0063] Specifically, the method extracts continuous location data from a user's exhibition behavior dataset and sequentially organizes it with timestamps. A sliding window filtering algorithm is then used to smooth the location data and remove transient, abnormal drift values. Furthermore, a micro-displacement determination mechanism is introduced to mark dwell points (e.g., those whose speed falls below a threshold and who remain there for longer than a preset time). Finally, the displacement data of the user's complete exhibition viewing process is structured and stored as a trajectory sequence along the timeline. Each trajectory point contains: a timestamp t, three-dimensional coordinates (x, y, z), a displacement vector Δd, and a dwell identifier r, forming a user behavior trajectory sequence.

[0064] A large number of historical user behavior trajectory sequences are collected and segmented and coded, and multidimensional features are extracted. These include average speed per segment, turn frequency, dwell point distribution density, maximum dwell time, and trajectory closure. A sample feature matrix is ​​constructed. Unsupervised clustering algorithms (such as DBSCAN or K-Means) are used to cluster trajectory patterns within this feature matrix, identifying several common mobility pattern categories, such as "sequential touring," "theme hopping," and "high-retention intensive reading." Typical behavior templates are constructed for each type of mobility pattern to serve as a reference for subsequent classification. The previously acquired user behavior trajectory sequences are segmented and coded, and multidimensional features are extracted. Features such as average speed per segment, turn frequency, dwell point distribution density, maximum dwell time, and trajectory closure of the target user are extracted, and a feature matrix is ​​constructed. Subsequently, typical behavior templates are traversed according to the feature matrix to match the target user's corresponding mobility pattern category. Behavior mining is performed based on the feature bias under this category, and behavioral indicators such as the residence-transfer ratio, interest center drift rate, and average gaze density are extracted. These behavioral indicators are fused to generate a spatiotemporal behavior feature vector, which serves as the basic input for subsequent heat map modeling and exhibit matching, significantly improving the accuracy and adaptability of user behavior modeling.

[0065] Furthermore, step S320 includes:

[0066] Step S321: performing feature analysis based on the associated tag set in combination with the three-dimensional point cloud data to obtain point cloud geometric features.

[0067] Step S322: performing feature analysis based on the associated tag set in combination with the multispectral image data to obtain spectral texture features.

[0068] Step S323: performing multimodal feature fusion on the point cloud geometric features and the spectral texture features according to the associated label set to construct an associated feature vector.

[0069] Step S324: performing hierarchical node analysis according to the associated feature vector to determine the plurality of exhibit node information.

[0070] Specifically, based on the established associated label set, a subset of three-dimensional point cloud data associated with each exhibit label is called. The voxel grid downsampling algorithm is then used to pre-process the original point cloud to compress redundant data and improve processing efficiency. Next, point cloud slice projection and principal direction fitting algorithms (such as PCA principal axis analysis) are used to extract local structural information, and indicators such as curvature distribution, boundary point density, and normal vector consistency in the point cloud are calculated. At the same time, through edge extraction and convex-concave analysis, a local geometric contour map is generated as the basic geometric representation for subsequent feature fusion. All geometric feature parameters are bound to their label structure according to the exhibit number and registered as point cloud geometric feature entries in the global exhibit information set. Through the refined analysis of the exhibit point cloud model, the geometric structure of calligraphy and painting exhibits in the spatial dimension can be accurately identified, effectively supporting subsequent node attribute reasoning and exhibit status recognition, and enhancing the exhibition system's intelligent understanding of the physical object form.

[0071] By retrieving the multispectral image bound to the exhibit label, image registration and illumination normalization are performed on each band of the image in turn. Subsequently, algorithms such as GLCM (Gray Level Co-occurrence Matrix) and LBP (Local Binary Pattern) are used to extract image texture directionality, roughness, granularity, and other information, and the main components of surface color and color saturation are analyzed in combination with color histograms. In addition, the moisture distribution and dye residue of the paper substrate are identified in the near-infrared and ultraviolet bands to further determine the potential material and aging state. Finally, all image features are encoded as texture feature vectors, bound to their label structure according to the exhibit number, and registered as spectral texture feature entries in the global exhibit information set. By combining multi-band image analysis, the material properties, surface structure, and potential aging of calligraphy and painting can be effectively identified, thereby achieving high-dimensional modeling of the exhibit content and status, providing a perceptual basis for layered identification and personalized recommendations of exhibits.

[0072] Using the exhibit number as the index, the extracted point cloud geometric feature vectors and spectral texture feature vectors are subjected to modal normalization and dimensionality unification. The fusion algorithm adopts a feature splicing and weighted fusion strategy, in which a priority weight factor (such as the state degradation risk score) is introduced for preservation state features. At the same time, the three-dimensional center coordinates of the exhibit are used as the spatial anchor point and combined with the fusion vector to generate the associated feature vector structure: F = (x, y, z, S i , P), where x, y, z are the coordinates of the exhibit point cloud, S iis the spectral texture feature vector, and P is the preservation status score.

[0073] Based on the associated feature vectors, a structured deconstruction operation is performed according to the preset exhibit node template. First, the basic attribute layer (such as exhibit name, number, material, size, and coordinates) is analyzed, and then derived attributes (such as surface aging level, light sensitivity, and mobile adaptability) are mined. Finally, through content-based semantic label matching (such as style category, historical period, author, etc.), the edges connecting it to other exhibits are constructed. All node information is organized as graph-structured data and injected into the exhibit graph database, forming a structured exhibition network with exhibits as nodes and attributes and relationships as edges, namely a coordinate-label mapping network. Through hierarchical parsing, each exhibit is transformed into a knowledge graph node with semantic understanding capabilities, achieving deep modeling of exhibition objects from the physical to the semantic level, effectively supporting personalized navigation, intelligent recommendations, and automatic exhibit association and interaction.

[0074] Furthermore, step S340 includes:

[0075] Step S341: traverse the gaze focus heat map to extract eye movement feature data, and calculate the area dwell time based on the eye movement feature data.

[0076] Step S342: Screening the coordinates of the exhibits according to the multiple matching confidence levels to generate a candidate exhibit set.

[0077] Step S343: When the area residence time is greater than or equal to the preset time, the exhibit identification module is activated to traverse the candidate exhibit set for number identification to determine the target exhibit number.

[0078] Specifically, the gaze focus heat map is traversed at the pixel level to locate the high-heat areas (i.e., areas with color intensity greater than the set threshold), and the corresponding eye movement event sequences are extracted. By performing time-domain aggregation on the start and end times of eye movement events (such as gaze point sequences) in these hot zones, the user's dwell time in each hot zone is calculated. In order to eliminate the interference of rapid scans, a minimum gaze threshold (such as 80ms) is set to filter out non-attention eye movement events to ensure the authenticity and validity of the dwell time. All calculation results are bound to the user's heat map index structure to provide a timeliness criterion for subsequent exhibit matching. Through high-precision eye movement dwell calculation, accurate identification of the user's real focus area is achieved, ensuring that the trigger conditions for activating the recognition module have sufficient human-computer interaction basis, and improving the reliability of target exhibit recognition and the fit with user intentions.

[0079] The coordinates of the exhibits are screened based on the multiple matching confidences calculated in step S330, and exhibits with matching confidences higher than a set threshold are included in the candidate exhibit set. At the same time, their spatial range is limited to not exceed the viewing cone radius (e.g., 2.5 meters) of the current user's gaze area.

[0080] The heat map region locations and the physical coordinates of candidate exhibits are combined to further screen for exhibits in overlapping areas of the candidate set. Upon confirming that a user has spent at least a preset dwell time in a particular heat map region, the exhibit recognition module is immediately activated. This module traverses the current set of candidate exhibits, focusing on the center coordinates of the high-heat zone in the heat map as the primary judgment window and calculating the spatial overlap ratio Rh between the candidate exhibit and each exhibit node. Each candidate exhibit's confidence score Cf and the heat zone overlap ratio Rh are then jointly scored (e.g., Score = Cf × Rh). The exhibit node with the highest score is selected as the current target exhibit, identified, and its number output. If multiple exhibits have similar scores, a fuzzy inference mechanism is used to further optimize the numbering decision based on image similarity and user preferences. To improve accuracy, a minimum confidence threshold (e.g., Cf > 0.65) is set. Objects below this threshold are marked as "uncertain" and further judgment is made based on accumulated data.

[0081] Furthermore, step S400 includes:

[0082] Step S410: dynamically synthesizing the multimodal sensor data according to the target exhibit number to obtain a multimodal synthetic data set.

[0083] Step S420: Performing a linkage display of calligraphy and painting for the target user based on the multimodal synthetic data set to generate display information, wherein the display information includes audio display information, brushstroke trajectory display information, and three-dimensional holographic image display information.

[0084] Step S430: activating the exhibit interaction module to perceive the target user's multi-dimensional interaction data set in real time, where the multi-dimensional interaction data set includes voice interaction data, action interaction data, and eye movement interaction data.

[0085] Step S440: Perform interaction based on the voice interaction data in combination with the audio presentation information to generate a first interaction perception parameter.

[0086] Step S450: performing interaction based on the action interaction data combined with the brushstroke trajectory display information to generate a second interaction perception parameter.

[0087] Step S460: performing interaction based on the eye movement interaction data in combination with the three-dimensional holographic image display information to generate a third interaction perception parameter.

[0088] Step S470: performing a navigation analysis of the calligraphy and painting exhibition for the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, and generating a personalized calligraphy and painting exhibition navigation path.

[0089] Specifically, after receiving the target exhibit number, the corresponding data for that exhibit is found within the existing multimodal sensor data. A time synchronization mechanism and coordinate normalization method are used to fuse the data from each modality. Tensor joint modeling or a cross-attention mechanism are used to enhance inter-modal complementarity, ultimately generating a unified multimodal synthetic dataset for subsequent display.

[0090] Based on a multimodal synthetic dataset, combined with pre-stored audio and brushstroke reconstruction data related to the exhibits, a linked display interface is invoked to generate three types of display information: audio display information, brushstroke trajectory display information, and 3D holographic image display information. The audio of the exhibit explanation (audio display information) can be output through audio equipment; the placement, force, and direction of the brush strokes in the artwork (brushstroke trajectory display information) can be replayed on the screen; and the physical appearance of the exhibit (3D holographic image display information) can be reconstructed using a 3D holographic projection system. The display information is synchronized with the user's spatial orientation, ensuring the coordinated integration and consistent rhythm of all types of information.

[0091] Activate the exhibit interaction module to conduct full-cycle monitoring of the target user, collecting a multi-dimensional interaction dataset consisting of voice interaction data, gesture interaction data, and eye movement interaction data. The module receives instructions such as explanation requests and questions through a voice channel (e.g., a microphone) (voice interaction data); recognizes user gestures and clicks through a camera or motion capture device (gesture interaction data); and records gaze points and pulsation rates through an eye tracker (eye movement interaction data). These data are timestamped and stored in a multi-dimensional interaction dataset, providing input for subsequent personalized recognition and navigation path adjustments.

[0092] Next, information fusion and response matching are performed on the three types of interaction data to generate corresponding behavioral response parameters, namely the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter. These three parameters correspond to the three channels of voice, action, and eye movement, respectively. For the voice channel, the user's voice interaction data is compared with the audio explanation content through a semantic understanding model to generate the first interaction perception parameter. For the action channel, the user's gesture trajectory is mapped to the stroke animation path, and the spatial trajectory overlap rate and the action rhythm matching degree are calculated to generate the second interaction perception parameter. For the eye movement channel, the interest weight is calculated based on the overlap time between the gaze point and the holographic projection area, and the third interaction perception parameter is output. The interaction perception parameters of the three channels comprehensively characterize the user's active attention dimension to the exhibits, providing clear weight support for the next step of guide optimization, making the basis for personalized recommendations more objective and multidimensional.

[0093] The first, second, and third interaction perception parameters are weighted and fused to construct a user's immediate interest vector, which is then matched against exhibit nodes for similarity. Combining the exhibition hall's spatial topology with the distribution of exhibit locations, a graph traversal algorithm generates the shortest, most relevant, personalized guided tour path. Path planning results are automatically updated based on multi-dimensional interaction weights. For example, users who prefer audio content are prioritized for detailed explanations, while those who prefer brushwork imitation are prioritized for interactive calligraphy areas. By integrating perception parameters with a coordinate-label mapping network, a dynamic coupling between guided tour content and user interests is achieved, transforming exhibition routes from static planning to proactive guidance, significantly enhancing the personalized visiting experience and in-depth exhibition engagement.

[0094] Furthermore, step S470 includes:

[0095] Step S471: Perform dynamic interest analysis on the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to construct a user interest vector.

[0096] Step S472: Optimize the path according to the user interest vector and the exhibit coordinate system to construct a moving path.

[0097] Step S473: Mapping the moving path to the exhibition hall global coordinate system for dynamic guidance marking to generate the personalized calligraphy and painting exhibition guide path.

[0098] Specifically, the user interest vector is a multi-dimensional vector that represents the user's current interest state, including the user's interest intensity in different exhibits, exhibition areas, exhibit categories, etc. The user interest vector is generated by weighting the first, second, and third interaction perception parameters. The weight of each type of perception parameter is dynamically adjusted based on the user's interaction frequency, interaction duration, interaction depth, etc. For example, if a user asks for background information about an exhibit multiple times, the weight of the voice interaction parameter of the exhibit will be increased; if the user stays in front of a certain exhibit for a long time, the weight of the exhibit in the interest vector will be increased accordingly. In this way, the user's real-time interest changes can be accurately captured.

[0099] Using path planning algorithms (such as the A-STAR algorithm and Dijkstra's algorithm), exhibits are prioritized based on user interest vectors. Then, based on the exhibit coordinate system, the exhibits are rearranged according to the shortest path principle to generate a user's tour path. This path planning process comprehensively considers factors such as the distance between exhibits, user interests, and the spatial and temporal relationships of exhibits to ensure an efficient and organized visit. During the user's visit, the path is adjusted and optimized in real time to adapt to changing user interests, enhancing interactivity and engagement.

[0100] Based on the optimized user path, the global exhibition hall coordinate system is used to map it to the exhibition space, generating dynamic guide markers. These dynamic guide markers are displayed in real time via displays, indicator lights, or floor projections within the exhibition hall, guiding the user to their next target exhibit. Voice, image, or map prompts can also be pushed to mobile devices to help users quickly reach the next exhibit. By continuously tracking the user's spatial location, the path is adjusted in real time to ensure that the user remains on the correct exhibit path. By integrating the user's optimized path with the exhibition hall's global coordinate system and providing real-time feedback through dynamic guide markers, the accuracy and convenience of personalized guided tours are greatly improved, optimizing the overall visitor experience.

[0101] In summary, the interactive method for calligraphy and painting exhibitions provided by the embodiments of the present invention has the following beneficial effects:

[0102] Through sensing devices, the audience's behavior at the exhibition site is collected in real time to build a user viewing behavior dataset, providing basic data support for subsequent identification and personalized recommendations. By traversing the exhibition area for multi-dimensional image calibration, a coordinate-label mapping network is constructed between the exhibit space coordinates and semantic information, creating conditions for accurate matching between user behavior and exhibits. By traversing the coordinate-label mapping network set based on the user viewing behavior dataset and activating the exhibit recognition module, the target exhibit number currently being focused on by the user is automatically identified, realizing the association conversion from behavior data to exhibit entity. According to the target exhibit number, calligraphy and painting are displayed to the target user in a linked manner, and display information is generated. The exhibit interaction module is activated to conduct in-depth interaction with the target user based on the display information, generating a personalized calligraphy and painting exhibition guide path for the target user, and enhancing the immersive viewing experience.

[0103] Overall, the embodiments of the present invention achieve real-time perception of user interests and personalized guided services in calligraphy and painting exhibitions by constructing a user behavior dataset, a coordinate-label mapping network for image calibration, and an exhibit identification and personalized linkage display mechanism based on behavior data, significantly improving the interactivity, intelligence level and user immersive experience of the exhibition.

[0104] Example 2, as Figure 3 As shown, based on the same inventive concept as the aforementioned embodiment 1, an embodiment of the present invention provides an interactive system for calligraphy and painting exhibitions, the system comprising:

[0105] The behavior data collection module 10 is used to collect target users in real time based on the exhibition area, and synchronously obtain user behavior data and exhibition environment data through a multimodal sensor array to construct a user exhibition behavior data set.

[0106] The image calibration module 20 is used to traverse the exhibit area and perform multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data and texture feature data of the exhibit surface, and establishes a coordinate-label mapping network through spatiotemporal association.

[0107] The target identification module 30 is used to traverse the coordinate-label mapping network set according to the user exhibition viewing behavior data set to match exhibit nodes, activate the exhibit identification module, and identify the target exhibit number.

[0108] The interactive display module 40 is used to display calligraphy and painting to the target user in a linked manner according to the target exhibit number, generate display information, start the exhibit interaction module to conduct in-depth interaction with the target user in combination with the display information, generate a personalized calligraphy and painting exhibition guide path for the target user, and adjust the guide path in real time to adapt to the user's dynamic behavior.

[0109] Furthermore, the behavior data collection module 10 of the embodiment of the present invention is further configured to perform the following steps:

[0110] A multimodal sensor array is deployed around the calligraphy and painting exhibits in the exhibition area, and the target user is captured in real time by the multimodal sensor array to obtain a user sensor data set; a mapping relationship between the exhibit coordinate system and the exhibition hall global coordinate system is established, and the user sensor data is converted into exhibition viewing behavior metadata; the exhibition viewing behavior metadata is subjected to spatiotemporal alignment processing according to the mapping relationship to generate a three-dimensional behavior vector; semantic parsing is performed based on the three-dimensional behavior vector to generate a semantic association graph, and user behavior analysis is performed according to the semantic association graph to construct multi-source behavior event data, and the multi-source behavior event data is added to the user exhibition viewing behavior dataset.

[0111] Furthermore, the image calibration module 20 of the embodiment of the present invention is further configured to perform the following steps:

[0112] The multimodal sensor array traverses the exhibition area to synchronously collect exhibits, thereby obtaining an exhibit sensor data set, wherein the exhibit sensor data set includes multispectral image data of the exhibit surface and three-dimensional point cloud data; the exhibit sensor data set is mapped to the exhibition hall global coordinate system to generate an exhibit space topology map; the three-dimensional point cloud data and the multispectral image data of the exhibit surface are associated and mapped according to the exhibit space topology map to generate an associated label set; the calligraphy and painting exhibits are calibrated based on the associated label set to construct the coordinate-label mapping network.

[0113] Furthermore, the target recognition module 30 of the embodiment of the present invention is further configured to perform the following steps:

[0114] Based on the user viewing behavior dataset, motion trajectory analysis is performed to extract the spatiotemporal behavior feature vector; based on the user viewing behavior dataset, line of sight tracking is performed to draw a line of sight focus heat map; based on the associated label set, multiple exhibit node information is extracted; the coordinate-label mapping network is traversed to perform matching calculations on the spatiotemporal behavior feature vector and the exhibit node information to obtain multiple matching confidence levels; a judgment is made according to the multiple matching confidence levels combined with the line of sight focus heat map, and the exhibit recognition module is activated according to the judgment result to obtain the target exhibit number.

[0115] Furthermore, the target recognition module 30 of the embodiment of the present invention is further configured to perform the following steps:

[0116] Based on the user exhibition behavior dataset, spatial displacement analysis is performed to obtain a behavior trajectory sequence of the target user; trajectory classification is performed according to the behavior trajectory sequence to obtain multiple movement pattern categories; behavior mining is performed based on the multiple movement pattern categories to obtain a spatiotemporal behavior feature vector.

[0117] Furthermore, the target recognition module 30 of the embodiment of the present invention is further configured to perform the following steps:

[0118] Based on the associated tag set, feature analysis is performed in combination with the three-dimensional point cloud data to obtain point cloud geometric features; based on the associated tag set, feature analysis is performed in combination with the multispectral image data to obtain spectral texture features; the point cloud geometric features and the spectral texture features are subjected to multimodal feature fusion according to the associated tag set to construct an associated feature vector; hierarchical node parsing is performed according to the associated feature vector to determine the multiple exhibit node information.

[0119] Furthermore, the target recognition module 30 of the embodiment of the present invention is further configured to perform the following steps:

[0120] The gaze focus heat map is traversed to extract eye movement feature data, and a region dwell time is calculated based on the eye movement feature data; exhibit coordinates are screened based on the multiple matching confidence levels to generate a candidate exhibit set; and when the region dwell time is greater than or equal to a preset time, the exhibit recognition module is activated to traverse the candidate exhibit set for number recognition to determine the target exhibit number.

[0121] Furthermore, the interactive display module 40 of the embodiment of the present invention is further configured to perform the following steps:

[0122] According to the target exhibit number, multimodal sensor data is called for dynamic synthesis to obtain a multimodal synthesis data set; based on the multimodal synthesis data set, calligraphy and painting are linked displayed to the target user to generate display information, and the display information includes audio display information, brush stroke trajectory display information, and three-dimensional holographic image display information; the exhibit interaction module is started to perceive the target user's multidimensional interaction data set in real time, and the multidimensional interaction data set includes voice interaction data, action interaction data, and eye movement interaction data; based on the voice interaction data combined with the audio display information, interaction is performed to generate a first interaction perception parameter; based on the action interaction data combined with the brush stroke trajectory display information, interaction is performed to generate a second interaction perception parameter; based on the eye movement interaction data combined with the three-dimensional holographic image display information, interaction is performed to generate a third interaction perception parameter; based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, a guided analysis of the calligraphy and painting exhibition is performed on the target user to generate a personalized calligraphy and painting exhibition guided path.

[0123] Furthermore, the interactive display module 40 of the embodiment of the present invention is further configured to perform the following steps:

[0124] Based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, a dynamic interest analysis of the target user is performed to construct a user interest vector; path optimization is performed according to the user interest vector in combination with the exhibit coordinate system to construct a movement path; the movement path is mapped to the exhibition hall global coordinate system for dynamic guidance marking to generate the personalized calligraphy and painting exhibition guide path.

[0125] Through the above detailed description of an interactive method applied to calligraphy and painting exhibitions in this specification, those skilled in the art can clearly understand an interactive system applied to calligraphy and painting exhibitions in this embodiment. For the system disclosed in Example 2, since it corresponds to the method disclosed in Example 1 and has corresponding functional modules and beneficial effects, the relevant details can be referred to the method part description.

[0126] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An interactive method applied to calligraphy and painting exhibitions, characterized in that: The method comprises: Real-time data collection of target users is carried out based on the exhibition area. User behavior data and exhibition environment data are simultaneously acquired through a multimodal sensor array to construct a user viewing behavior dataset. Traversing the exhibit area to perform multi-dimensional image calibration, the multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data and texture feature data of the exhibit surface, and establishing a coordinate-label mapping network through spatiotemporal association; According to the user exhibition behavior data set, the coordinate-label mapping network set is traversed to match exhibit nodes, and an exhibit identification module is activated to identify the target exhibit number; Performing a linked display of calligraphy and painting for the target user according to the target exhibit number, generating display information, activating the exhibit interaction module to conduct in-depth interaction with the target user based on the display information, generating a personalized calligraphy and painting exhibition guide path for the target user, and adjusting the guide path in real time to adapt to the user's dynamic behavior; The method includes: traversing the coordinate-label mapping network set according to the user exhibition viewing behavior data set to activate the exhibit identification module and identify the target exhibit number. Performing motion trajectory analysis based on the user viewing behavior dataset, extracting spatiotemporal behavior feature vectors, performing eye tracking based on the user viewing behavior dataset, and drawing an eye focus heat map; Extract multiple exhibit node information based on associated tag sets; Traversing the coordinate-label mapping network to perform matching calculations on the spatiotemporal behavior feature vector and the exhibit node information to obtain multiple matching confidences; A determination is performed according to the multiple matching confidences in combination with the sight focus heat map, and the exhibit identification module is activated according to the determination result to obtain the target exhibit number.

2. The interactive method for calligraphy and painting exhibition according to claim 1, characterized in that: Based on the exhibition area, target users are collected in real time. User behavior data and exhibition environment data are simultaneously acquired through a multimodal sensor array to construct a user viewing behavior dataset. The method includes: Deploying a multimodal sensor array around the calligraphy and painting exhibits in the exhibition area, and capturing target users in real time through the multimodal sensor array to obtain a user sensor data set; Establishing a mapping relationship between the exhibit coordinate system and the exhibition hall global coordinate system, and converting the user sensor data into exhibition viewing behavior metadata; Performing spatiotemporal alignment processing on the exhibition viewing behavior metadata according to the mapping relationship to generate a three-dimensional behavior vector; Semantic analysis is performed based on the three-dimensional behavior vector to generate a semantic association graph, user behavior analysis is performed based on the semantic association graph to construct multi-source behavior event data, and the multi-source behavior event data is added to the user exhibition viewing behavior dataset.

3. The interactive method for calligraphy and painting exhibition according to claim 2, characterized in that: Traversing the exhibit area to perform multi-dimensional image calibration, the multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data and texture feature data of the exhibit surface, and establishing a coordinate-label mapping network through spatiotemporal association. The method includes: The multimodal sensor array traverses the exhibit area to synchronously collect exhibits to obtain an exhibit sensor data set, wherein the exhibit sensor data set includes multispectral image data and three-dimensional point cloud data of the exhibit surface; Mapping the exhibit sensor data set to the exhibition hall global coordinate system to generate an exhibit space topology map; Correlation mapping is performed on the three-dimensional point cloud data and the multispectral image data of the exhibit surface according to the exhibit space topology map to generate the correlation label set; The calligraphy and painting exhibits are calibrated based on the associated label set, and the coordinate-label mapping network is constructed.

4. The interactive method for calligraphy and painting exhibition according to claim 1, characterized in that: Performing motion trajectory analysis based on the user exhibition viewing behavior dataset to extract spatiotemporal behavior feature vectors, the method includes: Perform spatial displacement analysis based on the user exhibition behavior dataset to obtain the target user's behavior trajectory sequence; Trajectory classification is performed according to the behavior trajectory sequence to obtain multiple movement pattern categories, and behavior mining is performed based on the multiple movement pattern categories to obtain a spatiotemporal behavior feature vector.

5. The interactive method for calligraphy and painting exhibition according to claim 1, characterized in that: The method for extracting multiple exhibit node information based on the associated tag set includes: Perform feature analysis based on the associated tag set and the three-dimensional point cloud data to obtain point cloud geometric features; Performing feature analysis based on the associated tag set in combination with the multispectral image data to obtain spectral texture features; Performing multimodal feature fusion on the point cloud geometric features and the spectral texture features according to the associated label set to construct an associated feature vector; Hierarchical node analysis is performed according to the associated feature vector to determine the plurality of exhibit node information.

6. The interactive method for calligraphy and painting exhibition according to claim 1, characterized in that: Performing a determination based on the multiple matching confidence levels in combination with the sight focus heat map, activating the exhibit identification module based on the determination result, and obtaining a target exhibit number, the method includes: Traversing the gaze focus heat map to extract eye movement feature data, and calculating the region dwell time based on the eye movement feature data; Screening the coordinates of the exhibits according to the multiple matching confidence levels to generate a set of candidate exhibits; When the residence time in the area is greater than or equal to a preset time, the exhibit identification module is activated to traverse the candidate exhibit set to perform number identification and determine the target exhibit number.

7. The interactive method for calligraphy and painting exhibition according to claim 2, characterized in that: Performing a linked display of calligraphy and painting for a target user according to the target exhibit number, generating display information, activating an exhibit interaction module to perform in-depth interaction with the target user based on the display information, generating a personalized calligraphy and painting exhibition guide path for the target user, and adjusting the guide path in real time to adapt to the user's dynamic behavior, the method includes: Calling the multimodal sensor data according to the target exhibit number to perform dynamic synthesis to obtain a multimodal synthetic data set; Performing a linkage display of calligraphy and painting for a target user based on the multimodal synthetic data set to generate display information, wherein the display information includes audio display information, brushstroke trajectory display information, and three-dimensional holographic image display information; Activate the exhibit interaction module to perceive a multi-dimensional interaction data set of a target user in real time, wherein the multi-dimensional interaction data set includes voice interaction data, action interaction data, and eye movement interaction data; Generate a first interaction perception parameter based on the voice interaction data combined with the audio presentation information; Performing interaction based on the action interaction data combined with the brushstroke trajectory display information to generate a second interaction perception parameter; Generate a third interaction perception parameter based on the eye movement interaction data and the three-dimensional holographic image display information; Based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter, a navigation analysis of the calligraphy and painting exhibition is performed on the target user to generate a personalized calligraphy and painting exhibition navigation path.

8. The interactive method for calligraphy and painting exhibition according to claim 7, characterized in that: The method includes: performing a guide analysis of a calligraphy and painting exhibition for a target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to generate a personalized calligraphy and painting exhibition guide path. Performing dynamic interest analysis on the target user based on the first interaction perception parameter, the second interaction perception parameter, and the third interaction perception parameter to construct a user interest vector; Optimize the path according to the user interest vector and the exhibit coordinate system to construct a moving path; The moving path is mapped to the exhibition hall global coordinate system for dynamic guidance marking to generate the personalized calligraphy and painting exhibition guide path.

9. An interactive system applied to calligraphy and painting exhibitions, characterized in that: The system is used to execute the interactive method for calligraphy and painting exhibitions according to any one of claims 1 to 8, comprising: The behavioral data collection module is used to collect data on target users in real time based on the exhibition area. It uses a multimodal sensor array to synchronously obtain user behavior data and exhibition environment data to construct a user viewing behavior dataset. An image calibration module is used to traverse the exhibit area and perform multi-dimensional image calibration. The multi-dimensional image calibration includes multi-spectral image data, three-dimensional point cloud data, and texture feature data of the exhibit surface, and establishes a coordinate-label mapping network through spatiotemporal association; A target identification module is used to traverse the coordinate-label mapping network set according to the user exhibition viewing behavior data set to match exhibit nodes, activate the exhibit identification module, and identify the target exhibit number; The interactive display module is used to display calligraphy and painting to the target user in a linked manner according to the target exhibit number, generate display information, start the exhibit interaction module to conduct in-depth interaction with the target user based on the display information, generate a personalized calligraphy and painting exhibition guide path for the target user, and adjust the guide path in real time to adapt to the user's dynamic behavior.

Citation Information

Patent Citations

  • Information visual management system and method based on digital twinning

    CN119271899A

  • Exhibition first push item optimization method combining user portrait and user positioning

    CN120067459A