Interactive museum handheld intelligent guide method and system

By identifying palm movements and emotional recognition technology, combined with realism rendering, the problems of insufficient convenience, vividness and visual effects of the existing intelligent guide system for cultural activities are solved, and the virtual guide with independent control is realized, which enhances the vividness and immersion of the user's experience.

CN120353340APending Publication Date: 2025-07-22HEFEI UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510489730.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing intelligent guide system for cultural activities has shortcomings in the vividness of user experience, the convenience of interaction and the authenticity of visual effects. It lacks convenient interaction methods, single information presentation form and realistic rendering technical support, resulting in insufficient vividness, autonomy and immersion of user experience.

Method used

By identifying palm movements to evoke and close virtual characters, combining emotional recognition technology and real-life rendering technology, AR devices are used to obtain three-dimensional point cloud data to establish a local coordinate system, using reverse kinematics algorithms to adjust the posture of virtual characters, and dynamically adjust the rendering accuracy of virtual characters to achieve convenient control and emotional response of virtual characters.

Benefits of technology

It improves the convenience and vividness of interaction, enhances the immersion and autonomy of users, improves the realism and emotional interaction of virtual guides, and optimizes the flexibility and personalized experience of the guide.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353340A_ABST
    Figure CN120353340A_ABST
Patent Text Reader

Abstract

The invention discloses an interactive museum handheld intelligent guide method and system, and the method comprises the steps: obtaining the three-dimensional point cloud data of a palm of a user in real time through a camera of AR equipment, and building a local coordinate system with the geometric center of the palm center as an original point; determining corresponding action judgment conditions for calling and closing the virtual character; when the palm center normal vector rotation angle meets a preset threshold value range and lasts for a first preset duration, determining a calling instruction, activating a virtual character to be bound to a local coordinate system, and dynamically rendering and displaying the virtual character in the AR equipment; the posture of the virtual character is adjusted in real time through an inverse kinematics algorithm, so that the visual gravity center of the virtual character is aligned with the sight line direction of the user; detecting the double-palm contact area of the user, and when the contact area proportion exceeds a set threshold value and lasts for a second preset duration; if yes, determining the instruction as a closing instruction, and triggering the disappearance animation of the virtual character. According to the invention, the convenience of interaction is optimized, the vividness of user experience is improved, and the flexibility and autonomy of navigation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of augmented reality and intelligent navigation, and more specifically, to an interactive museum handheld intelligent guide method and system. Background Art

[0002] Currently, with the development of augmented reality (AR) technology, its application in the field of cultural activity navigation has gradually attracted attention. For example, patent document CN110009793A (publication date: July 12, 2019) discloses an intelligent guide system for cultural activities. This system realizes the navigation function through AR glasses. When the camera detects the image feature information of the target object, it will search for the corresponding cultural information in the database and render it on the AR glasses.

[0003] However, this existing technology has some limitations. First, this system mainly relies on the description methods of text and images. This single form of information presentation is relatively boring and difficult to meet the needs of tourists for vivid and interesting experiences. Second, this system lacks a convenient interaction method. Users cannot flexibly control the appearance and hiding of the guide according to their own needs, which affects the autonomy and fluency of the viewing experience. In addition, the existing technology also needs to be improved in the rendering effect of virtual characters. Lack of the support of realistic rendering technology results in the visual effect of the virtual guide not being realistic enough, making it difficult for users to obtain an immersive experience.

[0004] In summary, although the existing technology has realized the basic AR navigation function in the field of cultural activity navigation, it has deficiencies in terms of the vividness of the user experience, the convenience of interaction, and the authenticity of the visual effect. Summary of the Invention

[0005] In view of this, the present invention provides an interactive museum handheld intelligent guide method and system, which can solve the deficiencies of the existing intelligent guide system for cultural activities. The present invention calls out and closes virtual characters on the user's palm by recognizing palm movements, and combines emotion recognition technology and realistic rendering technology to improve the interaction experience, and is applied to the intelligent navigation scenario of cultural places such as museums.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, an embodiment of the present invention provides an interactive museum handheld intelligent guide method, including the following steps:

[0008] Real-time obtain the three-dimensional point cloud data of the user's palm through the camera of the AR device, and establish a local coordinate system with the geometric center of the palm as the origin;

[0009] Based on the local coordinate system, determine the corresponding action determination conditions for calling out and closing the virtual character;

[0010] When the rotation angle of the palm normal vector satisfies the preset threshold range and lasts for the first preset duration, it is determined as a call-out instruction, activating the virtual character bound to the local coordinate system and dynamically rendering and displaying it in the AR device;

[0011] The posture of the virtual character is adjusted in real time through the inverse kinematics algorithm to align its visual center of gravity with the user's line of sight direction;

[0012] The contact area of the user's two palms is detected. When the proportion of the contact area exceeds the set threshold and lasts for the second preset duration, it is determined as a close instruction, triggering the disappearance animation of the virtual character.

[0013] Furthermore, the following steps are also included:

[0014] Combining the eye movement tracking data with the physiological signals, dynamically adjusting the rendering precision and interaction behavior of the virtual character to respond to the user's emotional state.

[0015] Furthermore, when the rotation angle of the palm normal vector satisfies the preset threshold range and lasts for the first preset duration, it is determined as a call-out instruction, including:

[0016] When the rotation angle θ of the palm normal vector satisfies: 82° ≤ θ ≤ 88°, and the duration is 0.3 seconds to 0.5 seconds, and the movement trajectory satisfies the constraint conditions, the call-out of the virtual character is triggered;

[0017] The formula is:

[0018]

[0019] Among them, N represents the total number of video frames; k represents the video frame index; θ(t k ) represents the rotation angle of the palm normal vector of the kth frame; Δt represents the sliding time window, and t0 represents the time point when the acceleration integration starts, that is: the initial moment when the user starts to move the palm; represents the real-time position of the palm origin in the local coordinate system, and the integrand is the L2 norm of the acceleration vector of the palm movement; 15mm / S 2 represents the upper limit of integration, which is the maximum allowable acceleration accumulation.

[0020] Furthermore, the contact area of the user's two palms is detected. When the proportion of the contact area exceeds the set threshold and lasts for the second preset duration, it is determined as a close instruction; including:

[0021] When the determination threshold for detecting the contact area of the user's two palms is more than 78% of the total number of pixels and the duration is 1.2 seconds, it is determined as a close instruction, triggering the closing of the virtual character;

[0022] The formula is:

[0023]

[0024] N w represents the range for verifying persistence, which is the total number of video frames corresponding within 1.2 seconds; k represents the index of the video frame; R norm (k) is the proportion of the double-palm contact area of the k-th frame in the previous N w sequence; δ(·) is a binary discrimination function that equals 1 when the condition is met and 0 otherwise; 0.95 is the confidence level.

[0025] Furthermore, the posture of the virtual character is adjusted in real time through the inverse kinematics algorithm to align its visual center of gravity with the user's line of sight direction; including:

[0026] (a) Define the position c v of the visual center of gravity of the virtual character under the joint angle vector θ as c v (θ), and set the target position where O is the origin of the palm coordinate system, is the user's line of sight direction vector, and h is the visual distance adjustment factor;

[0027] (b) Construct an objective function to minimize the position error and the posture naturalness constraint:

[0028]

[0029] where θ neutral represents the neutral position of the joint angles of the virtual character, and λ is the posture naturalness constraint weight;

[0030] (c) Iteratively solve for the joint angles through the pseudo-inverse method of the Jacobian matrix:

[0031]

[0032] where θ k is the joint angle vector of the k-th frame, c v (θ k ) is the position of the visual center of gravity c v of the virtual character under the joint angle vector θ k , J is the Jacobian matrix of the position of the visual center of gravity c v with respect to the joint angle vector θ, is the pseudo-inverse of the Jacobian matrix, and α is the iteration step size;

[0033] (d) Set the iteration termination condition to any of the following conditions:

[0034] Δc = c v - c target , and the L2 norm of Δc ≤ 0.5 mm;

[0035] The number of iterations exceeds 10 times;

[0036] The joint angle exceeds the preset biomechanical limit range.

[0037] Furthermore, the posture of the virtual character is adjusted in real time through the inverse kinematics algorithm to align its visual center of gravity with the user's line of sight direction; it also includes: compensating and calculating the binding error; specifically as follows:

[0038] a) Establish a spatial residual model to calculate the position deviation and normal vector angle deviation of the root node of the virtual character in the local coordinate system:

[0039]

[0040] Among them, O v (t) - O represents the position deviation of the root node of the virtual character in the local coordinate system, and O v (t) represents the real-time coordinates of the root node of the virtual character, and O represents the origin, that is, the binding target coordinates; represents the normal vector angle deviation, represents the vector in the Z-axis direction of the virtual character, represents the initial palm normal vector;

[0041] b) Based on the adaptive PID control law, dynamically compensate for the position and angle deviations, and the control law expression is:

[0042]

[0043] e(t) represents the error vector; K p represents the proportional gain diagonal matrix; K i represents the integral gain, and the integral window τ = 0.1s; represents the transformation matrix from the world coordinate system to the local coordinate system; the integral term represents integrating the error history of the world coordinate system after projecting it into the local coordinate system through the transformation matrix t represents the current time, and ξ represents the integration variable; the differential term represents a predictive supplement to the error change rate in the local coordinate system to suppress jitter;

[0044] c) According to the control output Δq(t), adjust the position and posture of the root node of the virtual character in real time until the following convergence conditions are met:

[0045] The L2 norm of Ox(t) - O ≤ 0.2mm, and The absolute value of ≤ 0.5°.

[0046] Furthermore, combine the eye movement tracking data with the physiological signals to dynamically adjust the rendering accuracy and interaction behavior of the virtual character to respond to the user's emotional state; including:

[0047] (1) The eye movement tracking module acquires the change rate of the user's pupil diameter ΔPupil at a sampling rate of not less than 120 Hz, and the PPG sensor collects the heart rate HR in real time.

[0048] (2) Calculate the emotional index EI based on the following formula to quantify the user's emotional state:

[0049]

[0050] Where, ΔPupil represents the change rate of the pupil diameter, HR represents the real-time heart rate, Pupil base represents the reference pupil diameter of the user at rest, HR rest represents the user's resting heart rate, HR max represents the user's maximum heart rate;

[0051] (3) Dynamically adjust the rendering parameters and interaction behaviors according to the emotional index range:

[0052] When EI < 0.4, reduce the virtual character material resolution to 512×512, and set the skeletal animation frame rate to 30 fps;

[0053] When 0.4 ≤ EI < 0.7, enable 2K resolution materials and increase the skeletal animation frame rate to 60 fps;

[0054] When EI ≥ 0.7, activate the subsurface scattering technology to simulate the skin texture, increase the material resolution to 8K, set the skeletal animation frame rate to 90 fps, and trigger the active interaction actions of the virtual character;

[0055] (4) If it is detected that the user has not gazed at the virtual character for 3 consecutive seconds, automatically pause the calculation of the emotional index and enter the low-power rendering mode.

[0056] Furthermore, based on the local coordinate system, determine the corresponding action determination conditions for calling out and closing the virtual character; it also includes an accidental touch protection mechanism:

[0057] Only respond to the instruction when the palm starts to flip from the ergonomic natural position;

[0058] Set a cooling period for continuous operations. If the interval between two call-out instructions is less than 1 second, ignore subsequent instructions.

[0059] In a second aspect, an interactive museum handheld intelligent guide system provided by an embodiment of the present invention adopts the interactive museum handheld intelligent guide method described in any item of the first aspect. The system includes:

[0060] The palm motion capture module is used to obtain the three-dimensional point cloud data of the user's palm in real time through the camera of the AR device, establish a local coordinate system with the geometric center of the palm as the origin, and calculate the rotation angle of the palm normal vector and the contact area of both palms;

[0061] Avatar binding module, used to dynamically bind the avatar to the local coordinate system and adjust its posture through the inverse kinematics algorithm so that its visual center of gravity is aligned with the user's line of sight;

[0062] The gesture-driven control module activates particle special effect rendering when responding to the call-out command, and triggers the space collapse animation when responding to the close command.

[0063] Furthermore, it also includes:

[0064] An emotional interaction module, based on the eye tracker and biosensor integrated in the AR device, for identifying the user's emotional state and adjusting the rendering parameters of the virtual character;

[0065] The realistic rendering engine uses subsurface scattering technology to simulate the skin texture of virtual characters and realizes ambient light adaptation based on the local lighting model.

[0066] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:

[0067] (1) Optimize the convenience of interaction

[0068] The existing technology lacks a convenient interaction method, and users cannot flexibly control the appearance and hiding of the guide according to their own needs, resulting in insufficient autonomy and fluency in the viewing experience. The present invention realizes the convenient calling out and closing of the virtual guide by recognizing palm movements (such as calling out the virtual character with the palm facing up and closing the virtual character with the palm closed). Users can switch between the self-viewing mode and the guided explanation mode anytime and anywhere, thereby improving the naturalness and convenience of interaction.

[0069] (2) Improve the vividness of user experience

[0070] The existing technology mainly displays information through text and images, which is relatively simple in form and difficult to attract tourists' attention, and cannot meet tourists' needs for vivid and interesting experiences. The present invention introduces virtual characters as guides to provide tourists with explanation services in a more vivid and vivid way, thereby enhancing users' sense of participation and immersion.

[0071] (3) Enhance the realism of virtual guides

[0072] The prior art lacks vividness in the presentation of virtual characters and the support of realistic rendering technology, resulting in a visual effect of the virtual guide that makes it difficult for users to obtain an immersive experience. The present invention adopts realistic rendering technology to make the virtual guide visually closer to real people, further enhancing the user's immersion and acceptance of the virtual guide.

[0073] (4) Improve the intelligence of emotional interaction

[0074] The prior art fails to fully consider the emotional needs of users and lacks emotional interaction functions. The present invention identifies the emotional state by observing the user's eyes, enabling the virtual guide to make corresponding actions according to the user's mood. For example, it can provide more detailed explanations when the user is confused or give short rest suggestions when the user is tired, thereby enhancing the intelligence and personalization of the tour guide service.

[0075] (5) Improve the flexibility and autonomy of the tour guide

[0076] The prior art lacks flexibility during the tour guide process. Users often need to rely on the information provided by the system throughout the process and cannot selectively learn according to their own interests and rhythms. The present invention uses the method of a palm-held virtual guide, enabling users to call out or close the guide at any time, independently choose whether to require the tour guide service, and at the same time can select specific exhibits or areas according to their own interests for in-depth understanding, improving the flexibility and autonomy of the tour guide.

[0077] In summary, the present invention aims to solve the deficiencies of the existing intelligent guide system for cultural activities in terms of user experience, interaction convenience, visual effect, and emotional interaction through innovative interaction methods, realistic rendering technology, and emotional interaction functions, providing users with a more vivid, convenient, intelligent, and personalized tour guide service. Description of the Drawings

[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0079] Figure 1 It is a flowchart of the interactive museum palm-held intelligent guide method provided by the present invention.

[0080] Figure 2 It is a block diagram of the interactive museum palm-held intelligent guide system provided by the present invention.

[0081] Figure 3 It is a core interaction flowchart of the intelligent guide system provided by the present invention.

[0082] Figure 4 It is the flow chart of the emotion-assisted optimization type of the intelligent guidance system provided by the present invention. Specific implementation manners

[0083] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0084] Referring to Figure 1 As shown, an interactive museum handheld intelligent guidance method disclosed in an embodiment of the present invention includes the following steps S1 to S5:

[0085] S1. Real-time obtain the three-dimensional point cloud data of the user's palm through the camera of the AR device, and establish a local coordinate system with the geometric center of the palm as the origin.

[0086] For example, the three-dimensional point cloud data of the user's palm is obtained in real time through the ToF depth camera (resolution 640×480, frame rate 60fps) of the user wearing the AR glasses.

[0087] Define the user's right hand as the reference hand (the left hand is processed by mirror symmetry), and establish the local coordinate system of the palm:

[0088] Origin O: The geometric center of the palm depression (determined by point cloud curvature analysis, curvature threshold κ≥0.15mm -1 );

[0089] Z axis : The initial normal vector of the palm (the reference direction is the opposite direction of gravity );

[0090] X axis Along the direction of the index finger (the first axis of the principal component analysis PCA of the point cloud);

[0091] Y axis Complete the construction of the right-handed system.

[0092] S2. Based on the local coordinate system, determine the corresponding action determination conditions for calling out and closing the virtual character;

[0093] In the local coordinate system of the palm, determine the gesture actions and other constraint conditions such as related time corresponding to the appearance and closing of the virtual character. The gesture actions include, for example: palm rotation, palms together, fingers spread, fist clenching, etc.

[0094] In addition, an accidental touch protection mechanism is also set up, including: taking the rotation of the palm as an example, only responding to the instruction when the palm starts to flip from the ergonomic natural position; further, a cooling period for continuous operation can also be set. If the interval between two call-out instructions is less than 1 second, subsequent instructions will be ignored.

[0095] S3. When the rotation angle of the palm normal vector meets the preset threshold range and lasts for the first preset duration, it is determined as a call-out instruction, the virtual character bound to the local coordinate system is activated, and is dynamically rendered and displayed in the AR device;

[0096] Taking the calculated rotation angle of the palm normal vector (determination threshold: 85°±3°, duration 0.3 - 0.5 seconds) as the call-out instruction;

[0097] Condition 1: Normal vector rotation angle determination

[0098] Calculate the palm normal vector in real time and the initial normal vector of the included angle:

[0099]

[0100] Angle threshold function:

[0101]

[0102] Condition 2: Movement trajectory constraint

[0103] The movement trajectory of the palm needs to meet the smoothness condition (excluding jitter interference):

[0104]

[0105] represents the real-time position of the palm origin in the local coordinate system. The integrand is the L2 norm of the acceleration vector of the palm movement. Δt is the sliding time window, and t0 represents the time point when the acceleration integration starts, that is: the initial moment when the user starts to move the palm, which can be expressed as the time when the palm starts to flip.

[0106] Constraint connotation:

[0107] Integration upper limit 15mm / s 2 : The maximum allowable acceleration accumulation (equivalent to the palm accelerating from rest to a speed of 4.5mm / s in 0.3 seconds).

[0108] Excluded scenarios:

[0109] Sudden jitter (such as hand muscle tremors, instantaneous acceleration > 50mm / s 2 )

[0110] Unintentional rapid movement (such as a sudden movement to avoid an obstacle).

[0111] Final determination:

[0112]

[0113] Among them, N = 18 frames, representing the total number of video frames; k represents the video frame index; Δt = 0.3 s. Verify the temporal stability of the rotation angle through a sliding window (N = 18 frames, Δt = 0.3 s). θ(t k ) represents the rotation angle of the palm normal vector of the k-th frame in the sequence of the first 18 frames.

[0114] S4. Adjust the posture of the virtual character in real time through the inverse kinematics algorithm so that its visual center of gravity is aligned with the user's line of sight direction;

[0115] Bind the root bone node of the virtual character to the origin of the palm coordinate system, and compensate the spatial mapping error to ≤2 mm; Iteratively solve the joint angles based on the pseudo-inverse method of the Jacobian matrix until the visual center of gravity is aligned with the target position (error ≤0.5 mm). The goal is: adjust the joint angle vector θ ∈ Rn of the virtual character so that its visual center of gravity c v is aligned with the target position c target in the local coordinate system of the palm. The specific process includes:

[0116] (a) Define the position c v of the visual center of gravity of the virtual character under the joint angle vector θ v (θ), and set the target position where O is the origin of the palm coordinate system, is the user's line of sight direction vector, and h is the visual distance adjustment factor;

[0117] (b) Construct an objective function to minimize the position error and the naturalness constraint of the posture, so that the virtual character's visual c v is aligned with the user's line of sight direction vector :

[0118]

[0119] Among them, θ neutral represents the neutral position of the joint angle of the virtual character, λ is the weight of the naturalness constraint of the posture (default value 0.1); h is the visual distance adjustment factor (default value 0.3 m);

[0120] (c) Iteratively solve the joint angles through the pseudo-inverse method of the Jacobian matrix:

[0121]

[0122] Among them, θ kis the joint angle vector of the k-th frame (the dimension is the number of virtual character joints), c v (θ k ) is the visual center of gravity c of the virtual character v at the position under the joint angle vector θ k , J is the Jacobian matrix of the visual center of gravity c v position with respect to the joint angle vector θ, is the pseudoinverse of the Jacobian matrix, α is the iteration step (learning rate);

[0123] The Jacobian matrix is constructed as follows:

[0124]

[0125] Pseudoinverse operation:

[0126]

[0127] The damping coefficient D = 0.01, I is the identity matrix: to avoid numerical instability in singular configurations

[0128] (d) Single iteration step

[0129] Calculate the current visual center of gravity c v (θ k )(forward kinematics);

[0130] Construct the Jacobian matrix J;

[0131] Solve for the displacement error Δc = c v - c target ;

[0132] Update the joint angles:

[0133] Termination conditions:

[0134] Success condition: The L2 norm of Δc ≤ 0.5 mm

[0135] Failure condition: The number of iterations > 10 or the joint angles exceed the biomechanical limits.

[0136] Furthermore, it also includes: compensating and calculating the binding error; specifically as follows:

[0137] a) Establish a spatial residual model, and calculate the position deviation and normal vector angle deviation of the virtual character's root node in the local coordinate system:

[0138]

[0139] where, O v (t) - O represents the position deviation of the virtual character's root node in the local coordinate system, O v (t) ∈ R3 Indicates the real-time coordinates of the root node of the virtual character in the local coordinate system of the palm center. O represents the origin, i.e., the bound target coordinates; Indicates the angular deviation of the normal vector, Indicates the vector in the Z-axis direction of the virtual character, Indicates the initial normal vector of the palm center; that is: indicates the Z-axis of the virtual character and the normal vector of the palm center of the included angle.

[0140] b) Based on the adaptive PID control law (coordinate system normalization), dynamically compensate for the position and angle deviations. The expression of the control law is:

[0141]

[0142] where e(t) represents the error vector; the proportional term K p e(t): The immediate compensation based on the current error, quickly reducing the position deviation; K p represents the proportional gain diagonal matrix, and its components are independently adjusted according to the error direction (position / angle). Example value: K p = diag(1.2, 1.2, 1.2, 0.6) corresponding to the X / Y / Z translation and normal vector rotation errors of the local coordinate system of the palm center. K i represents the integral gain, the integral window τ = 0.1s, balancing real-time performance and anti-noise ability; represents the transformation matrix from the world coordinate system to the local coordinate system; the integral term represents integrating the error history of the world coordinate system projected onto the local coordinate system through the transformation matrix , t represents the current time, ξ represents the integration variable, eliminating the cumulative error caused by sensor drift or environmental interference. The differential term represents the predictive supplement to the rate of change of the error in the local coordinate system, suppressing jitter;

[0143] c) According to the control output Δq(t), adjust the position and posture of the root node of the virtual character in real time until the following convergence conditions are met:

[0144] O v (t) - The L2 norm of O ≤ 0.2mm, and the absolute value of ≤ 0.5°.

[0145] S5. Detect the contact area of the user's two palms. When the proportion of the contact area exceeds the set threshold and lasts for the second preset duration, or the waving speed of a single palm is greater than the set value, it is determined as a close command, triggering the disappearance animation of the virtual character.

[0146] For example, detect the contact area of the user's two palms. When the proportion of the contact area exceeds the set threshold (≥78% of the total number of pixels) and lasts for a specified duration (1.2 seconds), or the waving speed of a single palm is greater than 1.5 m / s, it is determined as a close command, triggering the disappearance animation of the virtual character. The disappearance animation includes:

[0147] Space collapse effect (the number of model patches decreases from 500,000 to 100 frame by frame, taking 0.2 seconds);

[0148] Light point fading particle special effect (the number of particles is 500 - 1000, and the life cycle is 0.3 seconds).

[0149] The specific process of detecting the contact area of the user's two palms is as follows:

[0150] 1) Extract the contact area of the two palms and perform binary processing on the depth image:

[0151]

[0152] Among them, z hand is the reference depth of the palm, with a value between 20 cm and 80 cm, and σ z = 3 cm is the noise tolerance; a binary image is obtained for each frame, where 1 represents that the ray emitted along this pixel into the scene can hit the user's palm, and 0 means it doesn't. For each frame, the above binary images are generated for the left and right hands respectively. The number of pixels with a value of 1 is calculated as the number of pixels of the left and right hands, and then the number of pixel coordinates from which the rays can hit both the left and right hands simultaneously is calculated as the number of pixels in the overlapping area of the two palms.

[0153] 2) Calculate the proportion of the contact area:

[0154]

[0155] Among them, represents the number of pixels in the overlapping area of the two palms; respectively represent the total number of pixels of the left and right palms;

[0156] 3) Perform spatio-temporal persistence verification:

[0157] Adopt sliding window verification (window length T w = 1.2 s, confidence level p≥95%)

[0158]

[0159] Among them, N wIndicates the range for verifying persistence, the total number of video frames corresponding within 1.2 seconds. The parameter design basis: The minimum duration of a conscious human gesture is approximately 1.0 - 1.5 seconds (psychological research data). At 60fps sampling, 72 frames = 1.2 seconds, covering a complete gesture cycle; k represents the index of the video frame; For the first N w In the sequence, the ratio of the double-palm contact area of the k-th frame; δ(·) is a binary discrimination function, which is equal to 1 when the condition is met, otherwise equal to 0; 0.95 is the confidence level;

[0160] 4) Dynamic compensation mechanism: Used to eliminate individual palm size differences (such as the absolute area difference between an adult's palm and a child's palm). Only need to replace R in step 3) contact with R norm That's all.

[0161] Palm size adaptive calibration:

[0162] When the user uses it for the first time, perform a reference measurement:

[0163]

[0164] Real-time contact area normalization processing:

[0165]

[0166] Among them, A pixel is the actual area of a single pixel (calculated according to the depth value), and A base is the palm area measured as a reference when the user uses it for the first time.

[0167] Meaning of normalization:

[0168] When R norm ≥ 78% is determined as a valid closing instruction (independent of the absolute size of the palm).

[0169] Final determination:

[0170]

[0171]

[0172] In this step S5, when the palm flips to the summoned state, activate the particle effect rendering pipeline (500 - 1000 particles, fade-in time 0.3 seconds); when the double-palm contact reaches the closing condition, trigger the space collapse animation (the virtual character shrinks to a 5mm diameter light spot and then disappears, time-consuming 0.2 seconds).

[0173] Furthermore, as shown in Figure 1 Before step S5 is executed, this method may further include:

[0174] S6. Combine eye-tracking data with physiological signals to dynamically adjust the rendering accuracy and interaction behavior of virtual characters in response to the user's emotional state. This step includes two aspects: emotional-assisted decision-making optimization and realistic rendering optimization. The specific process is as follows:

[0175] (1) Obtain the user's pupil diameter change rate ΔPupil through the eye-tracking module at a sampling rate of no less than 120 Hz, and collect the heart rate HR in real time through the PPG sensor;

[0176] During the invocation of the virtual character, detect the user's pupil diameter change rate (such as a sampling rate of 120 Hz) through the eye-tracking module integrated in the AR device; when it is detected that the user continuously gazes at the virtual character for more than 3 seconds (determined as a high-interest state), automatically activate detailed enhancement rendering (for example, the material resolution is increased to 8K, and the skeletal animation frame rate is doubled).

[0177] (2) Calculate the emotional index EI based on the following formula to quantify the user's emotional state:

[0178]

[0179] where ΔPupil represents the pupil diameter change rate, HR represents the real-time heart rate, Pupil base represents the reference pupil diameter of the user at rest, HR rest represents the user's resting heart rate, HR max represents the user's maximum heart rate;

[0180] Realistic rendering optimization:

[0181] Construct a local lighting model in the palm local coordinate system, and collect the ambient light direction and intensity in real time;

[0182] Define a dynamic light transport equation based on the palm local coordinate system:

[0183]

[0184] where the point x is the point to be calculated, L o is the total intensity of the light emitted along the direction at the point x, L e is the self-emission at the point x, f r is the BSSRDF equation, indicating how much proportion of the light incident from the direction is scattered through the subsurface and then emitted along the direction at the point x, L i is the incident light, and the final cosine term, represents the normal vector at the point x.

[0185] The sub-surface scattering technology (SSS) is adopted to simulate the skin texture of virtual characters, making it consistent with the lighting effect of real palms.

[0186] Diffusion profile equation (indicating the light received around a point light source when a point light source is embedded in an infinitely large homogeneous participating medium):

[0187]

[0188] σ a : Absorption coefficient of the skin material of the virtual character, which can be set based on the absorption coefficient of human skin;

[0189] D: Diffusion constant; r is the distance to the point light; α is the brightness of the light source;

[0190] More specifically, the SSSS skin rendering technology can be adopted.

[0191] (3) Dynamically adjust the rendering parameters and interaction behaviors according to the emotional index range:

[0192] When EI < 0.4, reduce the material resolution of the virtual character to 512×512, and set the bone animation frame rate to 30fps;

[0193] When 0.4 ≤ EI < 0.7, enable the 2K resolution material, and increase the bone animation frame rate to 60fps;

[0194] When EI ≥ 0.7, activate the sub-surface scattering technology to simulate the skin texture, increase the material resolution to 8K, set the bone animation frame rate to 90fps, and trigger the active interaction actions of the virtual character;

[0195] (4) If it is detected that the user has not gazed at the virtual character for 3 consecutive seconds, automatically pause the calculation of the emotional index and enter the low-power rendering mode.

[0196] The serial numbers of the above steps S1 to S6 do not refer to the execution order, but are only for the convenience of describing the solution.

[0197] The interactive museum handheld intelligent guide method provided by the present invention can realize seamless switching between users' independent viewing and guided explanation. The presentation position error of the handheld virtual character ≤ 2mm, achieving millimeter-level spatial fusion; the opening and closing animation takes ≤ 0.5 seconds, conforming to the law of human instantaneous attention transfer; significantly improving the user experience. In addition, the dual-threshold determination (rotation angle + contact area) reduces the false trigger rate, improves the operation accuracy; also reduces the command response delay, realizing the improvement of system performance. And the local rendering technology can further reduce the GPU load, and the cooling period mechanism is adopted to reduce ineffective calculations, further reducing the power consumption.

[0198] The present invention is mainly used in cultural exhibition places such as museums, science and technology museums, and art galleries to provide visitors with personalized and immersive guided tour services. By wearing AR glasses, visitors can call out the virtual guide on their palms at any time through simple palm movements (such as calling out the virtual character with the palm facing up, and closing the virtual character with the palm closed) to obtain exhibit information, historical background and other knowledge explanations. This palm-calling and closing method greatly improves the user's operating convenience. Users can easily control the appearance and disappearance of the virtual guide without complex voice commands or touch operations, so as to focus more on visiting and learning the exhibits. At the same time, the virtual guide can make corresponding actions according to the emotional state of the visitor, further enhancing the fun and interactivity of the visit. In addition, through realistic rendering technology, the virtual guide can be presented with more realistic visual effects, further enhancing the user's sense of immersion.

[0199] Potential application areas also include tourist attractions such as theme parks and historical sites, as well as scenarios such as corporate exhibition halls and school education. Through customized content and interactive methods, targeted guided tours and learning experiences can be provided for different user groups.

[0200] Based on the same inventive concept, the present invention also provides an interactive museum handheld intelligent guide system, which adopts the interactive museum handheld intelligent guide method as the above embodiment. Since the principle of solving the problem by this system is similar to that of the above interactive museum handheld intelligent guide method, the implementation of this system can refer to the implementation of the above method, and the repeated parts will not be repeated.

[0201] Reference Figure 2 As shown, the system includes:

[0202] The palm motion capture module is used to obtain the three-dimensional point cloud data of the user's palm in real time through the camera of the AR device, establish a local coordinate system with the geometric center of the palm as the origin, and calculate the rotation angle of the palm normal vector and the contact area of both palms;

[0203] Avatar binding module, used to dynamically bind the avatar to the local coordinate system and adjust its posture through the inverse kinematics algorithm so that its visual center of gravity is aligned with the user's line of sight;

[0204] The gesture-driven control module activates particle special effects rendering when responding to the call-out command, and triggers the space collapse animation when responding to the close command;

[0205] An emotional interaction module, based on the eye tracker and biosensor integrated in the AR device, for identifying the user's emotional state and adjusting the rendering parameters of the virtual character;

[0206] The realistic rendering engine uses subsurface scattering technology to simulate the skin texture of virtual characters and realizes ambient light adaptation based on the local lighting model.

[0207] Example 1:

[0208] 1. The hardware components are as follows

[0209] (1) AR glasses main body

[0210] Housing material: magnesium alloy (thickness 1.2 mm, weight ≤ 80 g);

[0211] Optical module interface: MIPI CSI-2 dual-channel (bandwidth 4 Gbps).

[0212] (2) Depth vision module

[0213] ToF depth camera: Intel RealSense D455 (resolution 1280×720, frame rate 60 fps);

[0214] Infrared fill light: wavelength 850 nm (radiation power 1.2 mW / cm 2 )

[0215] (3) Gesture recognition processor

[0216] Core chip: NVIDIA Jetson Nano (computing power 472 GFLOPS);

[0217] Memory configuration: LPDDR4 4GB (bandwidth 25.6 GB / s).

[0218] (4) Display and feedback unit

[0219] Waveguide film: thickness 3 mm, refractive index 1.8;

[0220] Micro vibration motor: diameter 4 mm (response frequency 50 - 200 Hz).

[0221] 2. The core interaction process is as shown in Figure 3 shown

[0222] Step 101. Palm space positioning: Capture the three-dimensional point cloud of the palm through the ToF camera (point distance accuracy ±1 mm) and calculate the origin of the palm coordinate system (the geometric center point of the palm, algorithm error ≤ 0.5 mm)

[0223] Step 102: Action determination: Activation conditions (both conditions must be met):

[0224] Rotation angle θ of the palm normal vector: 82° ≤ θ ≤ 88° (threshold range)

[0225] Action duration t: 0.3 s ≤ t ≤ 0.5 s

[0226] Deactivation conditions (either condition is met):

[0227] The contact area ratio of both palms ≥ 78% and lasts for 1.2 s; or a single palm is quickly waved (linear velocity ≥ 1.5 m / s).

[0228] Step 103: Virtual character binding

[0229] Spatial mapping: Bind the root bone node of the virtual character to the origin of the palm coordinate system (error compensation algorithm);

[0230] Pose adjustment: Based on inverse kinematics (IK), calculate the joint angles of the character in real time so that the included angle between the line of sight direction of the character and the optical axis of the user's glasses ≤ 5°.

[0231] Step 104: Dynamic rendering

[0232] Fade-in animation: Simulate the process of the character emerging through a particle system (number of particles 800 ± 200, life cycle 0.3 s).

[0233] Collapse animation: The character model gradually decreases from 50,000 faces to 100 faces frame by frame, taking 0.2 s.

[0234] 3. Key technical parameters

[0235] Module Parameter Name Value Range Depth Vision Ranging Accuracy ±1mm@0.5m Gesture Recognition False Positive Rate ≤2%(Confidence Threshold 0.92) Rendering Engine Frame Rate Stability 90±5fps Power Consumption Control Continuous Working Duration ≥3h@2000mAh

[0236] Example 2:

[0237] On the basis of Example 1, add emotional assistance optimization, as shown in Figure 4 shown, is a schematic diagram of eye movement-rendering linkage, applicable to high-immersion scenarios.

[0238] 1. Hardware enhancement configuration

[0239] Eye movement tracking module: Tobii Pro Nano (sampling rate 120 Hz, accuracy 0.3°)

[0240] Biosensor: PPG heart rate module (accuracy ±2 bpm)

[0241] 2. Emotion-driven rendering process

[0242] Step 201: Physiological signal fusion

[0243] Calculate the emotion index EI:

[0244]

[0245] where ΔPupil is the change rate of pupil diameter and HR is the real-time heart rate.

[0246] Step 202: Dynamic adjustment of rendering

[0247] Low emotional state (EI < 0.4):

[0248] The texture map is downgraded to a resolution of 512×512;

[0249] The frame rate of the skeletal animation is reduced to 30 fps;

[0250] High emotional state (EI ≥ 0.7):

[0251] Activate the subsurface scattering (SSS) material (thickness 0.2 mm, scattering coefficient 0.8);

[0252] Enhance the eye specular map (increase the specular reflection intensity to 2.0).

[0253] 3. Performance comparison data

[0254] Emotional State Rendering Resolution Power Consumption Frame Rate Low (EI < 0.4) 1080p 3.2W 45fps Medium (0.4 ≤ EI < 0.7) 2K 4.1W 60fps High (EI≥0.7) 4K 5.8W 90fps

[0255] In addition, for the interactive museum handheld intelligent guide system provided by the present invention, a mobile terminal with AR function can also be selected. The mobile phone RGB camera (1080P@30fps) is used to replace the ToF camera, and the palm posture is estimated through key point detection (MediaPipeHands model). For example, the calling condition is changed to a five-finger open gesture (confidence ≥ 0.85); the closing condition is changed to a fist gesture (duration ≥ 0.8 s).

[0256] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0257] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An interactive museum handheld intelligent guide method, characterized in that, It includes the following steps: Obtain the three-dimensional point cloud data of the user's palm in real time through the camera of the AR device, and establish a local coordinate system with the geometric center of the palm as the origin; Based on the local coordinate system, determine the corresponding action determination conditions for calling and closing the virtual character; When the rotation angle of the palm normal vector satisfies the preset threshold range and lasts for the first preset duration, it is determined as a call instruction, and the virtual character is activated and bound to the local coordinate system, and is dynamically rendered and displayed in the AR device; Real-time adjust the posture of the virtual character through the inverse kinematics algorithm to align its visual center of gravity with the user's line of sight; Detect the contact area of the user's two palms. When the proportion of the contact area exceeds the set threshold and lasts for the second preset duration, it is determined as a close instruction, triggering the disappearance animation of the virtual character.

2. The interactive museum handheld intelligent guide method according to claim 1, wherein, It also includes the following steps: Combine eye tracking data with physiological signals to dynamically adjust the rendering accuracy and interaction behavior of the virtual character to respond to the user's emotional state.

3. The interactive museum handheld intelligent guide method according to claim 1, characterized in that, When the rotation angle of the palm normal vector satisfies the preset threshold range and lasts for the first preset duration, it is determined as a call instruction, including: When the rotation angle θ of the palm normal vector satisfies: 82° ≤ θ ≤ 88°, and the duration is 0.3 seconds to 0.5 seconds, and the movement trajectory meets the constraint conditions, trigger the call of the virtual character; The formula is: where N represents the total number of video frames; k represents the video frame index; θ(t k ) represents the rotation angle of the palm normal vector of the k-th frame; Δt represents the sliding time window, and t0 represents the time point when the acceleration integration starts, that is: the initial moment when the user starts to move the palm; represents the real-time position of the palm origin in the local coordinate system, and the integrand is the L2 norm of the acceleration vector of the palm movement; 15mm / S 2 represents the upper limit of integration, which is the maximum allowable acceleration accumulation.

4. The interactive museum handheld intelligent guide method according to claim 1, characterized in that, Detect the contact area of the user's two palms. When the proportion of the contact area exceeds the set threshold and lasts for the second preset duration, it is determined as a close instruction; including: When the determination threshold for detecting the contact area of the user's two palms is more than 78% of the total number of pixels and the duration is 1.2 seconds, it is determined as a close instruction, triggering the closing of the virtual character; The formula is: N w represents the range for verifying persistence, which is the total number of video frames corresponding within 1.2 seconds; k represents the index of the video frame; R norm (k) is the proportion of the double-palm contact area of the k-th frame in the previous N w sequence, which is equal to 1 when the binary discrimination function δ(·) satisfies the condition, otherwise it is equal to 0; 0.95 is the confidence level.

5. The interactive museum palm-top intelligent guide method according to claim 1, characterized in that, Real-time adjust the posture of the virtual character through the inverse kinematics algorithm to align its visual center of gravity with the user's line of sight; including: (a) Define the visual center of gravity c of the virtual character v The position c under the joint angle vector θ v (θ), and set the target position Among them, O is the origin of the palm coordinate system, is the user's line-of-sight direction vector, and h is the visual distance adjustment factor; (b) Construct an objective function to minimize the position error and the naturalness constraint of the posture; where θ neutral represents the neutral position of the joint angle of the virtual character, and λ is the weight of the pose naturalness constraint; (c) Iteratively solve the joint angles through the pseudo-inverse method of the Jacobian matrix; Among them, θ k is the joint angle vector of the k-th frame, and c v (θ k ) is the visual center of gravity c v of the virtual character at the joint angle vector θ k , J is the Jacobian matrix of the visual center of gravity c v position with respect to the joint angle vector θ, is the pseudo-inverse of the Jacobian matrix, and α is the iteration step size; (d) Set the iteration termination condition to any of the following conditions: Δc = c v -c target , the L2 norm of Δc ≤ 0.5 mm; The number of iterations exceeds 10 times; The joint angles exceed the preset biomechanical limit range.

6. The interactive museum handheld intelligent guide method according to claim 5, characterized in that Real-time adjust the posture of the virtual character through the inverse kinematics algorithm to align its visual center of gravity with the user's line of sight; it also includes: compensating and calculating the binding error; specifically as follows: a) Establish a spatial residual model, and calculate the position deviation and normal vector angle deviation of the root node of the virtual character in the local coordinate system; Among them, O v (t)-O represents the position deviation of the root node of the virtual character in the local coordinate system, and O v (t) represents the real-time coordinates of the root node of the virtual character, and O represents the origin, that is, the bound target coordinates; represents the normal vector angle deviation, represents the virtual character's Z-axis direction vector, represents the initial palm normal vector; b) Dynamically compensate the position and angle deviations based on the adaptive PID control law, and the expression of the control law is: e(t) represents the error vector; K p represents the proportional gain diagonal matrix; K i represents the integral gain, and the integration window τ = 0.1 s; represents the transformation matrix from the world coordinate system to the local coordinate system; the integral term represents integrating the error history in the world coordinate system after projecting it into the local coordinate system through the transformation matrix where t represents the current time and ξ represents the integration variable; the differential term represents making a predictive supplement to the error change rate in the local coordinate system to suppress jitter; c) According to the control output Δq(t), real-time adjust the position and posture of the root node of the virtual character until the following convergence conditions are met: O v (t) - The L2 norm of O ≤ 0.2 mm, and The absolute value of ≤ 0.5°.

7. The method for an interactive museum handheld intelligent guide according to claim 2, wherein, Combine eye tracking data with physiological signals to dynamically adjust the rendering accuracy and interaction behavior of the virtual character to respond to the user's emotional state; including: (1) Obtain the change rate of the user's pupil diameter ΔPupil through the eye tracking module at a sampling rate of not less than 120Hz, and collect the heart rate HR in real time through the PPG sensor; (2) Calculate the emotional index EI based on the following formula to quantify the user's emotional state: where ΔPupil represents the pupil diameter change rate, HR represents the real-time heart rate, Pupil base represents the reference pupil diameter of the user in the resting state, HR rest represents the resting heart rate of the user, HR max represents the maximum heart rate of the user; (3) Dynamically adjust the rendering parameters and interaction behavior according to the emotional index interval: When EI < 0.4, reduce the virtual character material resolution to 512×512 and set the skeletal animation frame rate to 30fps; When 0.4 ≤ EI < 0.7, enable 2K resolution materials and increase the skeletal animation frame rate to 60fps; When EI ≥ 0.7, activate the subsurface scattering technology to simulate skin texture, increase the material resolution to 8K, set the skeletal animation frame rate to 90fps, and trigger the active interaction actions of the virtual character; (4) If it is detected that the user has not gazed at the virtual character for 3 consecutive seconds, automatically pause the calculation of the emotional index and enter the low-power rendering mode.

8. The interactive museum handheld intelligent guide method according to claim 2, characterized in that, Based on the local coordinate system, determine the corresponding action determination conditions for calling out and closing the virtual character; also include an accidental touch protection mechanism: Respond to the instruction only when the palm starts to flip from the ergonomic natural position; Set a cooling period for continuous operations. If the interval between two call-out instructions is less than 1 second, ignore subsequent instructions.

9. An interactive museum handheld intelligent guide system, characterized in that, Adopt the interactive museum handheld intelligent guide method described in any one of claims 1-8. The system includes: A palm motion capture module for real-time obtaining the three-dimensional point cloud data of the user's palm through the camera of the AR device, establishing a local coordinate system with the geometric center of the palm as the origin, and calculating the rotation angle of the palm normal vector and the contact area of both palms; A virtual character binding module for dynamically binding the virtual character to the local coordinate system and adjusting its posture through the inverse kinematics algorithm to align its visual center of gravity with the user's line of sight; A gesture-driven control module that activates particle effect rendering when responding to a call-out instruction and triggers a space collapse animation when responding to a close instruction.

10. The interactive museum handheld intelligent guide system according to claim 9, characterized in that, Also include: An emotional interaction module based on the eye tracker and biosensor integrated in the AR device for identifying the user's emotional state and adjusting the rendering parameters of the virtual character; A photorealistic rendering engine that uses subsurface scattering technology to simulate the skin texture of the virtual character and realizes ambient light adaptation based on the local illumination model.

Citation Information

Patent Citations

  • Cultural activity intelligent guide system

    CN110009793A