Exhibition hall guide method and device based on multi-modal interaction and showing stand

By acquiring user information through multimodal interaction technology, dynamically matching the tour content and migrating the interactive behavior, the single interaction problem of the traditional exhibition hall tour method is solved, the efficiency and accuracy of the tour are improved, and the user experience is enhanced.

CN120805030APending Publication Date: 2025-10-17SHENHUA BAOSHEN RAILWAY GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510840448.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional digital exhibition hall navigation methods have a single interactive mode, which is difficult to meet the diverse needs of users, and the interaction efficiency is low and the accuracy is insufficient.

Method used

By obtaining the user's multimodal interaction information within a preset distance range of the exhibits, analyzing the user's interaction behavior and depth, pushing matching guide content, and migrating the interaction behavior and depth to the guide devices of adjacent exhibits when the user moves.

Benefits of technology

It improves the interactive efficiency of exhibition hall tours, the accuracy of tour information push and user experience, meets the diverse needs of users and saves interaction time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805030A_ABST
    Figure CN120805030A_ABST
Patent Text Reader

Abstract

The invention relates to an exhibition hall guiding method and device based on multi-modal interaction and a showing stand, and the method comprises the steps: obtaining multi-modal interaction information input by a user for an exhibit under the condition that the user is in a preset distance range of the exhibit, determining an interaction behavior of the user based on the multi-modal interaction information, and displaying the exhibition hall guiding device based on the interaction behavior of the user. The method comprises the steps of determining the interaction depth of a user, pushing guide content matched with the interaction depth, obtaining the movement track of the user, and migrating the interaction behavior and the interaction depth of the user to guide equipment corresponding to an adjacent exhibit under the condition that the movement track represents that the user moves to the adjacent exhibit of the exhibit. By adopting the method, the navigation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-modal interaction, in particular to a hall tour method and device based on multi-modal interaction, a computer device, a computer readable storage medium, a computer program product and a display stand. BACKGROUND

[0002] With the rapid development of information technology, digital exhibition halls have gradually become an important form of display and exhibition. In traditional digital exhibition hall tours, artificial explanation or simple voice devices are mainly relied on. Tour devices are equipped with touch screens or voice modules. For example, when a user actively interacts with a tour device, the tour device can only recognize specific voice or text information, which makes the interaction between the tour device and the user too single and difficult to meet the diversified needs of users. That is, the traditional exhibition tour method has the defects of low user interaction efficiency and insufficient tour accuracy. SUMMARY

[0003] Therefore, it is necessary to provide a hall tour method, device, apparatus, computer device, computer readable storage medium, computer program product and display stand based on multi-modal interaction, which can improve the tour accuracy.

[0004] In a first aspect, the present application provides a hall tour method based on multi-modal interaction, comprising:

[0005] When the user is within a predetermined distance range of an exhibit, multi-modal interaction information input by the user for the exhibit is acquired;

[0006] Based on the multi-modal interaction information, the interaction behavior of the user is determined;

[0007] Based on the interaction behavior of the user, the interaction depth of the user is determined, and tour content matching the interaction depth is pushed;

[0008] The motion trajectory of the user is acquired. When the motion trajectory represents the motion of the user to an adjacent exhibit of the exhibit, the interaction behavior and interaction depth of the user are migrated to a tour device corresponding to the adjacent exhibit.

[0009] In a second aspect, the present application also provides a hall tour device based on multi-modal interaction. The device comprises a control module, and a space perception module, a multi-modal interaction module and a tour display module connected to the control module, respectively;

[0010] The space perception module is configured to collect and send the position information and motion trajectory of the user to the control module;

[0011] The control module is configured to receive the position information, receive multi-modal interaction information input by the user for the exhibit through the multi-modal interaction module in a case where it is determined based on the position information that the user is within a preset distance range of the exhibit, determine an interaction behavior of the user based on the multi-modal interaction information, determine an interaction depth of the user based on the interaction behavior of the user, send a control instruction to the guide display module, the control instruction carrying guide content matched with the interaction depth, and migrate the interaction behavior and the interaction depth of the user to a guide device corresponding to a neighboring exhibit in a case where the motion trajectory represents that the user moves to the neighboring exhibit.

[0012] The guide display module is configured to receive the control instruction and push the guide content matched with the interaction depth.

[0013] In a third aspect, the present application further provides a display stand for carrying the multi-modal interaction-based exhibition hall guide device, the display stand comprising a foldable asymmetric truss, a plurality of hinges and a mobile chassis, the hinges being arranged at each connection point of the asymmetric truss, the asymmetric truss being folded and unfolded through the hinges, and the mobile chassis being arranged at the bottom of the asymmetric truss, and the mobile chassis being provided with a plurality of universal wheel sets at the bottom.

[0014] In a fourth aspect, the present application further provides a multi-modal interaction-based exhibition hall guide device, comprising:

[0015] The data acquisition module is configured to acquire multi-modal interaction information input by the user for the exhibit in a case where the user is within a preset distance range of the exhibit.

[0016] The interaction behavior determination module is configured to determine an interaction behavior of the user based on the multi-modal interaction information.

[0017] The content pushing module is configured to determine an interaction depth of the user based on the interaction behavior of the user and push guide content matched with the interaction depth.

[0018] The content migration module is configured to acquire a motion trajectory of the user and migrate the interaction behavior and the interaction depth of the user to a guide device corresponding to a neighboring exhibit in a case where the motion trajectory represents that the user moves to the neighboring exhibit.

[0019] In a fifth aspect, the present application further provides a computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing steps in the multi-modal interaction-based exhibition hall guide method embodiments when executing the computer program.

[0020] In a sixth aspect, the present application also provides a computer-readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for guiding a visitor in an exhibition hall based on multi-modal interaction.

[0021] In a seventh aspect, the present application also provides a computer program product, comprising a computer program which, when executed by a processor, implements the steps of the above-mentioned method for guiding a visitor in an exhibition hall based on multi-modal interaction.

[0022] The above-mentioned method, system and device for guiding a visitor in an exhibition hall based on multi-modal interaction, computer device, computer-readable storage medium and computer program product are different from the single interaction mode in the traditional scheme. In the case that the user is within the preset distance range of the exhibit, the multi-modal interaction information input by the user is acquired, and the interaction behavior of the user is analyzed based on the multi-modal interaction information, effectively alleviating the problem that the user's demand cannot be fully met due to the single interaction mode of the traditional guiding scheme. Further, the interaction depth of the user is dynamically determined according to the interaction behavior of the user, and the differentiated guiding content is pushed according to the interaction depth of the user, improving the adaptation degree of the guiding content to the user's demand, improving the guiding efficiency, and finally, the motion trajectory of the user is acquired. When it is detected that the user moves to the adjacent exhibit, the content migration between the guiding devices is triggered, the interaction behavior and the interaction depth of the user in the current exhibit are migrated to the guiding device corresponding to the adjacent exhibit, the interaction time of the user with other guiding devices is saved, and the interaction efficiency of the exhibition hall guiding, the accuracy of the guiding information pushing and the guiding experience of the user are effectively improved. The geometric configuration design and foldable property of the asymmetric truss in the above-mentioned display stand make the display stand can be unfolded and folded according to the space layout of the exhibition hall or the size of the exhibit, and the volume of the display stand can be greatly reduced in the folded state, improving the spatial flexibility and spatial adaptability of the display stand. The hinge structure is used at each connection point of the asymmetric truss, which can improve the exhibition efficiency of the display stand. The mobile chassis is further arranged at the bottom of the asymmetric truss, and a plurality of universal wheel sets are arranged at the bottom of the mobile chassis to realize the movement of the display stand. Therefore, when the display stand is used to carry the multi-modal interaction exhibition hall guiding system, the movement property is linked with the multi-modal interaction system, which can improve the guiding coverage and further improve the multi-modal interaction efficiency with the user. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0024] Figure 1An application environment diagram of a multi-modal interaction-based exhibition hall tour method in an embodiment;

[0025] Figure 2 A flowchart of a multi-modal interaction-based exhibition hall tour method in an embodiment;

[0026] Figure 3 A flowchart of a multi-modal interaction-based exhibition hall tour method in another embodiment;

[0027] Figure 4 A flowchart of a multi-modal interaction-based exhibition hall tour method in yet another embodiment;

[0028] Figure 5 A flowchart of a multi-modal interaction-based exhibition hall tour method in a detailed embodiment;

[0029] Figure 6 A structural block diagram of a multi-modal interaction-based exhibition hall tour device in an embodiment;

[0030] Figure 7 A structural block diagram of a multi-modal interaction module in an embodiment;

[0031] Figure 8 A structural block diagram of a showcase in an embodiment;

[0032] Figure 9 A structural block diagram of a multi-modal interaction-based exhibition hall tour apparatus in an embodiment;

[0033] Figure 10 An internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0034] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0035] It should be noted that the terms "first", "second", and the like used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" and any variations thereof used in the present application are intended to cover non-exclusive inclusion. The term "multiple" used in the present application refers to two or more. The term "and / or" used in the present application refers to one of the options or any combination of multiple options.

[0036] The multi-modal interaction-based exhibition hall tour method provided by the embodiments of the present application can be applied to, for example Figure 1The application environment shown. Among them, the interactive terminal 102 communicates with the control module 104 through the network. The data storage system can store the data required by the control module 104 to process. The data storage system can be integrated on the control module 104, or placed on the cloud or other network servers.

[0037] Specifically, the user can input multi-modal interaction information for the exhibit to the control module 104 through the interactive terminal 102 in the case that the user is in the preset distance range of the exhibit. The control module 104 determines the user's interaction behavior based on the multi-modal interaction information, and determines the user's interaction depth based on the user's interaction behavior. Then, the control module 104 can also obtain the motion trajectory of the user collected by the motion sensing device, and in the case that the motion trajectory represents the user's movement to the adjacent exhibit, the user's interaction behavior and interaction depth are migrated to the guide device corresponding to the adjacent exhibit.

[0038] Among them, the interactive terminal 102 can be but not limited to various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner, a smart vehicle device, a projection device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The control module 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0039] In an exemplary embodiment, as Figure 2 shown, a multi-modal interaction-based exhibition hall guide method is provided. Taking the control module 104 in the Figure 1 as an example, the method includes the following steps:

[0040] S100, in the case that the user is in the preset distance range of the exhibit, obtaining the multi-modal interaction information input by the user for the exhibit.

[0041] The preset distance range can be a circular area with the exhibit as the center and a specific radius. For example, the user is positioned by fusing a UWB (Ultra-Wideband) positioning technology and a binocular vision positioning technology, and it is identified whether the user is in the preset distance range of the exhibit. The multi-modal interaction information includes but is not limited to voice signals, touch signals, gesture signals, posture signals, etc. For example, the voice signals are collected by a ring-shaped 6-microphone array, the beamforming directional sound pickup is supported under 85 dB background noise, the gesture signals are collected by pressure grading operation of a capacitive touch screen (supporting 10-point touch) superimposed with a pressure-sensitive film, and the gesture signals are captured by a skeletal tracking camera. Eight kinds of operation gestures can be predefined, such as pinching and rotating a 3D model.

[0042] Exemplarily, the head coordinates of the user are collected by a skeletal tracking technology, and the positioning data obtained by UWB (Ultra-Wideband) base station triangulation is matched with a visual SLAM (Simultaneous Localization and Mapping) point cloud to confirm the audience motion vector. In combination with the head coordinates of the user and the audience motion vector, it is identified whether the user is in the preset distance range of the exhibit. When the user is in the preset distance range of the exhibit, it is confirmed that the user is in an effective interaction range. At this time, the sensors such as the microphone array, the touch screen, and the camera are synchronously activated. The time division multiplexing technology can be used to sample different modal signals to obtain multi-modal interaction information input by the user for the exhibit. In addition, the sampled signals can be preprocessed, for example, environmental noise is removed by a hardware filter and a Kalman filter to obtain multi-modal interaction information with high signal-to-noise ratio.

[0043] S200, determining an interaction behavior of the user based on the multi-modal interaction information.

[0044] The interaction behavior is an operation intention expressed by the user through multi-modal input, including but not limited to query behavior, interest behavior, etc. For example, the user asks a question “What is the historical background of this cultural relic?” by voice or clicks a “detailed introduction” button by touch.

[0045] Exemplarily, the multi-modal interaction information input by the user can be feature extracted. For example, acoustic features in voice information are extracted, keywords are identified, pressure features in touch information are extracted, touch objects are identified, and contour features of gesture information are extracted for similarity calculation with a predefined gesture library. Further, if the user inputs multiple multi-modal interaction information at the same time, a priority arbitration mechanism can be established in advance. For example, the touch operation is automatically paused when the user inputs voice, to prevent false touch. When a “rotation” gesture and an “enlargement” voice instruction are detected at the same time, semantic fusion (rotation + scaling) is performed to determine the fused user behavior.

[0046] S300, determine the interaction depth of the user based on the interaction behavior of the user, and push the guide content matched with the interaction depth.

[0047] The interaction depth is a hierarchical system divided according to the complexity and duration of the user behavior, and can include the following levels: L1 (primary): the user first gazes at the exhibit or touches it for a short time, triggering the push of the basic graphic and text introduction of the exhibit; L2 (intermediate): continuous interaction, triggering the push of the cycle disassembly animation of the exhibit; L3 (advanced): voice asking professional questions or activating expert mode, triggering the push of expert interview video. The guide content is information provided to the user about the exhibit, including but not limited to text, pictures, audio, video and other forms.

[0048] For example, according to the user interaction behavior determined in the above steps, the interaction depth of the user can be graded and determined in combination with the pre-set interaction depth evaluation rules. For example, if the user only gazes at the exhibit, the interaction depth is determined to be L1 level, if the user gazes at the exhibit for more than a certain time, it is determined to be L2 level, and if the user asks questions about the exhibit by voice, the interaction depth is determined to be L3 level. According to different interaction depths, the guide content matched with the interaction depth can be retrieved from the guide content database, for example, for L1 level interaction, push simple text introduction and basic pictures, for L2 level interaction, provide detailed animation explanation, for L3 level interaction, push professional academic materials, video lectures and other content, and display or play through the user terminal device.

[0049] S400, obtain the motion trajectory of the user, and in the case that the motion trajectory represents the motion of the user to the adjacent exhibit of the exhibit, migrate the interaction behavior and the interaction depth of the user to the guide device corresponding to the adjacent exhibit.

[0050] The motion trajectory is the path of position change formed during the movement of the user in the exhibition space, and the position information of the user can be obtained in real time through positioning technology such as UWB positioning to draw the motion trajectory. The adjacent exhibit refers to other exhibits adjacent to the current exhibit in spatial position in the exhibition layout. The guide device refers to a hardware device or software system for providing exhibit guide services to users, such as fixed guide terminals in the exhibition area.

[0051] Exemplarily, the position information of the user terminal device can be continuously collected, and the motion trajectory of the user can be generated according to the collected position data. When the motion trajectory shows that the user is moving towards a neighboring exhibit of the current exhibit, the interaction behavior and the interaction depth information of the user on the current exhibit are packaged and transmitted to the guide device corresponding to the neighboring exhibit, for example, through wireless communication technologies such as Wi-Fi and Bluetooth. After receiving the information, the guide device of the neighboring exhibit can prepare the corresponding guide content in advance according to the interaction behavior and the interaction depth of the user, and directly push the matched guide service when the user reaches the preset distance range of the neighboring exhibit, without the user's further interaction behavior and depth judgment. In addition to the interaction behavior and the interaction depth, the interaction record formed by the current guide device and the user interaction process can also be migrated to the guide device of the neighboring exhibit.

[0052] The above-mentioned exhibition hall guide method based on multi-modal interaction is different from the single interaction mode in the traditional scheme. In the case that the user is in the preset distance range of the exhibit, the multi-modal interaction information input by the user is obtained, and the interaction behavior of the user is analyzed based on the multi-modal interaction information, which effectively alleviates the problem that the user's demand is difficult to be fully met due to the single interaction mode of the traditional guide scheme. Further, the interaction depth of the user is dynamically determined according to the interaction behavior of the user, and the differentiated guide content is pushed according to the interaction depth of the user, which improves the adaptation degree of the guide content to the user's demand, improves the guide efficiency, and finally, the motion trajectory of the user is obtained. When it is detected that the user moves towards a neighboring exhibit, the content migration between guide devices is triggered, the interaction behavior and the interaction depth of the user on the current exhibit are migrated to the guide device corresponding to the neighboring exhibit, the interaction time of the user with other guide devices is saved, and the interaction efficiency of the exhibition hall guide, the accuracy of the guide information pushing, and the guide experience of the user are effectively improved.

[0053] In one embodiment, the multi-modal interaction information includes at least one of voice information, touch information, and gesture information, as shown in Figure 3 S200 includes:

[0054] S210, in the case that the multi-modal interaction information includes voice information and touch information, only based on the voice information, the interaction behavior of the user is determined.

[0055] S220, in the case that the multi-modal interaction information includes voice information and gesture information, the voice information and the gesture information are semantically fused, the fusion result is determined, and based on the fusion result, the interaction behavior of the user is determined.

[0056] The voice information can be collected by a microphone array and converted into a digital signal. The touch information is a physical signal such as pressure, position, sliding track, etc. generated by the user contacting the touch interface by a finger or other object, which can be captured in real time by a pressure sensor array. The gesture information is a visual dynamic feature generated by the user through a body action (such as finger bending, arm swinging), which can be collected by a visual camera.

[0057] For example, when the voice information and the touch information are simultaneously detected, the semantic features in the voice information can be extracted first, and the voice signal can be converted into a text sequence by a voice recognition engine, and then the intent keywords (such as “explain”, “zoom in”, “switch”, etc.) in the text can be analyzed. At this time, the touch information input by the user can not be analyzed. In addition, when the user inputs the voice information to activate the guide system, the touch information input by the user can also be temporarily suspended to prevent accidental touch.

[0058] When the user simultaneously inputs the voice information and the gesture information, the voice information can be converted into a text, and the contour features of the gesture information can be extracted and compared with a predefined gesture library to find a gesture action matching the gesture information. For example, when the voice instruction “zoom in” and the gesture action “rotate” are simultaneously detected, the semantic fusion of the voice information and the gesture information can be performed to determine that the interactive behavior of the user is “zoom in + rotate”.

[0059] In this embodiment, when the user simultaneously inputs the voice information and the touch information, the voice information input by the user is analyzed first, which can reduce accidental touch operation of the user and improve the accuracy of behavior recognition. When the user simultaneously inputs the voice information and the gesture information, the voice information and the gesture information input by the user are fused, which can more accurately identify the interactive behavior of the user.

[0060] In one embodiment, S300 includes: pushing the graphic guide data of the exhibit when the interaction depth is a first interaction depth; pushing the animation guide data of the exhibit when the interaction depth is a second interaction depth; and pushing the expert guide data of the exhibit when the interaction depth is a third interaction depth; wherein the first interaction depth is less than the second interaction depth, and the second interaction depth is less than the third interaction depth.

[0061] The first interaction depth is a shallow interaction stage in which the user has a preliminary interest in the exhibit, and the triggering condition can be that the user first gazes at the exhibit. The second interaction depth is a middle interaction stage in which the user enters information exploration, and the triggering condition can be that the user gazes at the exhibit for a time period longer than a threshold value, for example, longer than 5 seconds. The third interaction depth is a deep interaction stage in which the user conducts professional exploration, and the triggering condition can be that the user asks a voice question.

[0062] Exemplarily, the interaction behavior determined in the above step can be quantified according to an interaction depth grading rule in a preset interaction behavior database. For example, when the user first gazes at the exhibit, it can be considered that the interaction behavior of the user belongs to a first interaction depth, at which time the basic graphic-text information of the exhibit can be called from the graphic-text database, which can include high-definition pictures and textual descriptions of the exhibit. When the user continues to gaze at the exhibit for more than 5 seconds, it can be considered that the interaction behavior of the user belongs to a second interaction depth, at which time the three-dimensional disassembly animation of the exhibit can be called from the animation database, for example, to be played in a cycle of 30 seconds.

[0063] In this embodiment, through hierarchical guidance, different users can be provided with guidance data matched with them, so as to meet different guidance needs of the users, reduce cognitive overload of the users with shallow interaction depth when facing professional data, and meet the demand of the users with deep interaction depth for professional knowledge, thereby improving guidance accuracy and user satisfaction.

[0064] In one embodiment, the method before S100 further includes: sensing motion data of the user, predicting candidate visited exhibits of the user according to the motion data, and caching guidance data of the candidate visited exhibits, and in the case that the user inputs multi-modal interaction information for the candidate visited exhibits, pushing the guidance data of the candidate visited exhibits.

[0065] The motion data refers to dynamic characteristic data such as motion direction and motion speed of the user in the exhibition space. The candidate visited exhibits are next exhibits that the user is likely to go to, which are predicted according to the motion data of the user, and the number can be multiple, for example, 3 candidate visited exhibits.

[0066] Exemplarily, the motion vector of the user can be determined through matching of the UWB positioning data and the visual SLAM point cloud, and then the motion data such as the motion direction and the motion speed of the user can be extracted from the motion vector. Then, the motion data can be preprocessed by Kalman filtering or the like, and a deep learning model or the like is combined to predict the exhibits that the user is likely to visit in a future period of time, for example, according to the hall space topology graph and the motion direction and the motion speed of the user, the spatial angle between the user and each exhibit is determined, and the smaller the spatial angle is, the more likely the user is to visit the exhibit, and then the candidate visiting exhibits of the user are determined. Further, after the candidate visiting exhibits are determined, the control module can pre-cache the guide data of the candidate visiting exhibits, for example, according to the exhibit identifier of the candidate visiting exhibits, the corresponding guide data is called from the cloud server and cached in the data storage system. When the user inputs the multi-modal interaction information for the candidate visiting exhibits, the corresponding cached guide data can be directly pushed to the user, so as to shorten the response time.

[0067] In the embodiment, the pre-caching mechanism can effectively reduce the access delay of the guide data, and when the user performs multi-modal interaction for the candidate visiting exhibits, the interaction demand of the user can be quickly responded, and the guide efficiency and user satisfaction are improved.

[0068] In one embodiment, after the motion data of the user is perceived, the method further includes: determining whether the user enters the exhibit visiting area of the candidate visiting exhibit according to the motion data, and projecting the contour prompt projection of the candidate visiting exhibit in a case that the duration that the user enters the exhibit visiting area of the candidate visiting exhibit is greater than a preset duration threshold.

[0069] The exhibit visiting area is an area with the candidate visiting exhibit as the center and a preset distance as the radius, and can also be a square or an irregular area set according to the hall layout. The preset duration threshold is used to determine whether the user has real interest in the candidate exhibit, and can be set as 1-5 seconds by default, and can also be dynamically adjusted according to the complexity of the exhibit (for example, 5 seconds for large device art and 2 seconds for small cultural relics). The contour prompt projection is a two-dimensional contour spot of the candidate exhibit projected on the ground or wall surface by the AR projection device, and is used to guide the user to pay attention to the position of the exhibit.

[0070] Exemplarily, the user can be positioned by fusing the UWB technology and the binocular vision technology, the user coordinates obtained by the UWB positioning are fused with the scene depth information obtained by the binocular vision positioning to obtain the motion data of the user. For example, based on the UWB positioning result, the error caused by the multipath effect in the UWB positioning is corrected by matching the visual feature points extracted by the binocular vision system. When the binocular vision system detects a stable feature point in the scene, the stable feature point is associated and analyzed with the coordinate information obtained by the UWB positioning. If it is found that there is a deviation between the UWB positioning result and the spatial position presented by the visual feature point, the UWB positioning coordinates are adjusted by a weighted fusion algorithm to eliminate the error caused by the multipath effect, and high-precision fusion positioning of the user is realized. The user can be tracked by the above-mentioned fusion positioning technology to obtain the motion data of the user. When it is recognized according to the motion data that the head coordinates of the user enter the exhibit visiting area and the duration reaches 1.5 seconds, the AR projection is started. The AR projection device corresponding to the exhibit will project the AR content (such as a virtual model of the exhibit, a historical scene restoration, etc.) prepared in advance into the real scene according to the current position and viewing angle information of the user, to provide the user with an augmented reality visiting experience, so that the audience can more intuitively and vividly understand the information related to the exhibit.

[0071] In this embodiment, after the user enters the exhibit visiting area for a certain duration, the user's line of sight is guided by the AR projection without interfering with the user's independent visiting. Compared with the voice prompt and the screen prompt, this AR projection method can reduce the interference on the user and improve the user's satisfaction with the guide.

[0072] In one embodiment, as shown in Figure 4 the method further comprises:

[0073] S510, obtaining gait cycle and gaze hotspot distribution data of the user.

[0074] S520, determining the age of the user according to the gait cycle, and determining the interest data of the user according to the gaze hotspot distribution data.

[0075] S530, generating a user portrait based on the age of the user and the interest data of the user.

[0076] S540, in the case that the motion trajectory represents the movement of the user to the adjacent exhibit, migrating the user portrait to the guide device corresponding to the adjacent exhibit.

[0077] The gait cycle is the time interval from the landing of one side of the foot to the landing of the same side of the foot again when the user is walking, which can be obtained by collecting human motion micro-Doppler signals by a millimeter wave radar. The gaze hotspot distribution data refers to the user's line of sight landing point coordinate sequence collected by the head posture sensor and the eye tracking instrument, and is used to represent the frequency and duration of the user's gaze on different regions of the exhibit within a unit time.

[0078] Exemplarily, electromagnetic waves can be emitted by a millimeter wave radar, and human motion reflection signals can be received, and then micro-Doppler features can be extracted, and the starting point of the user's gait cycle can be identified from the micro-Doppler features, and the gait cycle of the user can be obtained by calculating the time interval of adjacent cycles. For the gaze hotspot distribution data, the pupil position of the user can be collected by the eye tracking instrument, and the user's line of sight direction can be corrected in real time by the head posture sensor, and then the distribution of the user's gaze hotspot on the exhibit can be analyzed according to the user's line of sight direction and pupil position, and the gaze hotspot distribution data of the user can be obtained. The user portrait includes but is not limited to the user's age group, interest label, interaction preference and other data.

[0079] Further, a gait cycle-age mapping database can be established in advance, for example, the gait cycle of a child is about 1.2 Hz, the gait cycle of an adult is about 0.8 Hz, and the gait cycle of an old person is about 0.6 Hz, and then the user's age can be quickly determined according to the collected gait cycle of the user and the above mapping database. The user's gaze hotspot distribution data can represent the frequency and duration of the user's gaze on different regions of the exhibit within a unit time, so according to the frequency and duration of the user's gaze on different regions of the exhibit within a unit time, the user's interest data such as the user's interest region and interest type for the exhibit can be quickly identified. Then, based on the user's age and the user's interest data, the user portrait can be generated, such as grouping the user by age, adding an interest label, predicting the user's guide preference, etc. according to the user's age and the user's interest data, for example, the user portrait of a child user can be "child + interested in bronze ware + prefer cartoon 3D animation guide + prefer slow guide voice", etc. In actual guide, different guide data can be pushed to the user according to the user portrait, for example, cartoon guide interface and slow guide voice can be used for child users. In addition, when the user's motion trajectory is identified to move to the adjacent exhibit of the current exhibit, the user portrait analyzed for the user can be migrated to the guide device of the adjacent exhibit, and after the adjacent exhibit guide device receives the user portrait, content matching the user portrait of the user can be preloaded, such as preloading a cartoon guide interface for a child user, and adjusting the speed of the guide voice to an appropriate speed. If the guide device of the current exhibit does not interact with the user within a certain time, it can switch to a standby state, for example, the power consumption is reduced from 120 watts to 15 watts, waiting for the user to input multi-modal interaction information to activate.

[0080] In this embodiment, by collecting the gait cycle and gaze hotspot distribution data of the user, the accuracy of age prediction can be improved based on the gait cycle, and the gaze hotspot distribution data can better reflect the interest of the user, thereby more accurately constructing the user portrait, accurately pushing personalized tour data to the user, and when the user moves, the user portrait can also be migrated to other tour devices, which can shorten the response time of the tour device and improve user tour satisfaction.

[0081] In order to make the exhibition hall tour method based on multi-modal interaction provided in the present application more clearly, a detailed embodiment and accompanying drawings will be combined below to explain the detailed embodiment, which includes the following steps: Figure 5

[0082] S501, sensing motion data of a user, predicting candidate visited exhibits of the user according to the motion data, and caching tour data of the candidate visited exhibits, and in the case that the user inputs multi-modal interaction information for the candidate visited exhibits, pushing the tour data of the candidate visited exhibits.

[0083] S502, judging whether the user enters an exhibit visiting area of the candidate visited exhibits according to the motion data, and in the case that the time length of the user entering the exhibit visiting area of the candidate visited exhibits is greater than a preset time length threshold, projecting a contour prompt projection of the candidate visited exhibits.

[0084] S503, in the case that the user is within a preset distance range of the exhibits, acquiring multi-modal interaction information input by the user for the exhibits, and in the case that the multi-modal interaction information includes voice information and touch information, determining the interaction behavior of the user based on only the voice information.

[0085] S504, in the case that the multi-modal interaction information includes voice information and gesture information, performing semantic fusion on the voice information and the gesture information, determining a fusion result, and determining the interaction behavior of the user based on the fusion result.

[0086] S505, determining the interaction depth of the user based on the interaction behavior of the user, and pushing tour content matched with the interaction depth.

[0087] S506, acquiring a motion trajectory of the user, and in the case that the motion trajectory represents that the user moves to adjacent exhibits of the exhibits, migrating the interaction behavior and the interaction depth of the user to a tour device corresponding to the adjacent exhibits.

[0088] S507, acquiring gait cycle and gaze hotspot distribution data of the user, determining the age of the user according to the gait cycle, and determining interest data of the user according to the gaze hotspot distribution data, and S generating a user portrait based on the age of the user and the interest data of the user; ​

[0089] S508, in the case that the motion trajectory represents that the user moves to a neighboring exhibit of the exhibit, migrating the user portrait to a guide device corresponding to the neighboring exhibit.

[0090] It should be understood that, although each step in the flowchart involved in each of the above embodiments is shown in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by the combination are within the scope of protection of the present application.

[0091] Based on the same inventive concept, the embodiments of the present application also provide a multi-modal interaction-based exhibition hall guide device for implementing the multi-modal interaction-based exhibition hall guide method described above. As shown in Figure 6 The multi-modal interaction-based exhibition hall guide device 600 includes a control module 610, and a space perception module 620, a multi-modal interaction module 630 and a guide display module 640 connected to the control module 610, respectively;

[0092] The space perception module 620 is configured to collect and send the position information and the motion trajectory of the user to the control module;

[0093] The control module 610 is configured to receive the position information, in the case that the user is determined to be within the preset distance range of the exhibit based on the position information, receive the multi-modal interaction information input by the user through the multi-modal interaction module 630 for the exhibit, determine the interaction behavior of the user based on the multi-modal interaction information, determine the interaction depth of the user based on the interaction behavior of the user, send a control instruction to the guide display module 640, the control instruction carrying the guide content matched with the interaction depth; and in the case that the motion trajectory represents that the user moves to a neighboring exhibit of the exhibit, migrating the interaction behavior and the interaction depth of the user to a guide device corresponding to the neighboring exhibit;

[0094] The guide display module 640 is configured to receive the control instruction and push the guide content matched with the interaction depth.

[0095] The problem-solving implementation provided by the system is similar to the implementation described in the above method, so the specific definition in one or more of the following embodiments of the multi-modal interaction-based exhibition guide device can refer to the definition of the multi-modal interaction-based exhibition guide method in the above, which will not be repeated here.

[0096] In one embodiment, as shown in Figure 7 The multi-modal interaction module 630 includes a voice collection sub-module 631, a touch input sub-module 632, and a gesture recognition sub-module 633.

[0097] The voice collection sub-module 631 includes a microphone array arranged in a ring shape and is configured to collect voice information input by the user for the exhibits.

[0098] The touch input sub-module 632 includes a capacitive touch screen and a pressure-sensitive film arranged in a stack and is configured to collect touch information input by the user for the exhibits.

[0099] The gesture recognition sub-module 633 is configured to collect gesture information input by the user for the exhibits.

[0100] The multi-modal interaction information includes at least one of voice information, touch information, and gesture information.

[0101] For example, the voice collection sub-module 631 includes microphone units arranged in a ring array, for example, six microphones arranged at equal intervals around the central axis of the exhibits, forming a dead-angle-free sound pickup area. The voice collection sub-module 631 supports adaptive beamforming technology, which can dynamically adjust the sound pickup direction to enhance the strength of the user's voice signal and suppress background noise from the back or side of the exhibits. Even in an environment with 85dB noise, the voice intelligibility can still be greater than or equal to 0.8. Moreover, the guide device can be locally deployed with a BERT-Tiny model, which can map the voice information input by the user to 17 pre-defined categories, such as common problem categories like "era query", "author information", etc.

[0102] The touch input sub-module 632 adopts a laminated structure design and is composed of a capacitive touch screen in the upper layer and a pressure-sensitive film in the lower layer. The capacitive touch screen is based on the capacitive sensing principle and is composed of a transparent conductive layer and an insulating layer, and can detect the capacitive change generated when a user's finger or a conductive object approaches. The pressure-sensitive film is made of piezoresistive material and can sense the size and distribution of the touch pressure. When the user touches the surface of the exhibit, the capacitive touch screen first detects the touch position and determines the touch point by scanning the capacitive change on the surface of the touch screen. At the same time, when the pressure-sensitive film is subjected to pressure, the resistance value of the piezoresistive material inside changes, and the size of the touch pressure is calculated by measuring the resistance change. When the touch pressure is 0-5 Newton, the user's light touch is recognized, and the graphic browsing is triggered. When the touch pressure is 5-10 Newton, the user's long press is recognized, and the expert mode of the guide device is activated. The touch position coordinates and pressure information are fused to obtain the touch information. The guide device can define a user touch gesture library in advance, such as pinching and rotating, and respond to the user's touch operation according to the user's touch position and touch pressure.

[0103] The gesture recognition sub-module 633 can be composed of a camera and an image processing unit, which can obtain the depth information of the scene in real time and capture the user's gesture image at the same time. The image processing unit is used for real-time processing and analysis of the gesture image and depth data. Exemplarily, the camera collects the user's gesture image and scene depth data, the image processing unit identifies the hand region in the user's gesture image and extracts the key point coordinates of the user's hand, combines the scene depth data, calculates the posture and motion trajectory of the user's hand in the three-dimensional space, and then identifies the user's gesture action, such as waving and rotating.

[0104] In this embodiment, the ring array distributed microphone unit eliminates the pickup blind area of the traditional single microphone, cooperates with the beamforming technology to realize high signal-to-noise ratio collection of the voice signal in the noisy environment of the exhibition hall, improves the interaction efficiency and accuracy of the voice information, the laminated design of the capacitive touch screen and the pressure-sensitive film realizes the double perception of the touch position and the pressure, can obtain more rich touch information, and improves the interaction efficiency and accuracy of the touch information. The gesture recognition sub-module can accurately capture the user's gesture action in a complex exhibition environment without the need for direct contact between the user and the device, so that the user can operate in a larger space range, and the degree of freedom and flexibility of user interaction is improved.

[0105] In one embodiment, as Figure 8As shown, the present application also provides a display rack 800 for carrying the above-mentioned multi-modal interaction-based exhibition hall guide device 600, which comprises a foldable asymmetric truss 810, a plurality of hinges 820 arranged at each connection point of the asymmetric truss 810, and a mobile chassis 830. The asymmetric truss 810 is folded and unfolded through the hinges 820. The mobile chassis 830 is arranged at the bottom of the asymmetric truss 810, and a plurality of universal wheel sets 840 are arranged at the bottom of the mobile chassis 830.

[0106] The geometric structure of the asymmetric truss 810 can be designed using a parameterized topological optimization algorithm. The geometric structure breaks the traditional symmetrical layout, ensures the structural strength (load-bearing capacity ≥ 50 kg), and realizes the lowest possible self-weight (empty weight 18 kg). After the asymmetric truss 810 is folded, it is compressed to 12.5% of the unfolded state.

[0107] The hinges 820 arranged at each connection point of the asymmetric truss 810 can be made of high-strength stainless steel. The hinges 820 can be equipped with bidirectional damping bearings and spring return mechanisms. When the asymmetric truss 810 is folded, the spring return mechanism releases elastic potential energy, and the resistance of the damping bearing controls the folding of the asymmetric truss 810 along the preset trajectory. When unfolded, the damping can be overcome by manual force, and the spring mechanism automatically locks the connection points to form a rigid connection structure, completing the unfolding of the asymmetric truss 810. In addition, the hinges 820 can be six-degree-of-freedom electromagnetic locking hinges. The joint angle of the asymmetric truss 810 is monitored in real time by a Hall sensor, and a servo motor is used to realize one-key unfolding. The asymmetric truss 810 is automatically unfolded and the screen posture is calibrated within 120 seconds. If the millimeter wave radar detects an obstacle within 5 centimeters when the asymmetric truss 810 is unfolded, the unfolding action is paused to achieve anti-collision protection.

[0108] The mobile chassis 830 is provided with a universal wheel set 840 at the bottom, which can be composed of four independent steering wheels. Each universal wheel is equipped with an independent DC servo motor, allowing each wheel to be independently controlled in terms of speed and steering direction. The universal wheel set 840 can be integrated with a piezoelectric sensor, which can be distributed on the contact part between the wheel and the ground. When the universal wheel set 840 rolls on the ground, the pressure change characteristics of different ground materials (tiles, carpets, slopes, etc.) are different, which will cause the piezoelectric ceramic to generate different intensity of charge signals, i.e. the ground material change can be identified. At this time, the control system can dynamically adjust the output torque of the four drive motors according to the preset ground material-torque mapping table. For example, on the tile floor, the motor output is medium torque to ensure mobility and stability; on the carpet floor, the torque is appropriately increased to overcome the resistance due to increased friction.

[0109] In addition, the mobile chassis 830 at the bottom can also be provided with a laser radar to emit a laser beam to detect the surrounding environment. When the laser beam encounters an obstacle, it is reflected back to the laser radar. By measuring the round-trip time of the laser, the point cloud data of the surrounding environment can be obtained. At the same time, by using the SLAM (simultaneous localization and mapping) algorithm, the continuous point cloud data is matched and fused to construct a three-dimensional map of the exhibition hall in real time. During the map construction process, the exhibition hall space can also be grid processed to mark the passable area, obstacle position and other information, and generate a passable area heat map of the exhibition hall to intuitively show the difficulty of passing and congestion in different areas.

[0110] In this embodiment, the geometric configuration design and foldable characteristics of the asymmetric truss in the display stand make the display stand can be unfolded and folded according to the space layout of the exhibition hall or the size of the exhibits, and the volume of the display stand can be greatly reduced in the folded state, thereby improving the space flexibility and space adaptability of the display stand. The connection points of the asymmetric truss adopt hinge structures, which can improve the exhibition efficiency of the display stand. The bottom of the asymmetric truss is also provided with a mobile chassis. The mobile chassis is provided with a plurality of universal wheel groups at the bottom to realize the movement of the display stand. Therefore, when the display stand is used to carry the multi-modal interactive exhibition guide system, the movement characteristics and the multi-modal interactive system are linked, which can improve the guide coverage and further improve the multi-modal interaction efficiency with the user.

[0111] In one embodiment, the hinge 820 is an electromagnetic locking hinge.

[0112] Based on the same inventive concept, the embodiments of the present application also provide a multi-modal interactive exhibition guide device for implementing the multi-modal interactive exhibition guide method described above. The problem-solving implementation scheme provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more multi-modal interactive exhibition guide device embodiments provided below can refer to the limitations of the multi-modal interactive exhibition guide method described above, which will not be repeated here.

[0113] In one exemplary embodiment, as Figure 9 shown, a multi-modal interactive exhibition guide device 900 is provided, which includes a data acquisition module 910, an interaction behavior determination module 920, a content pushing module 930 and a content migration module 940, wherein:

[0114] The data acquisition module 910 is configured to acquire multi-modal interaction information input by a user for an exhibit when the user is within a preset distance range of the exhibit.

[0115] The interaction behavior determination module 920 is configured to determine the interaction behavior of the user based on the multi-modal interaction information.

[0116] The content pushing module 930 is configured to determine the interaction depth of the user based on the interaction behavior of the user, and push the guide content matched with the interaction depth.

[0117] The content migration module 940 is configured to acquire the motion trajectory of the user, and in a case where the motion trajectory indicates that the user moves to a neighboring exhibit of the exhibit, migrate the interaction behavior and the interaction depth of the user to a guide device corresponding to the neighboring exhibit.

[0118] In an embodiment, the multi-modal interaction information includes at least one of voice information, touch information, and gesture information, and the interaction behavior determination module 920 is configured to, in a case where the multi-modal interaction information includes voice information and touch information, determine the interaction behavior of the user based on only the voice information, and in a case where the multi-modal interaction information includes voice information and gesture information, perform semantic fusion on the voice information and the gesture information, determine a fusion result, and determine the interaction behavior of the user based on the fusion result.

[0119] In an embodiment, the content pushing module 930 is further configured to, in a case where the interaction depth is a first interaction depth, push text and image guide data of the exhibit, in a case where the interaction depth is a second interaction depth, push animation guide data of the exhibit, and in a case where the interaction depth is a third interaction depth, push expert guide data of the exhibit, where the first interaction depth is less than the second interaction depth, and the second interaction depth is less than the third interaction depth.

[0120] In an embodiment, the multi-modal interaction-based exhibition guide device 900 is further configured to perceive motion data of the user, predict a candidate visited exhibit of the user according to the motion data, cache guide data of the candidate visited exhibit, and in a case where the user inputs multi-modal interaction information for the candidate visited exhibit, push the guide data of the candidate visited exhibit.

[0121] In an embodiment, the multi-modal interaction-based exhibition guide device 900 is further configured to determine, according to the motion data, whether the user enters an exhibit visiting area of the candidate visited exhibit, and in a case where a duration for which the user enters the exhibit visiting area of the candidate visited exhibit is greater than a preset duration threshold, project a contour prompt projection of the candidate visited exhibit.

[0122] In an embodiment, the multi-modal interaction-based exhibition guide device 900 is further configured to acquire gait cycle data and gaze hotspot distribution data of the user, determine an age of the user according to the gait cycle data, and determine interest data of the user according to the gaze hotspot distribution data, generate a user portrait based on the age of the user and the interest data of the user, and in a case where the motion trajectory indicates that the user moves to a neighboring exhibit of the exhibit, migrate the user portrait to a guide device corresponding to the neighboring exhibit.

[0123] The various modules in the above-mentioned exhibition hall guide device based on multi-modal interaction can be implemented wholly or partially by software, hardware, or a combination thereof. The various modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the various modules.

[0124] In an exemplary embodiment, a computer device, which can be a server, has an internal structure diagram as shown in Figure 10 The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store multi-modal interaction information and other data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with terminals outside through a network connection. The computer program is executed by the processor to implement a multi-modal interaction-based exhibition hall guide method.

[0125] In an exemplary embodiment, a computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-mentioned multi-modal interaction-based exhibition hall guide method embodiments.

[0126] In an embodiment, a computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in the above-mentioned multi-modal interaction-based exhibition hall guide method embodiments.

[0127] In an embodiment, a computer program product includes a computer program, and the computer program is executed by a processor to implement the steps in the above-mentioned multi-modal interaction-based exhibition hall guide method embodiments.

[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use, and processing of related data need to comply with relevant regulations.

[0129] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0130] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.

[0131] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A method for guiding exhibition halls based on multimodal interaction, characterized in that: The method comprises: When the user is within a preset distance range of the exhibit, obtaining multimodal interaction information input by the user for the exhibit; Determining the user's interaction behavior based on the multimodal interaction information; Determining the user's interaction depth based on the user's interaction behavior, and pushing navigation content that matches the interaction depth; A motion trajectory of the user is obtained, and when the motion trajectory indicates that the user moves toward an adjacent exhibit of the exhibit, the interaction behavior and interaction depth of the user are transferred to a guide device corresponding to the adjacent exhibit.

2. The method according to claim 1, characterized in that The multimodal interaction information includes at least one of voice information, touch information, and gesture information. The determining the user's interaction behavior based on the multimodal interaction information includes: In a case where the multimodal interaction information includes voice information and touch information, determining the user's interaction behavior based only on the voice information; In a case where the multimodal interaction information includes voice information and gesture information, semantic fusion is performed on the voice information and gesture information to determine a fusion result, and based on the fusion result, the user's interaction behavior is determined.

3. The method according to claim 2, characterized in that The pushed navigation content matching the interaction depth includes: When the interaction depth is the first interaction depth, pushing graphic and text guide data of the exhibit; When the interaction depth is the second interaction depth, pushing the animation guide data of the exhibit; When the interaction depth is the third interaction depth, pushing expert guided tour data of the exhibit; The first interaction depth is smaller than the second interaction depth, and the second interaction depth is smaller than the third interaction depth.

4. The method according to claim 1, wherein Before obtaining the multimodal interaction information input by the user regarding the exhibit, the method further includes: Sense user's motion data; predicting candidate exhibits to be visited by the user based on the motion data, and caching guide data of the candidate exhibits to be visited; When the user inputs multimodal interaction information for the candidate exhibit to be visited, navigation data of the candidate exhibit to be visited is pushed.

5. The method according to claim 4, characterized in that After sensing and obtaining the user's motion data, the method further includes: determining, based on the motion data, whether the user has entered an exhibit viewing area of ​​a candidate exhibit; When the time duration during which the user enters the exhibit viewing area of ​​the candidate exhibit is greater than a preset time duration threshold, a contour prompt projection of the candidate exhibit is projected.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Obtain the user's gait cycle and gaze hotspot distribution data; determining the user's age based on the gait cycle, and determining the user's interest data based on the gaze hotspot distribution data; Generate a user profile based on the user's age and the user's interest data; In a case where the motion trajectory represents that the user moves toward an adjacent exhibit of the exhibit, the user portrait is transferred to a guide device corresponding to the adjacent exhibit.

7. An exhibition hall guide device based on multimodal interaction, characterized in that: The device includes a control module, and a space perception module, a multimodal interaction module and a navigation display module respectively connected to the control module; The space perception module is configured to collect and send the user's location information and motion trajectory to the control module; The control module is configured to receive the location information, and if it is determined based on the location information that the user is within a preset distance range of the exhibit, receive multimodal interaction information input by the user for the exhibit through the multimodal interaction module, determine the user's interaction behavior based on the multimodal interaction information, determine the user's interaction depth based on the user's interaction behavior, and send a control instruction to the navigation display module, wherein the control instruction carries navigation content matching the interaction depth; and, when the motion trajectory represents that the user is moving toward an adjacent exhibit, migrating the user's interaction behavior and interaction depth to a guide device corresponding to the adjacent exhibit; The navigation display module is configured to receive the control instruction and push the navigation content matching the interaction depth.

8. The device according to claim 7, characterized in that The multimodal interaction module includes a voice acquisition submodule, a touch input submodule and a gesture recognition submodule; The voice collection submodule includes a microphone array arranged in a ring manner, and is configured to collect voice information input by the user regarding the exhibits; The touch input submodule includes a capacitive touch screen and a pressure-sensitive film that are stacked, and is configured to collect touch information input by the user on the exhibit; The gesture recognition submodule is configured to collect gesture information input by the user for the exhibit; The multimodal interaction information includes at least one of voice information, touch information and gesture information.

9. A display stand for carrying the exhibition hall guide device based on multimodal interaction according to claim 7 or 8, characterized in that: The display stand includes a foldable asymmetric truss, multiple hinges and a mobile chassis. The hinges are arranged at each connection point of the asymmetric truss, and the asymmetric truss completes folding and unfolding actions through the hinges. The mobile chassis is arranged at the bottom of the asymmetric truss, and multiple universal wheel sets are arranged at the bottom of the mobile chassis.

10. The display stand according to claim 9, characterized in that: The hinge is an electromagnetic locking hinge.