Smart building system interaction experience method and system combined with VR
By collecting user behavior data in VR interactive terminals to generate virtual-real mapping data, constructing intent evolution trajectories and driving the reorganization of virtual building scenes, the intuitiveness and personalization issues of existing building system interaction methods are solved, improving user experience and interaction efficiency.
Patent Information
- Application Number
- CN202510944994.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing building system interaction methods lack intuitiveness and immersion, making it difficult for users to fully understand the functions of the equipment. Furthermore, the interaction process cannot be dynamically adjusted according to user behavior, resulting in a less rich and personalized user experience.
By collecting continuous behavioral sequences of users in physical space through VR interactive terminals, virtual behavior mapping data is generated, an intent evolution trajectory is constructed, dynamic reconstruction instructions are generated, virtual building scenes are driven to perform real-time spatial reorganization, and multimodal feedback signals are provided to optimize the prediction model of intent evolution trajectory.
It enables dynamic reconfiguration of virtual building scenes, enhancing the realism and immersion of the interaction, providing a personalized user experience, and improving user satisfaction and interaction efficiency.
Smart Images

Figure CN120848725B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of virtual reality, in particular to a smart building system interaction experience method and system combined with VR. BACKGROUND
[0002] In the current field of building system interaction experience, traditional interaction methods mainly rely on fixed facilities in physical space and simple operation interfaces. For example, users control devices in the building, such as lights, air conditioners, access control, etc., by operating physical buttons, touch screens, or using remote controls. The above-mentioned interaction methods have many limitations. On the one hand, users need to be familiar with the operation methods of various devices, increasing the learning cost. On the other hand, the interaction process lacks intuitiveness and immersion, and users cannot fully and deeply understand the functions and running status of the building system.
[0003] With the development of virtual reality (VR) technology, although some cases of applying VR to building display have appeared, these applications mostly only present the physical structure of the building in a simple virtual way. Users can only passively watch the virtual scene and cannot interact with the virtual scene in depth, nor can they dynamically adjust the virtual scene according to the real-time behavior and intention of the user, resulting in a lack of richness and individualization of user experience, which cannot meet the growing demand of users for intelligent and interactive building experience. SUMMARY
[0004] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide a smart building system interaction experience method combined with VR, which comprises:
[0005] Collecting a continuous behavior sequence of a user in a physical space through a VR interaction terminal to generate virtual-real behavior mapping data, the virtual-real behavior mapping data containing the association between the physical behavior of the user and the functional area of the virtual building scene;
[0006] Building an intention evolution trajectory of the user in the virtual building scene based on the virtual-real behavior mapping data, the intention evolution trajectory containing the intention change trend in the time dimension and the area attention path in the space dimension;
[0007] Generating a dynamic reconstruction instruction of the virtual building scene according to the intention evolution trajectory, the dynamic reconstruction instruction containing a space layout adjustment rule and an interaction element activation strategy;
[0008] Driving the virtual building scene to perform real-time space reorganization according to the dynamic reconstruction instruction, and generating a multi-modal feedback signal, the multi-modal feedback signal containing visual scene update data and haptic interaction response data;
[0009] Collecting real-time response behavior of the user to the multi-modal feedback signal, and optimizing a prediction model of the intention evolution trajectory based on the real-time response behavior.
[0010] In still another aspect, the embodiment of the present application also provides a smart building system interactive experience system combined with VR, which comprises a processor and a machine readable storage medium, the machine readable storage medium is connected with the processor, the machine readable storage medium is used for storing programs, instructions or codes, and the processor is used for executing the programs, instructions or codes in the machine readable storage medium to realize the above method.
[0011] Based on the above aspects, the embodiment of the present application generates virtual-real behavior mapping data by collecting continuous behavior sequences of the user in the physical space, can accurately establish the association between the physical behavior of the user and the functional area of the virtual building scene, constructs the intention evolution trajectory based on the virtual-real behavior mapping data, comprehensively considers the intention change trend in the time dimension and the area attention path in the space dimension, and can comprehensively and dynamically grasp the intention of the user. According to the intention evolution trajectory, dynamic reconstruction instructions are generated, the real-time spatial reorganization of the virtual building scene according to the intention of the user is realized, the virtual scene is more in line with the needs of the user. The multi-modal feedback signal containing visual scene update data and haptic interaction response data is generated, rich sensory experience is provided for the user, and the sense of reality and immersion of the interaction is enhanced. The real-time response behavior of the user to the multi-modal feedback signal is collected and the prediction model of the intention evolution trajectory is optimized, which can continuously improve the prediction accuracy of the system for the intention of the user, thereby providing more personalized and intelligent building system interactive experience for the user, and significantly improving the user satisfaction and interaction efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is the execution flow diagram of the smart building system interactive experience method combined with VR provided by an embodiment of the present application.
[0013] Figure 2 is the schematic diagram of exemplary hardware and software components of the smart building system interactive experience system combined with VR provided by an embodiment of the present application. DETAILED DESCRIPTION
[0014] The present application will be described in detail below with reference to the accompanying drawings of the specification, Figure 1 is the flow diagram of the smart building system interactive experience method combined with VR provided by an embodiment of the present application, and the smart building system interactive experience method combined with VR will be described in detail below.
[0015] Step S110: collecting continuous behavior sequences of the user in the physical space through the VR interactive terminal, generating virtual-real behavior mapping data, and the virtual-real behavior mapping data contains the association between the physical behavior of the user and the functional area of the virtual building scene.
[0016] In the intelligent building system interaction experience scene combined with VR, the user wears a VR interaction terminal to enter a physical space corresponding to a virtual building scene, such as a simulated large shopping center physical space. After the VR interaction terminal is started, it begins to collect the continuous behavior sequence of the user in the physical space. Various sensors and components inside the VR interaction terminal cooperate with each other to collect the behavior information of the user in all directions, so as to construct the association between the physical behavior of the user and the functional area of the virtual building scene.
[0017] Step S111: Start the behavior capture module of the VR interaction terminal, activate the built-in limb sensor and gesture recognition component, the limb sensor is used to collect the limb movement trajectory data of the user, and the gesture recognition component is used to collect the gesture action sequence data of the user.
[0018] When the user wears the VR interaction terminal, the VR interaction terminal automatically triggers the start of the behavior capture module. The limb sensor can sense the position change of each part of the user's body in the physical space. In the simulated large shopping center physical space, the user may walk between various store passages, enter the store to view goods, etc., and the limb sensor will continuously record the movement trajectory of the user's limbs in the physical space, forming the limb movement trajectory data. These limb movement trajectory data reflect the position movement of the user in the physical space.
[0019] At the same time, the gesture recognition component is also activated, which captures the user's hand movements by means of special sensing technology and image processing capability. In the shopping center, the user may point to a store sign with his fingers, make a gesture to select goods, etc., and the gesture recognition component will record these gesture actions in time sequence to form gesture action sequence data.
[0020] Step S112: Perform spatial coordinate conversion processing on the limb movement trajectory data, map the trajectory data under the physical space coordinate system to a three-dimensional movement path under the virtual building scene coordinate system, and the three-dimensional movement path includes a real-time position coordinate sequence of the user in the virtual building scene.
[0021] The collected limb movement trajectory data is initially based on the physical space coordinate system, in order to accurately present the movement of the user in the virtual building scene, spatial coordinate conversion is needed. First, the corresponding relationship between the physical space coordinate system and the virtual building scene coordinate system is determined. This involves calibrating the origin, coordinate axis direction and scale of the two coordinate systems. Through a series of position mapping operations, each point on the limb movement trajectory of the user in the physical space is converted to the coordinate system of the virtual building scene.
[0022] For example, in a simulated large shopping center physical space, a user moves from an entrance to a dining area, which has a series of corresponding coordinate points in the physical space coordinate system. After coordinate conversion, these coordinate points form a new coordinate sequence in the virtual building scene coordinate system, constituting a three-dimensional movement path of the user in the virtual building scene. The three-dimensional movement path contains a real-time position coordinate sequence of the user in the virtual building scene, accurately reflecting the movement trajectory of the user in the virtual space.
[0023] Step S113: key frame extraction processing is performed on the gesture action sequence data, starting frame, peak frame and ending frame of the gesture action are identified, gesture contour features and motion vectors of each key frame are extracted, and the type and execution strength of the gesture action are determined.
[0024] For the collected gesture action sequence data, key frame extraction processing is performed. First, each frame of gesture image is analyzed, and the differences between adjacent frames are compared to identify the key frame. When the difference between adjacent frames first reaches a certain degree, the frame is determined as the starting frame of the gesture action. After the starting frame, the difference accumulation between the subsequent frames is continuously analyzed, and when the difference accumulation reaches the maximum value, the corresponding frame is the peak frame. After the peak frame, when the difference between adjacent frames is continuously below a certain degree for multiple times, the last frame below the difference threshold is determined as the ending frame.
[0025] After determining the key frame, the gesture image of the key frame is processed. For the extraction of gesture contour features, an edge detection method is used. By detecting the mutation of pixel values in the gesture image, the boundary points of the gesture are found out, and the coordinate information of these boundary points is integrated to form the feature description of the gesture contour. The calculation of the motion vector is by comparing the position changes of the gesture between the peak frame and the starting frame and between the peak frame and the ending frame. The position offset of the gesture between the peak frame and the starting frame is compared to determine the motion direction of the gesture action; and the position offset between the peak frame and the ending frame is combined to calculate the amplitude of the motion.
[0026] The extracted gesture contour features are compared with the pre-set gesture template library. The gesture template library contains the feature descriptions of various common gestures, and the type of the current gesture action is determined by matching the similarity. For the execution strength of the gesture action, the motion amplitude is normalized to be within a specific range, which is used as the execution strength parameter of the gesture action.
[0027] Step S1131: frame difference calculation is performed on the gesture action sequence data, the pixel change amount between adjacent frames is obtained, and the frame where the pixel change amount first exceeds the pre-set threshold is marked as the starting frame of the gesture action.
[0028] In processing the gesture action sequence data, inter-frame difference calculation is performed on two adjacent gesture images. By comparing the pixel value difference of corresponding pixel points in two images, the pixel change amount between each pair of adjacent frames is obtained. A preset threshold is set, and when the pixel change amount between a certain adjacent frame first exceeds the threshold, the frame is marked as the starting frame of the gesture action. For example, in a simulated large shopping center, the user raises his hand to make a gesture pointing to a certain store, and in the gesture action sequence data, when the pixel change amount between adjacent frames first reaches the preset change degree, the corresponding frame is the starting frame.
[0029] Step S1132: In the frame sequence after the starting frame, the cumulative value of the pixel change amount is continuously calculated, and when the cumulative value reaches the maximum value, the corresponding frame is marked as the peak frame of the gesture action.
[0030] From the starting frame, the subsequent frame sequence is continuously processed. The cumulative value of the pixel change amount between adjacent frames is continuously calculated. As the gesture action proceeds, the pixel change amount is continuously accumulated. When the cumulative value reaches the maximum value, it means that the gesture action has reached the highest point of intensity or amplitude, and the corresponding frame is marked as the peak frame of the gesture action. For example, when the user's hand gesture pointing to the store gradually reaches the maximum amplitude, the corresponding frame is the peak frame.
[0031] Step S1133: In the frame sequence after the peak frame, when the pixel change amount is continuously lower than the preset threshold for multiple frames, the last frame whose change amount is lower than the threshold is marked as the ending frame of the gesture action.
[0032] After the peak frame is determined, the frame sequence after the peak frame is monitored. When the pixel change amount between adjacent frames is continuously lower than the preset threshold for multiple frames, it means that the gesture action has approached the end. The last frame whose pixel change amount is lower than the threshold is marked as the ending frame of the gesture action. For example, the user's hand gesture pointing to the store ends, and the hand begins to lower, and when the pixel change amount between adjacent frames is continuously less than the preset threshold for multiple times, the last frame is the ending frame.
[0033] Step S1134: Extract the gesture image data of the starting frame, the peak frame and the ending frame, perform edge detection processing on the gesture image data, obtain the boundary point coordinate set of the gesture contour, and convert the boundary point coordinate set into a contour feature vector as the gesture contour feature.
[0034] After the gesture image data of the start frame, the peak frame and the end frame are extracted, edge detection processing is performed on the images. The edge detection finds the boundary of the gesture by identifying the mutation of pixel values in the image. After detecting the boundary points of the gesture contour, the coordinate information of the boundary points is collected to form a boundary point coordinate set. Then, the boundary point coordinate set is converted into a contour feature vector with a specific dimension and format, which is the gesture contour feature and can accurately describe the shape and contour information of the gesture.
[0035] Step S1135: Calculate the pixel displacement between the peak frame and the start frame to generate a motion direction vector of the gesture action, and calculate the length of the motion direction vector as the motion amplitude of the gesture action by combining the pixel displacement between the peak frame and the end frame.
[0036] By comparing the positions of the gesture in the peak frame and the start frame, the pixel displacement between the two is calculated. According to the displacement, the motion direction of the gesture action is determined, thereby generating a motion direction vector of the gesture action. At the same time, the length of the motion direction vector is further calculated by combining the pixel displacement between the peak frame and the end frame. The length represents the motion amplitude of the gesture action in the whole process, reflecting the strength and range of the gesture action.
[0037] Step S1136: Compare the gesture contour feature with the preset gesture template library to determine the type of gesture action, and normalize the motion amplitude as the execution strength parameter of the gesture action.
[0038] The extracted gesture contour feature is compared with the preset gesture template library. The gesture template library stores the feature descriptions of various common gestures. By calculating the similarity between the gesture contour feature and the template features in the template library, the most matching template is found, thereby determining the type of gesture action. For example, it can be determined as a "pointing" or "clicking" gesture type.
[0039] The calculated motion amplitude is normalized. The motion amplitude is adjusted to a specific range to make it comparable and standardized. The normalized motion amplitude is used as the execution strength parameter of the gesture action, which can accurately reflect the execution strength of the gesture action.
[0040] Step S114: Align the three-dimensional movement path and the gesture action type by timestamp to generate a behavior unit containing spatial position information and action feature information, and each behavior unit corresponds to a complete physical behavior segment.
[0041] The three-dimensional movement path and gesture action type obtained above are aligned according to timestamps. The timestamps record each data collection time, and the position information on the three-dimensional movement path and the gesture action type at the corresponding time are associated through the timestamps. In this way, at the same time point, there is the position information of the user in the virtual building scene and the corresponding gesture action type, forming a behavior unit.
[0042] Each behavior unit corresponds to a complete physical behavior segment. For example, in the simulated large shopping center, the user walks from a resting area to a clothing store while making a gesture of pointing to the store sign. After the corresponding three-dimensional movement path and gesture action type are aligned according to timestamps, a behavior unit is formed, which completely describes the user's physical behavior in this time period.
[0043] Step S115: Sequence integration processing is performed on the behavior units to combine them in chronological order to form a continuous behavior sequence of the user in the physical space, each behavior unit in the continuous behavior sequence being associated with a corresponding time marker and action feature parameter.
[0044] The generated behavior units are sequence-integrated. The behavior units are arranged and combined in chronological order. Each behavior unit is associated with a corresponding time marker and action feature parameter. The time marker accurately records the time of the behavior, and the action feature parameter includes gesture action type, execution strength, etc.
[0045] In the simulated large shopping center, the user may have multiple continuous behaviors, such as walking to the dining area first, making a gesture of checking the menu, and then walking to the entertainment area and making a gesture of selecting an entertainment project. Combining the behavior units corresponding to these behaviors in chronological order forms a continuous behavior sequence of the user in the physical space, clearly presenting the user's behavior process in a period of time.
[0046] Step S116: The continuous behavior sequence is matched with the function area division data of the virtual building scene to determine the virtual function area identifier corresponding to each behavior unit, calculate the interaction intention probability of the behavior unit based on the gesture action type, and generate virtual-real behavior mapping data containing the virtual area identifier, the interaction intention probability, and the time marker.
[0047] The generated continuous behavior sequence is matched with the function area division data of the virtual building scene. The virtual building scene is divided into different function areas, such as shopping area, dining area, resting area, etc., each function area having a corresponding identifier. By comparing the position information of the user in the behavior unit with the position range of the function area, the virtual function area identifier corresponding to each behavior unit is determined.
[0048] The interaction intention probability of the behavior unit is calculated according to the gesture action type. Different gesture action types reflect different interaction intentions of the user. For example, pointing to a store sign may indicate an intention to enter the store, and making a hand gesture to order food may indicate an intention to consume food and drink. According to preset rules and statistical data, the interaction intention probability of each behavior unit is calculated.
[0049] The virtual area identifier, interaction intention probability, and time marker are integrated together to generate virtual-real behavior mapping data. The virtual-real behavior mapping data contains the association between the user's physical behavior and the functional area of the virtual building scene.
[0050] Step S120: Based on the virtual-real behavior mapping data, an intention evolution trajectory of the user in the virtual building scene is constructed, which contains the intention change trend in the time dimension and the area attention path in the space dimension.
[0051] After obtaining the virtual-real behavior mapping data, an intention evolution trajectory of the user in the virtual building scene is constructed based on the virtual-real behavior mapping data. The intention evolution trajectory can comprehensively reflect the intention change of the user at different times and spaces, which is of great significance for understanding user needs and optimizing the virtual building scene.
[0052] Step S121: The virtual-real behavior mapping data is parsed to extract the virtual area identifier, interaction intention probability, and time marker of each behavior unit, and arranged in sequence according to the time marker to form a behavior unit sequence.
[0053] The virtual-real behavior mapping data is parsed to extract the virtual area identifier, interaction intention probability, and time marker of each behavior unit. The virtual area identifier indicates the functional area in the virtual building scene where the user's behavior occurs, the interaction intention probability reflects the possibility of the user's interaction intention at that behavior, and the time marker records the time when the behavior occurs.
[0054] According to the order of the time marker, these behavior units are arranged to form a behavior unit sequence. In the simulated large shopping center, through the behavior unit sequence, it can be clearly observed which functional areas the user is in at different time points and the corresponding interaction intention probability.
[0055] Step S122: The frequency and duration of each virtual area identifier in the behavior unit sequence are counted, and the virtual area identifier with the highest frequency and longest duration is determined as the current main attention area of the user.
[0056] The statistical analysis is performed on the sequence of behavior units. The frequency of each virtual region identifier appearing in the sequence, i.e. the number of times the user behaves in the virtual region, is counted. At the same time, the duration of the user in each virtual region is calculated, and the time markers of the behavior units are compared to determine the length of time the user stays in the region.
[0057] By comprehensively comparing the frequency and duration, the virtual region identifier with the highest frequency and the longest duration is found. In the simulated large shopping center, if the user appears in a certain clothing area the most number of times and stays the longest, then the virtual region identifier of the clothing area is determined as the current main focus region of the user, reflecting the user's current main interest.
[0058] Step S123: The interaction intention probability in the sequence of behavior units is weighted calculated, and the time marker is used as a weight factor to calculate the intention confidence of each behavior unit. The behavior units with an intention confidence higher than a preset threshold are screened as effective intention units.
[0059] The interaction intention probability in the sequence of behavior units is weighted calculated. The time marker is used as a weight factor because the behavior closer in time is more likely to reflect the user's current intention. Through a specific weighted calculation method, the intention confidence of each behavior unit is calculated by combining the interaction intention probability and the time marker.
[0060] A preset threshold is set to screen the calculated intention confidence. The behavior units with an intention confidence higher than the preset threshold are screened out as effective intention units. These effective intention units can better represent the user's real intention.
[0061] Step S124: The virtual region identifier change and the interaction intention probability change between adjacent effective intention units are analyzed, the intention transfer probability is calculated, and an intention transfer network is constructed, wherein the nodes of the intention transfer network are interaction intention types, and the edges are intention transfer probabilities.
[0062] The screened effective intention units are analyzed. The virtual region identifier change and the interaction intention probability change between adjacent effective intention units are compared. The change of the virtual region identifier reflects the user's transfer between different functional regions, and the change of the interaction intention probability embodies the user's intention change.
[0063] According to these changes, the intention transfer probability is calculated. The intention transfer probability represents the possibility of transferring from one interaction intention type to another interaction intention type. The interaction intention type is used as a node, and the intention transfer probability is used as the weight of the directed edge between nodes to construct an intention transfer network. The intention transfer network can clearly show the transfer relationship between user intentions.
[0064] Step S1241: Extract two adjacent effective intent units in the effective intent unit sequence, denoted as a preceding intent unit and a subsequent intent unit, and obtain the interactive intent type of the preceding intent unit and the interactive intent type of the subsequent intent unit.
[0065] Two adjacent effective intent units are extracted from the effective intent unit sequence, denoted as a preceding intent unit and a subsequent intent unit. The preceding intent unit represents the effective intent of the user at the previous time, and the subsequent intent unit represents the effective intent at the later time. The interactive intent types of the two intent units are obtained, for example, the preceding intent unit can be "viewing goods", and the subsequent intent unit can be "buying goods".
[0066] Step S1242: Count the number of combinations of all preceding intent types and subsequent intent types in the effective intent unit sequence, calculate the proportion of the occurrence frequency of each combination in the total combination frequency as the intent transition probability.
[0067] The combinations of all preceding intent types and subsequent intent types in the effective intent unit sequence are counted. The number of occurrences of each combination is counted, and then the proportion of the occurrence frequency of each combination in the total combination frequency is calculated. This proportion is the intent transition probability, which reflects the possibility of transition from one interactive intent type to another interactive intent type. For example, the number of combinations of "viewing goods" to "buying goods" is counted, and the total number of combinations is counted, and the proportion of the two is calculated to obtain the corresponding intent transition probability.
[0068] Step S1243: Construct an initial intent transition network with interactive intent types as network nodes and intent transition probabilities as directed edge weights between nodes.
[0069] Interactive intent types are used as network nodes, and each node represents a possible interactive intent. The calculated intent transition probability is used as the directed edge weight between nodes to construct an initial intent transition network. The initial intent transition network shows the transition relationship and possibility size between different interactive intent types.
[0070] Step S1244: Analyze the change of the virtual area identifier in the effective intent unit sequence, when the virtual area identifiers of adjacent effective intent units are different, adjust the weight value of the corresponding intent transition probability, and increase the probability weight of cross-region intent transition.
[0071] The change of the virtual area identifier in the effective intent unit sequence is analyzed. When the virtual area identifiers of adjacent effective intent units are different, it indicates that the user has transferred between different functional areas, and the above-mentioned cross-area transfer may be accompanied by a large change in the intent. At this time, the weight value of the corresponding intent transfer probability is adjusted, and the probability weight of the cross-area intent transfer is increased. For example, when the user transfers from the shopping area to the catering area, the corresponding intent transfer probability weight is increased to more accurately reflect the change in the user's intent.
[0072] Step S1245: When the change amplitude of the interactive intent probability of adjacent effective intent units exceeds a preset threshold, the intent transfer probability is adjusted again, and the transfer weight of the significantly changed intent is increased.
[0073] When the change amplitude of the interactive intent probability of adjacent effective intent units exceeds a preset threshold, it indicates that the user's intent has changed significantly. At this time, the intent transfer probability is adjusted again, and the transfer weight of the significantly changed intent is increased. For example, the user suddenly changes from a lower "viewing goods" intent probability to a higher "purchasing goods" intent probability, which exceeds the preset threshold, so the probability weight of the above-mentioned intent transfer is increased to more accurately capture the change in the user's intent.
[0074] Step S1246: The adjusted intent transfer probability is updated to the initial intent transfer network to generate an intent transfer network containing area association features.
[0075] The adjusted intent transfer probability is updated to the initial intent transfer network. By updating the weight value, the intent transfer network more accurately reflects the transfer of the user's intent, while considering the changes in the virtual area identifier and the significant changes in the interactive intent probability, generating an intent transfer network containing area association features. The intent transfer network not only considers the transfer probability between interactive intent types, but also incorporates the influence of the change of the virtual area identifier on the intent transfer, making the description of the evolution of the user's intent more accurate and comprehensive.
[0076] Step S125: Based on the intent transfer network and the current main focus area, the next possible interactive intent type and the corresponding virtual area of the user are predicted, and an intent evolution node is generated, which contains the predicted intent type, the target area identifier, and the prediction probability.
[0077] The next possible interactive intent type and the corresponding virtual area of the user are predicted using the constructed intent transfer network and the determined current main focus area. In the intent transfer network, the interactive intent type corresponding to the current main focus area is taken as the starting point, and the possibility of the user transferring to other intent types is calculated according to the intent transfer probability. In combination with the association between the virtual area and the interactive intent type, the target area identifier corresponding to each predicted intent type is determined.
[0078] The predicted intention type, target area identifier, and corresponding prediction probability are integrated together to generate an intention evolution node. For example, in a simulated large shopping center, if the current main focus area is the clothing area, according to the intention transfer network, it is predicted that the user's next step may have the intention to try on clothes in the fitting room, and the virtual area identifier where the fitting room is located is the target area identifier. At the same time, the probability of the above prediction is calculated to form an intention evolution node containing the predicted intention type, target area identifier, and prediction probability.
[0079] Step S126: The effective intention unit sequence and the intention evolution node are integrated in time sequence to form an intention evolution trajectory containing a historical intention trajectory and a future prediction direction, and each node in the intention evolution trajectory is associated with a corresponding time stamp, virtual area identifier, and intention confidence.
[0080] The effective intention unit sequence and the intention evolution node are integrated in time sequence. The effective intention unit sequence represents the user's past intention trajectory, and the intention evolution node embodies the prediction of the user's future intention. They are arranged in time sequence to form a complete intention evolution trajectory.
[0081] In the intention evolution trajectory, each node is associated with a corresponding time stamp, virtual area identifier, and intention confidence. The time stamp records the time when the intention occurs or is predicted, the virtual area identifier indicates the specific functional area where the intention occurs, and the intention confidence reflects the reliability of the intention. Through the intention evolution trajectory, the user's intention change trend and spatial dimension area focus path from the past to the future can be clearly observed.
[0082] Step S130: Generating dynamic reconstruction instructions of the virtual building scene according to the intention evolution trajectory, the dynamic reconstruction instructions containing spatial layout adjustment rules and interactive element activation strategies.
[0083] After obtaining the intention evolution trajectory of the user, dynamic reconstruction instructions of the virtual building scene are generated based on the intention evolution trajectory. The dynamic reconstruction instructions aim to adjust the virtual building scene in real time according to the user's intention change to provide a more user-demand-oriented interactive experience.
[0084] Step S131: Analyzing the intention evolution trajectory to extract the virtual area identifier of the current main focus area, the predicted intention type of the intention evolution node, and the target area identifier.
[0085] The intention evolution trajectory is analyzed to extract key information. The virtual area identifier of the current main focus area clearly indicates the functional area that the user is currently most interested in, the predicted intention type of the intention evolution node indicates the user's next possible interaction intention, and the target area identifier corresponds to the virtual area where the predicted intention may occur. For example, in a simulated large shopping center, after analyzing the intention evolution trajectory, it is determined that the current main focus area is the cosmetics area, the predicted intention type is to buy cosmetics, and the target area identifier is the checkout area.
[0086] Step S132: Retrieve the basic layout data corresponding to the current main focus area and the target area identifier from the area configuration library of the virtual building scene, which includes the position parameters of the fixed structures in the area, the initial state of the interactive elements, and the spatial connection relationship.
[0087] The basic layout data corresponding to the current main focus area and the target area identifier is retrieved from the area configuration library of the virtual building scene. The area configuration library stores detailed information of each functional area in the virtual building scene, and the basic layout data includes the position parameters of the fixed structures in the area, such as the positions of walls, columns, etc.; the initial state of the interactive elements, such as whether the button is clickable, whether the display stand is visible, etc.; and the spatial connection relationship, such as the position and connection method of the passageway between different areas.
[0088] Step S133: Match the predicted intention type with the preset reconstruction rule library to determine the target type of dynamic reconstruction, which includes area scaling reconstruction, element priority reconstruction, path guidance reconstruction, and function combination reconstruction.
[0089] According to the extracted predicted intention type, match with the preset reconstruction rule library. The reconstruction rule library contains reconstruction rules and target types corresponding to different predicted intention types. Through matching, the target type of dynamic reconstruction is determined. Possible target types include area scaling reconstruction, which adjusts the display size of the area; element priority reconstruction, which changes the display and operation priority of the interactive elements; path guidance reconstruction, which plans the path to guide the user to the target area; and function combination reconstruction, which combines the interactive elements of different areas to realize specific functions. For example, if the predicted intention type is to quickly find a certain commodity, the target type of path guidance reconstruction may be matched.
[0090] Step S134: If the target type is area scaling reconstruction, calculate the scaling coefficient according to the intention confidence of the current main focus area, generate a space layout adjustment rule that adjusts the display size of the area according to the scaling coefficient, and the higher the intention confidence, the larger the scaling coefficient.
[0091] When the target type of dynamic reconstruction is region scaling reconstruction, the scaling coefficient is calculated according to the intention confidence of the current main focus region. The intention confidence reflects the degree of user's attention and the reliability of the intention to the current main focus region. The higher the intention confidence, the greater the user's interest in the region, and the more prominent the region needs to be displayed, so the scaling coefficient is larger.
[0092] The intention confidence is converted into a scaling coefficient through a specific calculation method. Then, a spatial layout adjustment rule is generated according to the scaling coefficient, which is used to adjust the display size of the current main focus region. For example, in a simulated large shopping center, if the current main focus region is the electronic product area and the intention confidence is high, the scaling coefficient calculated is large, and the electronic product area is enlarged in the virtual building scene according to the scaling coefficient to attract the user's attention.
[0093] Step S135: If the target type is element priority reconstruction, extract the core interactive elements in the region according to the predicted intention type, set the display level of the core interactive elements higher than that of other elements, and generate an interactive element activation strategy to improve the visibility of the core interactive elements.
[0094] When the target type is element priority reconstruction, the core interactive elements are extracted from the region according to the predicted intention type. The core interactive elements are elements closely related to the predicted intention, for example, if the predicted intention type is to purchase goods, then the purchase button, price label, etc. of the goods are core interactive elements.
[0095] The display level of the core interactive elements is set higher than that of other elements, and the display order of the elements is adjusted to make the core interactive elements more prominent and visible. At the same time, an interactive element activation strategy is generated to improve the visibility of the core interactive elements, including setting the color, brightness, etc. of the core interactive elements, so that they are more eye-catching in the virtual building scene. For example, in the clothing area of a simulated large shopping center, if the predicted intention type is to view the details of a certain clothing, then the core interactive elements such as the clothing detail display button and the size selection box will be set to a higher display level and may change color to attract user clicks.
[0096] Step S136: If the target type is path guidance reconstruction, plan the optimal moving path according to the spatial position relationship between the current main focus region and the target region identifier, and generate a spatial layout adjustment rule to add visual guidance markers on the optimal moving path.
[0097] When the target type is path guidance reconstruction, first analyze the spatial position relationship between the current main focus region and the target region identifier. In the virtual building scene, each region has its position coordinates in space, and by comparing these coordinates, the relative position and distance between the two can be determined.
[0098] Based on the obstacle distribution data of the virtual building scene, the straight-line distance directly connecting the current main focus area and the target area is processed for obstacle avoidance. Through search algorithms such as A* algorithm, multiple candidate moving paths that avoid obstacles are found. Then, these candidate moving paths are evaluated, considering factors such as path length and number of turns, and the candidate moving path with shorter path length and fewer turns is determined as the optimal moving path.
[0099] Visual guidance markers are added to the optimal moving path to generate space layout adjustment rules. The visual guidance markers can be arrows, lines, etc., used to guide users to move from the current main focus area to the target area in the virtual building scene. For example, in a simulated large shopping center, if the user is currently in the dining area and predicts the next step to go to the entertainment area, the system will plan an optimal moving path that avoids other shops and obstacles, and add visual guidance markers such as arrows on the path to help the user find the route to the entertainment area.
[0100] Step S1361: Extract the center coordinates of the current main focus area and the center coordinates of the target area identification from the spatial map data of the virtual building scene, and calculate the straight-line distance and relative direction between the two center coordinates.
[0101] The center coordinates of the current main focus area and the target area are obtained from the spatial map data of the virtual building scene. The spatial map data records the position information of each area in the virtual building scene. By extracting the center coordinates of the two areas, the straight-line distance between them is calculated. Coordinate calculation methods can be used to determine the distance by comparing the difference between the two coordinates. At the same time, according to the relative position of the coordinates, the relative direction between the two areas is determined, such as due east, due west, northeast, etc.
[0102] Step S1362: Based on the obstacle distribution data of the virtual building scene, the straight-line distance is processed for obstacle avoidance, generating multiple candidate moving paths, which are continuous paths connecting the center coordinates of the current main focus area and the center coordinates of the target area identification and not passing through obstacles.
[0103] The straight-line distance calculated is processed using the obstacle distribution data of the virtual building scene. The obstacle distribution data records the position and range of all obstacles in the virtual building scene. When planning a moving path, these obstacles need to be avoided. Through search algorithms, starting from the center coordinates of the current main focus area, multiple continuous paths that can connect to the center coordinates of the target area without passing through obstacles are tried to find. These paths are candidate moving paths, which are continuous in space and bypass all obstacles.
[0104] Step S1363: Calculate the path length and the number of turns of each candidate moving path, and determine the candidate moving path with the highest priority as the optimal moving path. The shorter the path length and the fewer the number of turns, the higher the priority of the candidate moving path.
[0105] Each generated candidate moving path is evaluated. The length of each path is calculated, which is the sum of the lengths of the line segments on the path. At the same time, the number of turns in the path is counted, which reflects the degree of tortuosity of the path. The shorter the path length and the fewer the number of turns of the candidate moving path, the more convenient and efficient the user moves on this path, and thus the higher the priority.
[0106] By comparing the path length and the number of turns of each candidate moving path, the candidate moving path with the highest priority is determined as the optimal moving path. For example, in a simulated large shopping center, there are multiple candidate moving paths from the clothing area to the checkout area. After evaluation, a path with the shortest length and the fewest turns is selected as the optimal moving path.
[0107] Step S1364: Set guide marker points at equal intervals on the optimal moving path, and dynamically adjust the interval distance of the guide marker points according to the path length. The longer the path, the larger the interval distance.
[0108] On the determined optimal moving path, guide marker points are set at equal intervals. The guide marker points are used to guide the user to move along the optimal moving path in the virtual building scene. The interval distance of the guide marker points is dynamically adjusted according to the path length. The longer the path, the larger the interval distance to avoid too dense marker points; the shorter the path, the relatively smaller the interval distance.
[0109] Through a specific algorithm, the appropriate interval distance is calculated according to the path length, and then the guide marker points are set on the optimal moving path according to the interval distance. For example, in a simulated large shopping center, if the optimal moving path is long, the interval distance of the guide marker points is large; if the path is short, the interval distance is small to ensure that the guide marker points can effectively guide the user and not cause visual interference to the user.
[0110] Step S1365: Configure visual display parameters for each guide marker point, including marker shape, color, and flashing frequency. The marker shape is an arrow-shaped pointing target area to identify the direction, the color is a preset color with significant contrast with the surrounding environment, and the flashing frequency is set to a fixed period.
[0111] The visual display parameters are configured for each guide marker point. The marker shape is an arrow, and the arrow points to the direction of the target area identification, so that the user can intuitively know which direction to move. In terms of color, a preset color with significant contrast with the surrounding environment is selected, so that the guide marker point is more eye-catching in the virtual building scene and is easily noticed by the user. The flashing frequency is set to a fixed period, and the user's attention is further attracted through flashing.
[0112] For example, in a simulated large shopping center, the arrow shape of the guide marker point points to the target area, and the color may be set to bright red, and the flashing frequency is one flash per second, which can effectively guide the user to reach the target area along the optimal movement path.
[0113] Step S1366: The coordinate sequence of the optimal movement path, the position parameters of the guide marker points, and the visual display parameters are integrated into the path guide rule as a component of the spatial layout adjustment rule.
[0114] The coordinate sequence of the optimal movement path, the position parameters of the guide marker points, and the visual display parameters are integrated together to form the path guide rule. The path guide rule describes in detail how to display the optimal movement path and guide marker points in the virtual building scene. The path guide rule is used as a component of the spatial layout adjustment rule to guide the spatial layout adjustment of the virtual building scene. For example, in a simulated large shopping center, the coordinate sequence of the optimal movement path from the dining area to the entertainment area, the position and visual display parameters of the guide marker points are integrated into the path guide rule, and the guide markers are added in the virtual building scene according to the path guide rule to adjust the spatial layout and facilitate the user to find the route to the entertainment area.
[0115] Step S137: If the target type is functional combination reconstruction, the associated interactive elements of the current main focus area and the target area identification are combined and displayed according to the functional requirements associated with the predicted intention type, and an interactive element activation strategy for jointly activating the associated elements is generated.
[0116] When the target type of dynamic reconstruction is functional combination reconstruction, the associated interactive elements of the current main focus area and the target area identification are combined and displayed according to the functional requirements associated with the predicted intention type. Different predicted intention types correspond to different functional requirements, for example, if the predicted intention type is to compare goods, then the interactive elements related to the comparison of goods in the current main focus area and the target area are combined together.
[0117] The associated interactive elements are combined, and an interactive element activation strategy for jointly activating the associated elements is generated. The strategy includes setting the display mode, operation rules, etc. of the combined interactive elements, to ensure that users can conveniently use the combined elements to meet functional requirements. For example, in a simulated large shopping center, if the predicted intent type is to compare the performance of different brand mobile phones, the associated interactive elements of the mobile phone display area (the current main focus area) and the mobile phone parameter comparison area (the target area), such as mobile phone pictures, parameter lists, etc. are combined and displayed, and corresponding operation buttons are set to facilitate user comparison operations.
[0118] Step S138: Combine the spatial layout adjustment rules and the interactive element activation strategy into structured instruction data to generate dynamic reconstruction instructions for the virtual building scene.
[0119] The generated spatial layout adjustment rules and interactive element activation strategy are combined into structured instruction data. The spatial layout adjustment rules are used to adjust the spatial structure and layout of the virtual building scene, such as area scaling, adding guide markers, etc. The interactive element activation strategy is used to control the display and operation state of the interactive elements, such as improving the visibility of core interactive elements, jointly activating associated elements, etc.
[0120] The two parts of content are integrated according to a certain structure to form dynamic reconstruction instructions for the virtual building scene. The dynamic reconstruction instructions contain all the information needed for dynamic reconstruction of the virtual building scene.
[0121] Step S140: Drive the virtual building scene to perform real-time spatial reorganization according to the dynamic reconstruction instructions, and generate multi-modal feedback signals, which include visual scene update data and haptic interaction response data.
[0122] According to the generated dynamic reconstruction instructions, the virtual building scene is reorganized in real time. At the same time, in order to make users more intuitively feel the changes in the virtual building scene, multi-modal feedback signals are generated, including visual scene update data and haptic interaction response data.
[0123] Step S141: Analyze the dynamic reconstruction instructions to extract the target functional area that needs to be reorganized and the corresponding adjustment parameters.
[0124] The dynamic reconstruction instruction is parsed to extract the target functional area that needs to be reorganized and the corresponding adjustment parameters. The dynamic reconstruction instruction records in detail the adjustment operations required for each target functional area, such as the scaling factor of area scaling, the activation state of interactive elements, etc. By parsing the instruction, it is determined which functional areas need to be changed, as well as the specific adjustment method and parameters of each area. For example, in the simulation of a large shopping center, it is determined after parsing the instruction that the clothing area, dining area, etc. need to be reorganized, and the corresponding scaling factor, element display level, etc. adjustment parameters are obtained.
[0125] Step S142: Call the space reorganization engine of the virtual building scene, load the basic layout data of the target functional area, modify the position parameters in the basic layout data according to the space layout adjustment rule, including adjusting the display size of the area, updating the coordinates of the guide marker points, and modifying the relative positions of the interactive elements.
[0126] The space reorganization engine of the virtual building scene is called, which is responsible for the actual space reorganization operation of the virtual building scene. First, the basic layout data of the target functional area that needs to be reorganized is loaded, which contains the initial position and state information of the fixed structures, interactive elements, etc. in the area.
[0127] According to the space layout adjustment rule, the position parameters in the basic layout data are modified. If the space layout adjustment rule contains area scaling reconstruction, the display size of the area is adjusted according to the scaling factor to change the size of the area in the virtual building scene; if it contains path guide reconstruction, the coordinates of the guide marker points are updated to accurately display on the optimal moving path; at the same time, the relative positions of the interactive elements are modified to meet the layout requirements after adjustment. For example, in the simulation of a large shopping center, the clothing area is enlarged according to the space layout adjustment rule, the coordinates of the guide marker points are updated, and the relative positions of the clothing display racks and fitting rooms are adjusted.
[0128] Step S143: Update the state of the interactive elements in the target functional area according to the interactive element activation strategy, activate the operable attributes of the core interactive elements, set the semi-transparent display state of the non-core interactive elements, and close the response function of the irrelevant interactive elements.
[0129] According to the interactive element activation strategy, the state of the interactive elements in the target functional area is updated. For core interactive elements, activate their operable attributes to make them respond to user operations such as clicking, touching, etc. For example, in the dining area of the simulated large shopping center, if the predicted intent type is ordering, then the core interactive elements such as the menu and the order button are activated, and the user can click the menu to select dishes and click the order button to place an order.
[0130] For non-core interactive elements, set them to a semi-transparent display state, reduce their visibility in the virtual building scene, and at the same time, reduce their interactive response trigger range, only keep the basic display information and simplified interactive feedback. For irrelevant interactive elements, turn off their response function, so that they do not respond to user operations, and at the same time, set the display transparency to a preset hidden value, release the unnecessary texture resources and animation data occupied by them, and improve the performance of the virtual building scene.
[0131] For example, step S1431: Extract the identification list and priority order of core interactive elements from the interactive element activation strategy, and traverse all interactive elements in the target functional area to identify whether the identification of each interactive element is in the identification list of core interactive elements.
[0132] Extract the identification list and priority order of core interactive elements from the interactive element activation strategy. The identification list contains the unique identification of all core interactive elements determined, and the priority order indicates the importance of each core interactive element. Traverse all interactive elements in the target functional area, and check whether the identification of each interactive element is in the identification list of core interactive elements. For example, in the simulation of the electronic product area of a large shopping center, the product detail display button, the purchase button, etc. are determined as core interactive elements in the interactive element activation strategy, and their identifications are listed in the identification list. When traversing all interactive elements in the area, identify whether the identification of each element is in the identification list.
[0133] Step S1432: For interactive elements whose identifications are in the core interactive element list, enable their operable attributes, set the interactive response trigger range to the complete boundary area of the element, and load the detailed texture resources and interactive feedback animation data of the interactive element.
[0134] For interactive elements whose identifications are in the core interactive element list, enable their operable attributes so that they can respond to user operations. Set the interactive response trigger range to the complete boundary area of the element, so that the user can trigger the corresponding interactive operation by clicking or touching any part of the element. At the same time, load the detailed texture resources and interactive feedback animation data of the interactive element to make the interactive element display more realistic and vivid in the virtual building scene, and enhance the user's interactive experience. For example, in the simulation of the cosmetics area of a large shopping center, for core interactive elements such as lipstick color testing buttons, enable their operable attributes, and when the user clicks the button, the interactive operation of color testing can be triggered, and at the same time, the detailed texture of the lipstick and the color testing animation effect are loaded.
[0135] Step S1433: For the interactive elements identified not in the core interactive element list but belonging to the target functional area, determine them as non-core interactive elements, set their display transparency to a preset semi-transparent value, reduce the interactive response trigger range to the central area of the non-core interactive element, and retain the basic texture resource and simplified interactive feedback animation.
[0136] For the interactive elements identified not in the core interactive element list but belonging to the target functional area, determine them as non-core interactive elements. Set their display transparency to a preset semi-transparent value to make them less obvious in the virtual building scene and reduce the disturbance to user attention. Reduce the interactive response trigger range to the central area of the non-core interactive element, so that only when the user clicks or touches the central part of the element will limited interactive operation be triggered. Retain the basic texture resource and simplify the interactive feedback animation to maintain the basic display of the element and simple interactive feedback. For example, in the jewelry area of the simulated large shopping center, some decorative elements are determined as non-core interactive elements, set to semi-transparent display, reduce the interactive response trigger range, and only retain the basic appearance texture and simple flashing animation.
[0137] Step S1434: For the interactive elements not belonging to the target functional area, determine them as irrelevant interactive elements, turn off their interactive response function, set the display transparency to a preset hidden value, and release the unnecessary texture resources and animation data occupied by them.
[0138] For the interactive elements not belonging to the target functional area, determine them as irrelevant interactive elements. Turn off their interactive response function so that they do not respond to user operations and avoid user misoperation. Set the display transparency to a preset hidden value to completely hide them in the virtual building scene. Release the unnecessary texture resources and animation data occupied by them to reduce the resource consumption of the virtual building scene and improve the running efficiency of the system. For example, in the simulated large shopping center, when the clothing area is reorganized, the interactive elements of the catering area are determined as irrelevant interactive elements, their response function is turned off, they are hidden, and the resources occupied by them are released.
[0139] Step S1435: Record the status update results of all interactive elements, generate an element status update log, and use it as a reference basis for subsequent multi-modal feedback signal generation.
[0140] The state update results of all interactive elements are recorded, including the activation state of core interactive elements, the semi-transparent display state of non-core interactive elements, and the closed state of irrelevant interactive elements. These information is organized into an element state update log, which details the state changes of each interactive element. The element state update log serves as a reference for subsequent multi-modal feedback signal generation, providing basic information for generating accurate visual scene update data and haptic interaction response data. For example, in a simulated large shopping center, the element state update log records the state changes of each interactive element in the clothing area. When generating multi-modal feedback signals later, the visual display and haptic feedback can be adjusted based on these records.
[0141] Step S144: The spatial reorganization engine performs real-time rendering processing on the target functional area according to the modified layout data and the updated element state, generating visual scene update data containing spatial reorganization effects. The frame rate of the visual scene update data is consistent with the display refresh rate of the VR interactive terminal.
[0142] The spatial reorganization engine performs real-time rendering processing on the target functional area according to the modified layout data and the updated element state. During the rendering process, the adjusted area display size, guide marker position, and interactive element display state are integrated to generate visual scene update data containing spatial reorganization effects.
[0143] To ensure that users can smoothly experience the changes in the virtual building scene, the frame rate of the visual scene update data is consistent with the display refresh rate of the VR interactive terminal. In this way, when the user wears the VR interactive terminal, they can observe the virtual building scene undergoing real-time and smooth spatial reorganization, without experiencing picture freezing or flickering. For example, in a simulated large shopping center, after the entertainment area is reorganized, the spatial reorganization engine renders the reorganized entertainment area in real time. The visual scene update data is transmitted to the user at the same frame rate as the display refresh rate of the VR interactive terminal, allowing the user to feel the real-time changes in the entertainment area.
[0144] Step S145: According to the state of the interactive elements after spatial reorganization and the user's current behavior unit, calculate the haptic interaction response parameters, including the vibration intensity and vibration mode of the interaction point. The vibration intensity is positively correlated with the activation priority of the interactive element, and the vibration mode is one-to-one corresponding to the interactive intent type.
[0145] According to the spatially reorganized interactive element state and the user's current behavior unit, the haptic interaction response parameters are calculated. The activation priority of the interactive element reflects its importance and operability in the virtual building scene, and the vibration intensity is positively correlated with the activation priority of the interactive element. The higher the activation priority of the interactive element, the greater the vibration intensity generated by the haptic feedback module of the VR interaction terminal when the user interacts with it.
[0146] The vibration mode corresponds to the interaction intent type one by one, and different interaction intent types correspond to different vibration modes. For example, in the simulated large shopping center, if the user's interaction intent type is to click the product detail button, the corresponding vibration mode may be short and strong vibration; if it is to browse the product list, the vibration mode may be light and continuous vibration. Through a specific algorithm, the appropriate vibration intensity and vibration mode are calculated according to the interactive element state and the user behavior unit as the haptic interaction response parameters.
[0147] Step S146: Convert the haptic interaction response parameters into control signals of the VR interaction terminal haptic feedback module, align the control signals with the visual scene update data by timestamp, and generate a multi-modal feedback signal containing visual scene update data and haptic interaction response data.
[0148] The calculated haptic interaction response parameters are converted into control signals of the VR interaction terminal haptic feedback module. The haptic feedback module generates corresponding vibrations according to these control signals, allowing users to feel the interaction in the virtual building scene through touch.
[0149] The converted control signals are aligned with the visual scene update data by timestamp. The timestamp records the time when each data is generated. By aligning the timestamps, the visual scene update and the haptic interaction response are synchronized in time, allowing users to observe the changes in the virtual building scene while feeling the corresponding haptic feedback. The aligned visual scene update data and haptic interaction response data are integrated together to generate a multi-modal feedback signal containing visual scene update data and haptic interaction response data. For example, in the simulated large shopping center, when the user clicks on a core interactive element, the elements in the visual scene will have corresponding feedback display, and at the same time the haptic feedback module of the VR interaction terminal will generate vibration according to the haptic interaction response parameters. The above-mentioned synchronous feedback of vision and touch is transmitted to the user through the multi-modal feedback signal, enhancing the user's interaction experience.
[0150] Step S150: Collect the user's real-time response behavior to the multi-modal feedback signal, and optimize the prediction model of the intent evolution trajectory based on the real-time response behavior.
[0151] The real-time response behaviors of the user to the multi-modal feedback signals are collected, and the prediction model of the intention evolution trajectory is optimized by analyzing the response behaviors, so that the prediction model can more accurately predict the intention of the user, and the quality of the virtual building scene interaction experience is improved.
[0152] Step S151: The real-time response behaviors of the user after receiving the multi-modal feedback signals are continuously collected by the behavior capturing module of the VR interaction terminal, and the real-time response behaviors include new body movement trajectories, new gesture action sequences, and gaze staying positions of the user.
[0153] The real-time response behaviors of the user after receiving the multi-modal feedback signals are continuously collected by the behavior capturing module of the VR interaction terminal. The body sensors in the behavior capturing module continue to record new body movement trajectories of the user. The user may move in the virtual building scene according to the visual scene update data and the haptic interaction response data, such as walking towards the target area pointed by the guide marker. The gesture recognition component collects new gesture action sequences of the user. The user may make gestures such as clicking and selecting to interact with the interactive elements. At the same time, the gaze tracking technology is used to collect the gaze staying positions of the user to understand the key areas and elements that the user focuses on in the virtual building scene. For example, in the simulated large shopping center, after the user receives the multi-modal feedback signal about the reorganization of the clothing area, the behavior capturing module records the new body movement trajectories of the user, such as whether to walk towards the reorganized clothing display area; collects new gesture action sequences, such as whether to click the purchase button of the clothing; and records the gaze staying positions, such as whether to stay on a piece of clothing.
[0154] Step S152: Feature extraction is performed on the real-time response behaviors to obtain the degree of coincidence between the new body movement trajectories and the optimal movement path, the number of interactions between the new gesture action sequences and the core interactive elements, and the duration of the gaze staying positions on the core interactive elements.
[0155] The collected real-time response behaviors are subjected to feature extraction. For the new body movement trajectories, the degree of coincidence with the optimal movement path is calculated. The degree of coincidence can be calculated by comparing the coordinate sequence of the new body movement trajectories with the coordinate sequence of the optimal movement path to calculate the similarity between the two. For the new gesture action sequences, the number of interactions with the core interactive elements is counted to understand the frequency of user operations on the core interactive elements. For the gaze staying positions, the duration on the core interactive elements is recorded to reflect the degree of attention of the user to the core interactive elements. For example, in the simulated large shopping center, the degree of coincidence between the new body movement trajectories of the user and the optimal movement path to the fitting room is calculated; the number of times the user clicks the purchase button of the clothing and other core interactive elements is counted; and the duration of the user's gaze staying on a popular piece of clothing is recorded.
[0156] Step S153: Compare the degree of coincidence, the number of interactions, and the duration with the preset feedback evaluation threshold. When the degree of coincidence is higher than the set degree of coincidence threshold, the number of interactions reaches the set number of interactions threshold, and the duration exceeds the set duration threshold, determine that the real-time response behavior is effective positive feedback; otherwise, determine that it is invalid or negative feedback.
[0157] Compare the extracted degree of coincidence, the number of interactions, and the duration with the preset feedback evaluation threshold. Set the degree of coincidence threshold, the number of interactions threshold, and the duration threshold, which are determined according to a large amount of experimental data and user behavior analysis. When the degree of coincidence is higher than the set degree of coincidence threshold, the number of interactions reaches the set number of interactions threshold, and the duration exceeds the set duration threshold, it indicates that the user has made a positive response to the multi-modal feedback signal, and the real-time response behavior is determined to be effective positive feedback. For example, in the simulated large shopping center, if the degree of coincidence between the user's new limb movement trajectory and the optimal movement path is higher than the set threshold, the number of clicks on the core interactive element reaches the set number, and the duration of the gaze on the core interactive element exceeds the set duration, it is determined to be effective positive feedback. Conversely, if these conditions are not met, it is determined to be invalid or negative feedback.
[0158] Step S154: Based on the real-time response behavior of the effective positive feedback, extract the corresponding intention evolution node and dynamic reconstruction instruction parameter, increase the weight value of this type of intention evolution node in the prediction model, and optimize the transition probability of the intention transition network.
[0159] Based on the real-time response behavior of the effective positive feedback, extract the corresponding intention evolution node and dynamic reconstruction instruction parameter from the intention evolution trajectory and dynamic reconstruction instruction. These intention evolution nodes and parameters represent the predicted intentions and reconstruction operations that can cause the user to respond positively.
[0160] Increase the weight value of this type of intention evolution node in the prediction model, so that the prediction model is more inclined to predict these types of intentions. At the same time, optimize the transition probability of the intention transition network, adjust the size of each intention transition probability according to the situation of effective positive feedback, so that the model can more accurately reflect the transition law of user intentions. For example, in the simulated large shopping center, if the user makes effective positive feedback on the scaling reconstruction of the clothing area and the activation of the core interactive element, extract the corresponding intention evolution node and dynamic reconstruction instruction parameter, increase the weight of these intention evolution nodes, adjust the transition probability of related intentions in the intention transition network, and improve the prediction accuracy.
[0161] Step S155: Based on the real-time response behavior of invalid or negative feedback, analyze the prediction deviation reasons of the intention evolution trajectory, adjust the prediction probability calculation method of the intention evolution node, and modify the spatial layout adjustment rules or interactive element activation strategy in the dynamic reconstruction instruction.
[0162] For real-time response behavior of invalid or negative feedback, the prediction deviation reason of the intention evolution trajectory is analyzed. It may be that the predicted intention type is not accurate, or the spatial layout adjustment rule and the interaction element activation strategy in the dynamic reconstruction instruction do not meet the user's needs.
[0163] According to the analysis result, the prediction probability calculation method of the intention evolution node is adjusted, new factors are introduced or the weight distribution is adjusted, so that the prediction probability is more accurate. At the same time, the spatial layout adjustment rule or the interaction element activation strategy in the dynamic reconstruction instruction is modified, for example, the region scaling coefficient is adjusted, the selection of the core interaction element is changed, etc., so as to improve the interaction effect of the virtual building scene. For example, in the simulated large shopping center, if the user is not interested in the setting of the guide marker point, the interval distance and the visual display parameter of the guide marker point are adjusted after analyzing the reason, and the path guiding rule in the dynamic reconstruction instruction is modified.
[0164] Step S156: Update the optimized weight value, transition probability and calculation method to the prediction model of the intention evolution trajectory, verify the optimization effect through a new round of virtual-real behavior mapping data collection and intention evolution trajectory construction.
[0165] The optimized weight value, transition probability and calculation method are updated to the prediction model of the intention evolution trajectory. The updated prediction model can more accurately predict the user's intention.
[0166] Through a new round of virtual-real behavior mapping data collection and intention evolution trajectory construction, the continuous behavior sequence of the user in the physical space is collected again, the virtual-real behavior mapping data is generated, and the intention evolution trajectory is constructed. The new intention evolution trajectory is compared with the one before optimization to verify the optimization effect. If the optimized prediction model can more accurately predict the user's intention, and the user's interaction response to the virtual building scene is more positive, it means that the optimization has achieved good results; otherwise, further analysis is needed to continue optimization. For example, in the simulated large shopping center, after optimization, the behavior data of the user is collected again, the intention evolution trajectory is constructed, and whether the user's response behavior has improved is observed to verify the optimization effect of the prediction model.
[0167] Figure 2 A schematic diagram of exemplary hardware and software components of the intelligent building system interaction experience system 100 combining VR according to some embodiments of the present application is shown. For example, the processor 120 can be used in the intelligent building system interaction experience system 100 combining VR, and used to execute the functions in the present application.
[0168] The VR-based intelligent building system interactive experience system 100 can be a general server or a special-purpose server, both of which can be used to implement the VR-based intelligent building system interactive experience method of the present application. The present application only shows one server, but for the sake of convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0169] For example, the VR-based intelligent building system interactive experience system 100 can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, a ROM, or a RAM, or any combination thereof. The VR-based intelligent building system interactive experience system 100 can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof, for example. The method of the present application can be implemented according to these program instructions. The VR-based intelligent building system interactive experience system 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0170] For the sake of illustration, only one processor is described in the VR-based intelligent building system interactive experience system 100. However, it should be noted that the VR-based intelligent building system interactive experience system 100 in the present application can also include multiple processors, so the steps performed by one processor described in the present application can also be jointly performed or individually performed by multiple processors. For example, if the processor of the VR-based intelligent building system interactive experience system 100 performs steps A and B, it should be understood that steps A and B can also be jointly performed by two different processors or individually performed in one processor. For example, a first processor performs step A, a second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0171] In addition, the present application also provides a readable storage medium, wherein computer executable instructions are pre-set in the readable storage medium, and when a processor executes the computer executable instructions, the VR-based intelligent building system interactive experience method is implemented.
[0172] It should be noted that, in order to simplify the description of the present application and to help understand one or more embodiments of the present application, in the foregoing description of the embodiments of the present application, various features are sometimes combined into one embodiment, drawing, or description thereof.
Claims
1. A method for interactive experience of an intelligent building system combining VR, characterized in that, The method includes: The continuous sequence of user behavior in physical space is collected through VR interactive terminals to generate virtual behavior mapping data, which includes the correlation between user physical behavior and functional areas of virtual building scene; Based on the virtual behavior mapping data, a user's intent evolution trajectory in the virtual building scene is constructed. The intent evolution trajectory includes the intent change trend in the time dimension and the regional attention path in the spatial dimension. Based on the intent evolution trajectory, a dynamic reconstruction instruction for the virtual building scene is generated, which includes spatial layout adjustment rules and interactive element activation strategies. The virtual building scene is driven to perform real-time spatial reorganization according to the dynamic reconstruction instructions, and multimodal feedback signals are generated. The multimodal feedback signals include visual scene update data and tactile interaction response data. Collect users' real-time response behavior to the multimodal feedback signals, and optimize the prediction model of intent evolution trajectory based on the real-time response behavior; The dynamic reconstruction instruction for generating a virtual building scene based on the intent evolution trajectory includes: The intent evolution trajectory is analyzed to extract the virtual region identifier of the current main focus area, the predicted intent type of the intent evolution node, and the target region identifier; Retrieve the basic layout data corresponding to the current main focus area and target area identifiers from the area configuration library of the virtual building scene. The basic layout data includes the position parameters of fixed structures within the area, the initial state of interactive elements, and spatial connection relationships. Based on the predicted intent type matching preset reconstruction rule base, the target type of dynamic reconstruction is determined. The target type includes area scaling reconstruction, element priority reconstruction, path guidance reconstruction and function combination reconstruction. If the target type is regional scaling and reconstruction, the scaling factor is calculated based on the intent confidence of the current main focus area, and a spatial layout adjustment rule is generated to adjust the display size of the area according to the scaling factor. The higher the intent confidence, the larger the scaling factor. If the target type is element priority reconstruction, then extract the core interactive elements in the area according to the predicted intent type, set the display level of the core interactive elements to be higher than other elements, and generate an interactive element activation strategy to improve the visibility of the core interactive elements. If the target type is path-guided reconstruction, then based on the spatial relationship between the current main focus area and the target area identifier, the optimal movement path is planned, and spatial layout adjustment rules are generated to add visual guidance marks on the optimal movement path. If the target type is functional combination reconstruction, then based on the functional requirements associated with the predicted intent type, the associated interactive elements of the current main focus area and the target area identifier are combined and displayed to generate an interactive element activation strategy that jointly activates associated elements. The spatial layout adjustment rules and interactive element activation strategies are combined into structured instruction data to generate dynamic reconstruction instructions for the virtual building scene.
2. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The step of collecting continuous behavioral sequences of users in physical space through a VR interactive terminal to generate virtual behavior mapping data includes: The VR interactive terminal's behavior capture module is activated, and the built-in limb sensor and gesture recognition component are activated. The limb sensor is used to collect the user's limb movement trajectory data, and the gesture recognition component is used to collect the user's gesture action sequence data. The limb movement trajectory data is processed by spatial coordinate transformation to map the trajectory data in the physical space coordinate system to a three-dimensional movement path in the virtual building scene coordinate system. The three-dimensional movement path contains the user's real-time position coordinate sequence in the virtual building scene. The gesture sequence data is processed by keyframe extraction to identify the start frame, peak frame and end frame of the gesture, extract the gesture contour features and motion vectors of each keyframe, and determine the type and execution force of the gesture. The three-dimensional movement path and gesture action type are aligned according to timestamps to generate behavior units containing spatial location information and action feature information. Each behavior unit corresponds to a complete physical behavior segment. The behavioral units are processed by sequence integration and combined in chronological order to form a continuous sequence of user behavior in physical space. Each behavioral unit in the continuous sequence of behavior is associated with a corresponding time stamp and action feature parameters. The continuous behavior sequence is matched with the functional area division data of the virtual building scene to determine the virtual functional area identifier corresponding to each behavior unit. The interaction intent probability of the behavior unit is calculated by combining the gesture action type, and virtual behavior mapping data containing virtual area identifier, interaction intent probability and time stamp is generated.
3. The method for interactive experience of intelligent building systems combined with VR according to claim 2, characterized in that, The keyframe extraction process for the gesture sequence data includes identifying the start frame, peak frame, and end frame of the gesture, extracting the gesture contour features and motion vectors of each keyframe, and determining the type and intensity of the gesture. The inter-frame difference calculation is performed on the gesture action sequence data to obtain the pixel change between adjacent frames, and the frame where the pixel change first exceeds a preset threshold is marked as the starting frame of the gesture action. In the frame sequence after the starting frame, the cumulative value of pixel change is continuously calculated. When the cumulative value reaches the maximum value, the corresponding frame is marked as the peak frame of the gesture action. In the frame sequence following the peak frame, when the pixel change is below a preset threshold for multiple consecutive frames, the last frame with a change below the threshold is marked as the end frame of the gesture action. The gesture image data of the start frame, peak frame and end frame are extracted, the gesture image data is processed by edge detection, the set of boundary point coordinates of the gesture contour is obtained, and the set of boundary point coordinates is converted into a contour feature vector as the gesture contour feature. Calculate the pixel displacement between the peak frame and the start frame to generate the motion direction vector of the gesture. Combine the pixel displacement between the peak frame and the end frame to calculate the magnitude of the motion direction vector as the motion amplitude of the gesture. The gesture contour features are compared with a preset gesture template library to determine the type of gesture action. The motion amplitude is then normalized and used as the execution force parameter of the gesture action.
4. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The construction of the user's intent evolution trajectory in the virtual building scene based on the virtual behavior mapping data includes: The virtual behavior mapping data is parsed, and the virtual area identifier, interaction intent probability, and time stamp of each behavior unit are extracted and arranged in the order of time stamps to form a sequence of behavior units; The frequency and duration of each virtual region identifier in the behavioral unit sequence are statistically analyzed, and the virtual region identifier with the highest frequency and longest duration is determined as the user's current main focus area; The interaction intent probabilities in the sequence of behavioral units are weighted and calculated, with time stamp as the weighting factor. The intent confidence of each behavioral unit is calculated, and behavioral units with intent confidence higher than a preset threshold are selected as valid intent units. Analyze the changes in virtual region identifiers and interaction intent probabilities between adjacent valid intent units, calculate intent transfer probabilities, and construct an intent transfer network, where nodes in the intent transfer network represent interaction intent types and edges represent intent transfer probabilities. Based on the intent transfer network and the current main focus area, the possible type of user's next interaction intent and the corresponding virtual area are predicted, and an intent evolution node is generated. The intent evolution node includes the predicted intent type, target area identifier and predicted probability. The effective intent unit sequence and intent evolution nodes are integrated in chronological order to form an intent evolution trajectory that includes historical intent trajectories and future prediction directions. Each node in the intent evolution trajectory is associated with a corresponding timestamp, virtual region identifier, and intent confidence level.
5. The interactive experience method for intelligent building systems combined with VR according to claim 4, characterized in that, The analysis of changes in virtual region identifiers and interaction intent probabilities between adjacent valid intent units, calculation of intent transfer probabilities, and construction of an intent transfer network includes: Extract two adjacent valid intent units from the valid intent unit sequence, denoted as the preceding intent unit and the following intent unit, and obtain the interaction intent type of the preceding intent unit and the interaction intent type of the following intent unit; The frequency of all combinations of preceding and subsequent intent types in the effective intent unit sequence is counted, and the proportion of the frequency of each combination to the total frequency of combinations is calculated as the intent transfer probability. An initial intent transfer network is constructed using interaction intent types as network nodes and intent transfer probabilities as the weights of directed edges between nodes. Analyze the changes in virtual region identifiers in the sequence of valid intent units. When the virtual region identifiers of adjacent valid intent units are different, adjust the weight value of the corresponding intent transfer probability and increase the probability weight of cross-region intent transfer. When the change in the interaction intent probability of adjacent valid intent units exceeds a preset threshold, the intent transfer probability is adjusted a second time, increasing the transfer weight of intents with significant probability changes. The adjusted intent transfer probabilities are updated to the initial intent transfer network to generate an intent transfer network that includes regional association features.
6. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, If the target type is path-guided reconstruction, then based on the spatial relationship between the current primary focus area and the target area identifiers, the optimal movement path is planned, and spatial layout adjustment rules for adding visual guidance markers on the path are generated, including: Extract the center coordinates of the current main area of interest and the center coordinates of the target area identifier from the spatial map data of the virtual building scene, and calculate the straight-line distance and relative direction between the two center coordinates; Based on obstacle distribution data in a virtual building scene, obstacle avoidance processing is performed on the straight-line distance to generate multiple candidate movement paths. The candidate movement paths are all continuous paths that connect the center coordinates of the current main focus area and the center coordinates of the target area identifier without passing through obstacles. Calculate the path length and number of turns for each candidate movement path, and determine the candidate movement path with the highest priority as the optimal movement path. The shorter the path length and the fewer the number of turns, the higher the priority of the candidate movement path. Guide markers are set at equal intervals along the optimal movement path. The interval between the guide markers is dynamically adjusted according to the path length, with the interval increasing as the path length increases. Visual display parameters are configured for each guide marker point, including marker shape, color, and flashing frequency. The marker shape is an arrow pointing in the direction of the target area, the color is a preset color with significant contrast to the surrounding environment, and the flashing frequency is set to a fixed period. The coordinate sequence of the optimal movement path, the position parameters of the guide markers, and the visual display parameters are integrated into a path guidance rule, which is used as a component of the spatial layout adjustment rule.
7. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The process of driving the virtual building scene to perform real-time spatial reorganization according to the dynamic reconstruction instructions and generating multimodal feedback signals includes: The dynamic reconstruction instruction is analyzed to extract the spatial layout adjustment rules and interactive element activation strategies, and to determine the target functional areas that need to be reorganized and their corresponding adjustment parameters. The virtual building scene's spatial reorganization engine is invoked to load the basic layout data of the target functional area. Based on the spatial layout adjustment rules, the position parameters in the basic layout data are modified, including adjusting the area display size, updating the coordinates of the guide marker points, and modifying the relative positions of interactive elements. Based on the interaction element activation strategy, update the state of the interaction elements in the target functional area, activate the operable attributes of the core interaction elements, set the semi-transparent display state of the non-core interaction elements, and disable the response function of irrelevant interaction elements. The spatial reorganization engine performs real-time rendering processing on the target functional area based on the modified layout data and the updated element status, generating visual scene update data containing spatial reorganization effects. The frame rate of the visual scene update data is consistent with the display refresh rate of the VR interactive terminal. Based on the state of the interactive elements after spatial reorganization and the user's current behavioral unit, tactile interaction response parameters are calculated, including the vibration intensity and vibration mode of the interaction point. The vibration intensity is positively correlated with the activation priority of the interactive element, and the vibration mode corresponds one-to-one with the interaction intent type. The tactile interaction response parameters are converted into control signals for the tactile feedback module of the VR interactive terminal, and aligned with the visual scene update data by timestamp to generate a multimodal feedback signal containing both visual scene update data and tactile interaction response data.
8. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The process of collecting users' real-time response behavior to the multimodal feedback signals and optimizing the prediction model of intent evolution trajectory based on the real-time response behavior includes: The VR interactive terminal continuously collects the user's real-time response behavior after receiving multimodal feedback signals through the behavior capture module. The real-time response behavior includes the user's new limb movement trajectory, new gesture sequence, and gaze position. Feature extraction is performed on the real-time response behavior to obtain the degree of consistency between the new limb movement trajectory and the optimal movement path, the number of interactions between the new gesture sequence and the core interactive element, and the duration of the gaze position on the core interactive element. The matching degree, number of interactions, and duration are compared with preset feedback evaluation thresholds. When the matching degree is higher than the set matching degree threshold, the number of interactions reaches the set number threshold, and the duration exceeds the set duration threshold, the real-time response behavior is determined to be valid positive feedback; otherwise, it is determined to be invalid or negative feedback. Based on the real-time response behavior with effective positive feedback, the corresponding intent evolution nodes and dynamic reconstruction instruction parameters are extracted, the weight values of such intent evolution nodes in the prediction model are increased, and the transfer probability of the intent transfer network is optimized. Based on the real-time response behavior of invalid or negative feedback, analyze the reasons for the prediction deviation of the intent evolution trajectory, adjust the prediction probability calculation method of intent evolution nodes, and modify the spatial layout adjustment rules or interactive element activation strategies in the dynamic reconstruction instructions. The optimized weight values, transition probabilities, and calculation methods are updated to the prediction model of the intent evolution trajectory. The optimization effect is verified through a new round of virtual behavior mapping data collection and intent evolution trajectory construction.
9. An intelligent building system interactive experience system incorporating VR, characterized in that, The system includes a processor and a memory, the memory being connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the VR-integrated intelligent building system interactive experience method as described in any one of claims 1-8.
Citation Information
Patent Citations
User-interactivity enabled search filter tool optimized for virtualized worlds
US12159364B2
Systems and methods for language-based three-dimensional interactive environment construction and interaction
WO2025024353A2