Interactive experience method and system for intelligent building system in combination with VR
By collecting user behavior data to generate virtual-real mapping data, constructing intent evolution trajectories and dynamically reconstructing virtual building scenes, the problem of insufficient intuitiveness and immersion in traditional building system interaction methods is solved, realizing a personalized and intelligent interactive experience.
Patent Information
- Application Number
- CN202510944994.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing building system interaction methods lack intuitiveness and immersion, making it difficult for users to fully understand the functions of the devices. Traditional VR applications cannot deeply interact with virtual scenes, resulting in a less rich and personalized user experience.
By collecting user behavior sequences through VR interactive terminals, virtual behavior mapping data is generated, an intent evolution trajectory is constructed, dynamic reconstruction instructions are generated, real-time spatial reorganization and multimodal feedback of virtual scenes are realized, and the intent prediction model is optimized.
It achieves precise matching between virtual building scenes and user intentions, providing a personalized and intelligent interactive experience, and improving user satisfaction and interaction efficiency.
Smart Images

Figure CN120848725A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual reality technology, and more specifically, to an intelligent building system interactive experience method and system that combines VR. Background Technology
[0002] In the current field of building system interaction experience, traditional interaction methods mainly rely on fixed facilities in physical spaces and simple user interfaces. For example, users control building equipment such as lights, air conditioning, and access control by operating physical buttons, touch screens, or using remote controls. These interaction methods have several limitations. On the one hand, users need to become familiar with the operation of various devices, increasing the learning curve; on the other hand, the interaction process lacks intuitiveness and immersion, making it difficult for users to fully and deeply understand the functions and operating status of the building system.
[0003] With the development of virtual reality (VR) technology, although some cases of applying VR to building displays have emerged, most of these applications simply present the physical structure of buildings in a virtual way. Users can only passively watch the virtual scene and cannot interact deeply with it. Furthermore, the virtual scene cannot be dynamically adjusted according to the user's real-time behavior and intentions, resulting in a less rich and personalized user experience that fails to meet the growing demand for intelligent and interactive building experiences. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for interactive experience of an intelligent building system incorporating VR, the method comprising:
[0005] The continuous sequence of user behavior in physical space is collected through VR interactive terminals to generate virtual behavior mapping data, which includes the correlation between user physical behavior and functional areas of virtual building scene;
[0006] Based on the virtual behavior mapping data, a user's intent evolution trajectory in the virtual building scene is constructed. The intent evolution trajectory includes the intent change trend in the time dimension and the regional attention path in the spatial dimension.
[0007] Based on the intent evolution trajectory, a dynamic reconstruction instruction for the virtual building scene is generated, which includes spatial layout adjustment rules and interactive element activation strategies.
[0008] The virtual building scene is driven to perform real-time spatial reorganization according to the dynamic reconstruction instructions, and multimodal feedback signals are generated. The multimodal feedback signals include visual scene update data and tactile interaction response data.
[0009] The system collects real-time response behavior of users to the multimodal feedback signals and optimizes the prediction model of intent evolution trajectory based on the real-time response behavior.
[0010] In another aspect, embodiments of the present invention also provide an intelligent building system interactive experience system combined with VR, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or code. The processor is used to execute the programs, instructions or code in the machine-readable storage medium to implement the above-described method.
[0011] Based on the above, this invention generates virtual behavior mapping data by collecting continuous behavioral sequences of users in physical space. This enables precise establishment of the association between user physical behavior and functional areas of virtual building scenes. Based on the virtual behavior mapping data, an intent evolution trajectory is constructed, comprehensively considering the temporal trend of intent changes and the spatial focus path, allowing for a comprehensive and dynamic understanding of user intent. Dynamic reconstruction instructions are generated based on the intent evolution trajectory, enabling real-time spatial reorganization of the virtual building scene according to user intent, making the virtual scene more aligned with user needs. Multimodal feedback signals containing visual scene update data and tactile interaction response data are generated, providing users with a rich sensory experience and enhancing the realism and immersion of the interaction. Collecting real-time user responses to the multimodal feedback signals and optimizing the prediction model of the intent evolution trajectory continuously improves the system's accuracy in predicting user intent, thereby providing users with a more personalized and intelligent building system interaction experience, significantly improving user satisfaction and interaction efficiency. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the execution flow of the intelligent building system interactive experience method combined with VR provided in the embodiment of the present invention.
[0013] Figure 2 This is a schematic diagram of exemplary hardware and software components of the VR-integrated intelligent building system interactive experience system provided in the embodiments of the present invention. Detailed Implementation
[0014] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating an intelligent building system interactive experience method incorporating VR, according to an embodiment of the present invention. The following is a detailed description of this intelligent building system interactive experience method incorporating VR.
[0015] Step S110: Collect the user's continuous behavior sequence in the physical space through the VR interactive terminal to generate virtual behavior mapping data. The virtual behavior mapping data includes the correlation between the user's physical behavior and the functional areas of the virtual building scene.
[0016] In the interactive experience scenario of an intelligent building system incorporating VR, users wear VR interactive terminals to enter a physical space corresponding to the virtual building scene, such as a simulated large shopping mall. Once the VR interactive terminal is turned on, it begins to collect the user's continuous behavioral sequence within that physical space. Various sensors and components inside the VR interactive terminal work together to collect user behavioral information from all angles, in order to build a correlation between the user's physical behavior and the functional areas of the virtual building scene.
[0017] Step S111: Start the behavior capture module of the VR interactive terminal and activate the built-in limb sensor and gesture recognition component. The limb sensor is used to collect the user's limb movement trajectory data, and the gesture recognition component is used to collect the user's gesture action sequence data.
[0018] When a user puts on the VR interactive terminal, the terminal automatically triggers the activation of the behavior capture module. Among these, the body sensors can detect changes in the position of various parts of the user's body in physical space. In a simulated large shopping mall space, a user might walk between store aisles, enter stores to browse merchandise, etc. The body sensors continuously record the user's limb movement trajectories in the physical space, forming body movement trajectory data. This data reflects the user's positional movement within the physical space.
[0019] At the same time, the gesture recognition component is also activated, using special sensing technology and image processing capabilities to capture the user's hand movements. In a shopping mall, users may point to a store sign or make gestures to select products; the gesture recognition component will record these gestures in chronological order, forming gesture sequence data.
[0020] Step S112: Perform spatial coordinate transformation processing on the limb movement trajectory data to map the trajectory data in the physical space coordinate system to a three-dimensional movement path in the virtual building scene coordinate system. The three-dimensional movement path contains the user's real-time position coordinate sequence in the virtual building scene.
[0021] The collected limb movement trajectory data was initially based on a physical space coordinate system. To accurately represent the user's movement in the virtual building scene, a spatial coordinate transformation is required. First, the correspondence between the physical space coordinate system and the virtual building scene coordinate system is determined. This involves calibrating the origin, coordinate axis directions, and scale of the two coordinate systems. Through a series of position mapping operations, each point on the user's limb movement trajectory in physical space is transformed into the coordinate system of the virtual building scene.
[0022] For example, in a simulated large shopping mall physical space, a user moves from an entrance to a food court, corresponding to a series of coordinate points in the physical space coordinate system. After coordinate transformation, these coordinate points form a new coordinate sequence in the virtual building scene coordinate system, constituting the user's three-dimensional movement path within the virtual building scene. This three-dimensional movement path contains the user's real-time position coordinate sequence within the virtual building scene, accurately reflecting the user's movement trajectory in the virtual space.
[0023] Step S113: Perform keyframe extraction processing on the gesture action sequence data, identify the start frame, peak frame and end frame of the gesture action, extract the gesture contour features and motion vectors of each key frame, and determine the type and execution force of the gesture action.
[0024] For the acquired gesture sequence data, keyframe extraction processing is required. First, each frame of the gesture image is analyzed, and keyframes are identified by comparing the differences between adjacent frames. When the difference between adjacent frames first reaches a certain level, that frame is determined as the starting frame of the gesture. After the starting frame, the cumulative difference between subsequent frames is continuously analyzed. When the cumulative difference reaches its maximum value, the corresponding frame is the peak frame. After the peak frame, when the difference between adjacent frames falls below a certain level multiple times consecutively, the last frame below the difference threshold is determined as the ending frame.
[0025] After determining the keyframes, the gesture images of the keyframes are processed. For gesture contour feature extraction, edge detection is used. By detecting abrupt changes in pixel values in the gesture image, the boundary points of the gesture are identified, and the coordinate information of these boundary points is integrated to form a feature description of the gesture contour. The motion vector is calculated by comparing the changes in gesture position between the peak frame and the start frame, as well as between the peak frame and the end frame. Comparing the positional offset of the gesture between the peak frame and the start frame determines the direction of the gesture movement; then, combining the positional offset between the peak frame and the end frame, the amplitude of the movement is calculated.
[0026] The extracted gesture contour features are compared with a pre-defined gesture template library. This library contains feature descriptions of various common gestures; by matching similarity, the type of the current gesture is determined. For the execution force of the gesture, the amplitude of the movement is normalized to fall within a specific range, and this normalization is used as the execution force parameter for the gesture.
[0027] Step S1131: Perform inter-frame difference calculation on the gesture action sequence data to obtain the pixel change between adjacent frames, and mark the frame where the pixel change first exceeds a preset threshold as the starting frame of the gesture action.
[0028] When processing gesture sequence data, inter-frame difference calculation is performed on two adjacent gesture images. By comparing the pixel value differences of corresponding pixels in the two images, the pixel change between each pair of adjacent frames is obtained. A preset threshold is set; when the pixel change between adjacent frames first exceeds this threshold, that frame is marked as the starting frame of the gesture. For example, in a simulated large shopping mall, when a user raises their hand to point to a store, the frame corresponding to the first time the pixel change between adjacent frames reaches the preset change level in the gesture sequence data is the starting frame.
[0029] Step S1132: In the frame sequence after the starting frame, continuously calculate the cumulative value of pixel change. When the cumulative value reaches the maximum value, the corresponding frame is marked as the peak frame of the gesture action.
[0030] Starting from the initial frame, the subsequent frame sequence continues to be processed. The cumulative value of pixel changes between adjacent frames is continuously calculated. As the gesture progresses, the pixel changes accumulate. When the cumulative value reaches its maximum value, it indicates that the gesture has reached a peak in intensity or amplitude; the corresponding frame is marked as the peak frame of the gesture. For example, when a user raises their finger towards a store and gradually reaches its maximum amplitude, the corresponding frame is the peak frame.
[0031] Step S1133: In the frame sequence after the peak frame, when the pixel change amount is lower than the preset threshold for multiple consecutive frames, mark the last frame with a change amount lower than the threshold as the end frame of the gesture action.
[0032] After identifying the peak frame, the subsequent frame sequence is monitored. When the pixel change between adjacent frames is below a preset threshold for several consecutive frames, it indicates that the gesture is nearing its end. The last frame with a pixel change below the threshold is marked as the end frame of the gesture. For example, when a user raises their hand to point at a store and begins to lower it, the last frame is the end frame when the pixel change between adjacent frames is below the preset threshold multiple times.
[0033] Step S1134: Extract gesture image data from the start frame, peak frame, and end frame; perform edge detection processing on the gesture image data to obtain the set of boundary point coordinates of the gesture contour; and convert the set of boundary point coordinates into a contour feature vector as the gesture contour feature.
[0034] After extracting gesture image data from the start frame, peak frame, and end frame, edge detection is performed on these images. Edge detection identifies the boundaries of the gesture by recognizing abrupt changes in pixel values within the image. After detecting the boundary points of the gesture contour, the coordinate information of these boundary points is collected to form a boundary point coordinate set. Then, feature transformation is performed on this boundary point coordinate set, converting it into a contour feature vector with specific dimensions and format. This contour feature vector serves as the gesture contour feature, which accurately describes the shape and contour information of the gesture.
[0035] Step S1135: Calculate the pixel displacement between the peak frame and the start frame, generate the motion direction vector of the gesture, and combine the pixel displacement between the peak frame and the end frame to calculate the magnitude of the motion direction vector as the motion amplitude of the gesture.
[0036] By comparing the position of the gesture in the peak frame and the starting frame, the pixel displacement between them is calculated. This displacement is used to determine the direction of the gesture, thus generating a motion direction vector. Simultaneously, by combining the pixel displacement between the peak frame and the ending frame, the magnitude of the motion direction vector is further calculated. This magnitude represents the amplitude of the gesture throughout the entire process, reflecting the force and range of the gesture.
[0037] Step S1136: Compare the gesture contour features with a preset gesture template library to determine the type of gesture action, and normalize the motion amplitude as the execution force parameter of the gesture action.
[0038] The extracted gesture contour features are compared with a pre-defined gesture template library. This library stores feature descriptions of various common gestures. By calculating the similarity between the gesture contour features and the features of each template in the library, the best-matching template is found, thus determining the type of gesture. For example, it might be identified as a "pointing" or "clicking" gesture.
[0039] The calculated motion amplitude is then normalized. The motion amplitude is adjusted to a specific range to ensure comparability and standardization. The normalized motion amplitude, used as a parameter of the execution force of the gesture, accurately reflects the intensity of the gesture.
[0040] Step S114: Align the three-dimensional movement path and gesture action type according to the timestamp to generate a behavior unit containing spatial location information and action feature information. Each behavior unit corresponds to a complete physical behavior segment.
[0041] The previously obtained 3D movement paths and gesture types are aligned according to timestamps. Each timestamp records the moment of data acquisition, and these timestamps are used to associate the position information on the 3D movement path with the corresponding gesture type. In this way, at the same point in time, we have the user's position information and corresponding gesture type in the virtual building scene, forming a behavioral unit.
[0042] Each behavioral unit corresponds to a complete physical behavior segment. For example, in a simulated large shopping mall, a user walks from a rest area to a clothing store while making a gesture pointing to the store sign. The corresponding 3D movement path and gesture type, when aligned with timestamps, constitute a behavioral unit that fully describes the user's physical behavior during that time period.
[0043] Step S115: Perform sequence integration processing on the behavioral units and combine them in chronological order to form a continuous behavioral sequence of the user in the physical space. Each behavioral unit in the continuous behavioral sequence is associated with a corresponding time stamp and action feature parameters.
[0044] The generated behavioral units are then sequenced and combined. Each behavioral unit is arranged and combined in chronological order. Each behavioral unit is associated with a corresponding timestamp and action feature parameters. The timestamp accurately records the moment the behavior occurred, while the action feature parameters include information such as the type of gesture and the intensity of the action.
[0045] In a simulated large shopping mall, a user might perform multiple consecutive actions, such as first walking to the food court and gesturing to view the menu, then walking to the entertainment area and gesturing to select an entertainment activity. Combining these behavioral units in chronological order creates a sequence of continuous user behavior in the physical space, clearly presenting the user's behavioral process over a period of time.
[0046] Step S116: Match the continuous behavior sequence with the functional area division data of the virtual building scene, determine the virtual functional area identifier corresponding to each behavior unit, calculate the interaction intent probability of the behavior unit in combination with the gesture action type, and generate virtual behavior mapping data containing virtual area identifier, interaction intent probability and time stamp.
[0047] The generated continuous behavior sequences are matched with the functional area division data of the virtual building scene. The virtual building scene is divided into different functional areas, such as shopping areas, dining areas, and rest areas, each with a corresponding identifier. By comparing the user's location information in the behavior unit with the location range of the functional area, the virtual functional area identifier corresponding to each behavior unit is determined.
[0048] The probability of interaction intent for each behavioral unit is calculated by combining the type of gesture. Different gesture types reflect different user interaction intents. For example, pointing to a shop sign may indicate an intention to enter the shop, while making a gesture to order food may indicate an intention to consume food. Based on preset rules and statistical data, the probability of interaction intent for each behavioral unit is calculated.
[0049] By integrating virtual area identifiers, interaction intent probabilities, and time stamps, virtual behavior mapping data is generated. This virtual behavior mapping data contains the correlation between user physical behavior and functional areas of the virtual building scene.
[0050] Step S120: Construct the user's intent evolution trajectory in the virtual building scene based on the virtual behavior mapping data. The intent evolution trajectory includes the intent change trend in the time dimension and the regional attention path in the spatial dimension.
[0051] After obtaining the virtual behavior mapping data, a user's intent evolution trajectory in the virtual building scenario is constructed based on this data. The intent evolution trajectory can comprehensively reflect the changes in user intent at different times and in different spaces, which is of great significance for understanding user needs and optimizing the virtual building scenario.
[0052] Step S121: Parse the virtual behavior mapping data, extract the virtual region identifier, interaction intent probability and time stamp of each behavior unit, and arrange them in the order of time stamp to form a behavior unit sequence.
[0053] The virtual behavior mapping data is parsed to extract the virtual area identifier, interaction intent probability, and time stamp for each behavioral unit. The virtual area identifier indicates the functional area in the virtual building scene where the user's behavior occurs, the interaction intent probability reflects the likelihood of the user's interaction intent under that behavior, and the time stamp records the time when the behavior occurs.
[0054] These behavioral units are arranged in chronological order to form a behavioral unit sequence. In a simulated large shopping mall, this behavioral unit sequence clearly shows which functional areas a user is in at different times, and the corresponding probability of their interaction intent.
[0055] Step S122: Count the frequency and duration of each virtual region identifier in the behavior unit sequence, and determine the virtual region identifier with the highest frequency and the longest duration as the user's current main focus area.
[0056] Statistical analysis is performed on the behavioral unit sequences. The frequency of each virtual region identifier appearing in the sequence is counted, representing the number of times the user performs an action within that virtual region. Simultaneously, the duration of the user's activity in each virtual region is calculated, and the user's dwell time in that region is determined by comparing the time stamps of the behavioral units.
[0057] By comprehensively comparing frequency and duration, the virtual area markers with the highest frequency and longest duration are identified. In a simulated large shopping mall, if a user appears most frequently and stays in a certain clothing section the longest, then the virtual area marker for that clothing section is identified as the user's current primary focus area, reflecting the user's current main interests.
[0058] Step S123: Perform a weighted calculation on the interaction intent probabilities in the sequence of behavioral units, using time stamps as weighting factors, calculate the intent confidence of each behavioral unit, and select behavioral units with intent confidence higher than a preset threshold as valid intent units.
[0059] The interaction intent probabilities in the sequence of behavioral units are weighted and calculated. Time stamps are used as weighting factors because more recent behaviors are more likely to reflect the user's current intent. Using a specific weighting method, combining the interaction intent probability and time stamps, the intent confidence of each behavioral unit is calculated.
[0060] A preset threshold is set to filter the calculated intent confidence scores. Behavioral units with intent confidence scores higher than the preset threshold are selected as valid intent units. These valid intent units are more representative of the user's true intent.
[0061] Step S124: Analyze the changes in virtual region identifiers and interaction intent probabilities between adjacent valid intent units, calculate intent transfer probabilities, and construct an intent transfer network, where nodes in the intent transfer network represent interaction intent types and edges represent intent transfer probabilities.
[0062] The selected valid intent units are analyzed. Changes in virtual region identifiers and interaction intent probabilities are compared between adjacent valid intent units. Changes in virtual region identifiers reflect user movement between different functional areas, while changes in interaction intent probabilities reflect shifts in user intent.
[0063] Based on these changes, the intent transition probability is calculated. The intent transition probability represents the likelihood of transitioning from one interaction intent type to another. Using interaction intent types as nodes and intent transition probabilities as the weights of directed edges between nodes, an intent transition network is constructed. This intent transition network clearly demonstrates the transition relationships between user intents.
[0064] Step S1241: Extract two adjacent valid intent units from the valid intent unit sequence, denoted as the preceding intent unit and the following intent unit, and obtain the interaction intent type of the preceding intent unit and the interaction intent type of the following intent unit.
[0065] Extract two adjacent valid intent units from the valid intent unit sequence, denoted as the preceding intent unit and the following intent unit, respectively. The preceding intent unit represents the user's valid intent at an earlier time step, and the following intent unit represents the valid intent at a later time step. Obtain the interaction intent type of these two intent units; for example, the preceding intent unit might be "view product," and the following intent unit might be "purchase product."
[0066] Step S1242: Count the number of occurrences of all combinations of preceding and subsequent intent types in the effective intent unit sequence, and calculate the proportion of the occurrence frequency of each combination to the total combination frequency, as the intent transfer probability.
[0067] The system statistically analyzes all combinations of preceding and following intent types within the effective intent unit sequence. It counts the frequency of each combination and then calculates the proportion of each combination's frequency to the total frequency of all combinations. This proportion represents the intent transition probability, reflecting the likelihood of shifting from one interaction intent type to another. For example, by counting the frequency of the "view product" to "buy product" combination and the total frequency of all combinations, the ratio between the two yields the corresponding intent transition probability.
[0068] Step S1243: Construct an initial intent transfer network using the interaction intent type as the network node and the intent transfer probability as the weight of the directed edges between nodes.
[0069] The network uses interaction intent types as nodes, with each node representing a possible interaction intent. An initial intent transition network is constructed using the calculated intent transition probabilities as the weights of the directed edges between nodes. This initial intent transition network demonstrates the transition relationships and probability levels between different interaction intent types.
[0070] Step S1244: Analyze the changes in the virtual region identifiers in the sequence of valid intent units. When the virtual region identifiers of adjacent valid intent units are different, adjust the weight value of the corresponding intent transfer probability and increase the probability weight of cross-region intent transfer.
[0071] Analyze the changes in virtual region identifiers within the sequence of valid intent units. When adjacent valid intent units have different virtual region identifiers, it indicates that the user has moved between different functional areas. Such cross-regional moves may be accompanied by significant changes in intent. In this case, adjust the weight value of the corresponding intent transfer probability, increasing the probability weight of cross-regional intent transfers. For example, when a user moves from the shopping area to the dining area, the corresponding intent transfer probability weight will increase to more accurately reflect changes in user intent.
[0072] Step S1245: When the change in the interaction intent probability of adjacent valid intent units exceeds a preset threshold, the intent transfer probability is adjusted a second time to increase the transfer weight of intents with significant probability changes.
[0073] When the change in the interaction intent probability of adjacent valid intent units exceeds a preset threshold, it indicates a significant shift in the user's intent. At this point, the intent shift probability is adjusted a second time, increasing the weight of intents with significant probability changes. For example, if a user's intent probability suddenly shifts from a low "view product" to a high "buy product" probability, exceeding the preset threshold, the probability weight of this intent shift is increased to more accurately capture changes in user intent.
[0074] Step S1246: Update the adjusted intent transfer probability to the initial intent transfer network to generate an intent transfer network containing regional association features.
[0075] The adjusted intent transfer probabilities are then updated in the initial intent transfer network. By updating the weight values, the intent transfer network more accurately reflects the transfer of user intent, while also considering changes in virtual region identifiers and significant changes in interaction intent probabilities, generating an intent transfer network that incorporates region association features. This intent transfer network not only considers the transfer probabilities between interaction intent types but also incorporates the impact of virtual region identifier changes on intent transfer, making the description of user intent evolution more accurate and comprehensive.
[0076] Step S125: Based on the intent transfer network and the current main focus area, predict the possible interaction intent type of the user in the next step and the corresponding virtual area, and generate an intent evolution node. The intent evolution node includes the predicted intent type, the target area identifier and the predicted probability.
[0077] Using a constructed intent transition network and a defined current primary focus area, the system predicts the user's likely next interaction intent type and its corresponding virtual region. Within the intent transition network, starting with the interaction intent type corresponding to the current primary focus area, the probability of the user transitioning to other intent types is calculated based on the transition probabilities of each intent. Combining the association between virtual regions and interaction intent types, a target region identifier is determined for each predicted intent type.
[0078] The predicted intent type, target area identifier, and corresponding predicted probability are integrated to generate an intent evolution node. For example, in a simulated large shopping mall, if the current main focus area is the clothing area, the intent transition network predicts that the user may next intend to go to the fitting room to try on clothes. The virtual area identifier where the fitting room is located is the target area identifier. At the same time, the predicted probability is calculated to form an intent evolution node that includes the predicted intent type, target area identifier, and predicted probability.
[0079] Step S126: Integrate the effective intent unit sequence and intent evolution nodes in chronological order to form an intent evolution trajectory that includes historical intent trajectories and future prediction directions. Each node in the intent evolution trajectory is associated with a corresponding timestamp, virtual region identifier, and intent confidence.
[0080] The effective intent unit sequence and intent evolution nodes are integrated in chronological order. The effective intent unit sequence represents the user's past intent trajectory, while the intent evolution nodes reflect the prediction of the user's future intent. Arranging them in chronological order forms a complete intent evolution trajectory.
[0081] In this intent evolution trajectory, each node is associated with a corresponding timestamp, virtual region identifier, and intent confidence level. The timestamp records the moment the intent occurred or was predicted, the virtual region identifier indicates the specific functional area where the intent occurred, and the intent confidence level reflects the reliability of the intent. Through this intent evolution trajectory, the changing trend of user intent from the past to the future and the spatial dimension of regional attention path can be clearly observed.
[0082] Step S130: Generate a dynamic reconstruction instruction for the virtual building scene based on the intent evolution trajectory. The dynamic reconstruction instruction includes spatial layout adjustment rules and interactive element activation strategies.
[0083] After obtaining the user's evolving intent, dynamic reconstruction instructions for the virtual building scene are generated based on this. These dynamic reconstruction instructions aim to adjust the virtual building scene in real time according to changes in the user's intent, providing a more user-friendly interactive experience.
[0084] Step S131: Analyze the intent evolution trajectory and extract the virtual region identifier of the current main focus area, the predicted intent type of the intent evolution node, and the target region identifier.
[0085] The intent evolution trajectory is analyzed to extract key information. The virtual area identifier of the current primary focus area clearly identifies the functional area of greatest interest to the user. The predicted intent type of the intent evolution node indicates the user's possible next interaction intent, and the target area identifier corresponds to the virtual area where the predicted intent may occur. For example, in a simulated large shopping mall, after analyzing the intent evolution trajectory, it is determined that the current primary focus area is the cosmetics section, the predicted intent type is to purchase cosmetics, and the target area identifier is the checkout area.
[0086] Step S132: Retrieve the basic layout data corresponding to the current main focus area and target area identifiers from the area configuration library of the virtual building scene. The basic layout data includes the position parameters of fixed structures within the area, the initial state of interactive elements, and spatial connection relationships.
[0087] Retrieve the basic layout data corresponding to the currently primary focus area and target area identifiers from the area configuration library of the virtual building scene. The area configuration library stores detailed information about each functional area in the virtual building scene. The basic layout data includes the positional parameters of fixed structures within the area, such as the positions of walls and columns; the initial state of interactive elements, such as whether buttons are clickable and whether display shelves are visible; and spatial connection relationships, such as the location and connection method of passageways between different areas.
[0088] Step S133: Based on the predicted intent type matching of the preset reconstruction rule base, determine the target type of dynamic reconstruction. The target type includes region scaling reconstruction, element priority reconstruction, path guidance reconstruction, and function combination reconstruction.
[0089] Based on the extracted predicted intent type, it is matched against a pre-defined reconstruction rule base. The reconstruction rule base contains reconstruction rules and target types corresponding to different predicted intent types. Through matching, the target type for dynamic reconstruction is determined. Possible target types include area scaling reconstruction (adjusting the display size of an area); element priority reconstruction (changing the display and operation priority of interactive elements); path guidance reconstruction (planning a path to guide the user to the target area); and function combination reconstruction (combining interactive elements from different areas to achieve a specific function). For example, if the predicted intent type is to quickly find a product, it might match the target type of path guidance reconstruction.
[0090] Step S134: If the target type is region scaling reconstruction, calculate the scaling factor based on the intent confidence of the current main focus area, and generate a spatial layout adjustment rule to adjust the display size of the region according to the scaling factor. The higher the intent confidence, the larger the scaling factor.
[0091] When the target type of dynamic reconstruction is region scaling reconstruction, the scaling factor is calculated based on the intent confidence of the currently primary area of interest. Intent confidence reflects the user's level of attention to the currently primary area of interest and the reliability of their intent. The higher the intent confidence, the greater the user's interest in that area, and the more prominently that area needs to be displayed; therefore, the larger the scaling factor.
[0092] Using a specific calculation method, intent confidence is converted into a scaling factor. Then, a spatial layout adjustment rule is generated based on this scaling factor, which is used to adjust the display size of the currently primary area of interest. For example, in a simulated large shopping mall, if the current primary area of interest is the electronics section and the intent confidence is high, the calculated scaling factor will be large. Based on this scaling factor, the electronics section will be enlarged in the virtual building scene to attract the user's attention.
[0093] Step S135: If the target type is element priority reconstruction, then extract the core interactive elements in the area according to the predicted intent type, set the display level of the core interactive elements to be higher than other elements, and generate an interactive element activation strategy to improve the visibility of the core interactive elements.
[0094] When the target type is element priority reconstruction, core interactive elements are extracted from the region based on the predicted intent type. Core interactive elements are elements closely related to the predicted intent. For example, if the predicted intent type is to buy a product, then the product's buy button, price tag, etc., are core interactive elements.
[0095] The core interactive elements are prioritized and displayed above other elements. By adjusting the display order of these elements, the core interactive elements are made more prominent and visible. Simultaneously, activation strategies are generated to enhance the visibility of these core interactive elements, including setting their color, brightness, and other visual attributes to make them more eye-catching in the virtual building scene. For example, in a simulated clothing section of a large shopping mall, if the predicted intent is to view details of a particular garment, then the details display button, size selection box, and other core interactive elements will be placed at a higher display level and may even have their colors changed to attract user clicks.
[0096] Step S136: If the target type is path-guided reconstruction, then based on the spatial relationship between the current main focus area and the target area identifier, plan the optimal movement path and generate spatial layout adjustment rules for adding visual guidance marks on the optimal movement path.
[0097] When the target type is path-guided reconstruction, the spatial relationship between the current primary focus area and the target area identifier is analyzed first. In the virtual building scene, each area has its spatial coordinates. By comparing these coordinates, the relative position and distance between the two are determined.
[0098] Based on obstacle distribution data from a virtual building scenario, obstacle avoidance is performed on the straight-line distance directly connecting the current primary area of interest and the target area. A search algorithm, such as the A* algorithm, is used to find multiple candidate movement paths that avoid obstacles. These candidate paths are then evaluated, taking into account factors such as path length and the number of turns, and the candidate movement path with the shortest path length and the fewest turns is determined as the optimal movement path.
[0099] Visual guidance markers are added to the optimal movement path to generate spatial layout adjustment rules. These markers can be arrows, lines, etc., used to guide users from their current primary focus area to their target area within the virtual building scene. For example, in a simulated large shopping mall, if a user is currently in the food and beverage area and the system predicts they will next go to the entertainment area, it will plan an optimal movement path that avoids other shops and obstacles, and add visual guidance markers such as arrows along this path to help the user find their way to the entertainment area.
[0100] Step S1361: Extract the center coordinates of the current main area of interest and the center coordinates of the target area identifier from the spatial map data of the virtual building scene, and calculate the straight-line distance and relative direction between the two center coordinates.
[0101] The center coordinates of the current primary focus area and target area are obtained from the spatial map data of the virtual building scene. The spatial map data records detailed location information for each area within the virtual building scene. By extracting the center coordinates of these two areas, the straight-line distance between them is calculated. This distance can be determined by comparing the difference between the two coordinates using coordinate calculation methods. Simultaneously, the relative direction between the two areas, such as due east, due west, or northeast, is determined based on the relative positions of the coordinates.
[0102] Step S1362: Based on the obstacle distribution data of the virtual building scene, obstacle avoidance processing is performed on the straight-line distance to generate multiple candidate movement paths. The candidate movement paths are all continuous paths that connect the center coordinates of the current main focus area and the center coordinates of the target area identifier without passing through obstacles.
[0103] Using obstacle distribution data from a virtual building scene, the calculated straight-line distances are processed. The obstacle distribution data records the location and extent of all obstacles in the virtual building scene. When planning a movement path, these obstacles need to be avoided. A search algorithm starts from the center coordinates of the current primary area of interest and attempts to find multiple continuous paths that connect to the center coordinates of the target area without passing through obstacles. These paths are candidate movement paths; they are continuous in space and bypass all obstacles.
[0104] Step S1363: Calculate the path length and number of turns for each candidate movement path, and determine the candidate movement path with the highest priority as the optimal movement path. The shorter the path length and the fewer the number of turns, the higher the priority of the candidate movement path.
[0105] Each generated candidate movement path is evaluated. The length of each path is calculated, which is the sum of the lengths of all line segments along the path. Simultaneously, the number of turns in the path is counted; the number of turns reflects the path's complexity. Candidate movement paths with shorter lengths and fewer turns indicate that movement is more convenient and efficient for the user, and therefore have higher priority.
[0106] By comparing the path length and number of turns of each candidate movement path, the candidate movement path with the highest priority is determined as the optimal movement path. For example, in a simulated large shopping mall, there are multiple candidate movement paths from the clothing area to the checkout area. After evaluation, the path with the shortest path length and the fewest turns is selected as the optimal movement path.
[0107] Step S1364: Set up guide markers at equal intervals on the optimal movement path. The interval between the guide markers is dynamically adjusted according to the path length. The longer the path, the greater the interval.
[0108] Guide markers are placed at equal intervals along the determined optimal movement path. These markers guide the user along the optimal path within the virtual building scene. The interval between the guide markers is dynamically adjusted based on the path length. Longer paths require larger intervals to avoid overly dense marker placement; shorter paths require relatively smaller intervals.
[0109] Using a specific algorithm, an appropriate interval distance is calculated based on the path length, and then guide markers are set at this interval distance along the optimal movement path. For example, in a simulated large shopping mall, if the optimal movement path is long, the interval distance of the guide markers is larger; if the path is short, the interval distance is smaller, to ensure that the guide markers can effectively guide users without causing visual interference.
[0110] Step S1365: Configure visual display parameters for each guide marker point. The visual display parameters include marker shape, color and flashing frequency. The marker shape adopts an arrow shape pointing in the direction of the target area. The color adopts a preset color with significant contrast to the surrounding environment. The flashing frequency is set to a fixed period.
[0111] Configure visual display parameters for each guide marker. The marker shape is arrow-shaped, with the arrow pointing in the direction of the target area, so users can intuitively know which direction to move in. For color, choose a preset color with significant contrast to the surrounding environment to make the guide markers more eye-catching in the virtual building scene and easier for users to notice. The flashing frequency is set to a fixed period to further attract user attention through flashing.
[0112] For example, in a simulated large shopping mall, the arrow shape of the guide markers points to the target area, the color may be set to bright red, and the flashing frequency is once per second, which can effectively guide users to reach the target area along the optimal movement path.
[0113] Step S1366: Integrate the coordinate sequence of the optimal movement path, the position parameters of the guide markers, and the visual display parameters into a path guidance rule, which is a component of the spatial layout adjustment rule.
[0114] The optimal movement path's coordinate sequence, guide marker location parameters, and visual display parameters are integrated to form a path guidance rule. This rule details how to display the optimal movement path and guide markers in a virtual building scene. The path guidance rule is used as part of the spatial layout adjustment rule to guide adjustments to the spatial layout of the virtual building scene. For example, in a simulated large shopping mall, the optimal movement path's coordinate sequence from the dining area to the entertainment area, the location of guide markers, and visual display parameters are integrated into a path guidance rule. Guide markers are added in the virtual building scene according to this rule, and the spatial layout is adjusted to facilitate users finding their way to the entertainment area.
[0115] Step S137: If the target type is functional combination reconstruction, then according to the functional requirements associated with the predicted intent type, the associated interactive elements of the current main focus area and the target area identifier are combined and displayed to generate an interactive element activation strategy for jointly activating associated elements.
[0116] When the target type of dynamic reconstruction is functional combination reconstruction, the associated interactive elements of the current main focus area and the target area are combined and displayed according to the functional requirements associated with the predicted intent type. Different predicted intent types correspond to different functional requirements. For example, if the predicted intent type is to compare products, then the interactive elements related to product comparison in the current main focus area and the target area need to be combined together.
[0117] These related interactive elements are combined, and an activation strategy for jointly activating these related elements is generated. This strategy includes setting the display method and operation rules of the combined interactive elements to ensure that users can easily use these combined elements to meet their functional needs. For example, in a simulated large shopping mall, if the predicted intent type is to compare the performance of different brands of mobile phones, the related interactive elements of the mobile phone display area (the current main focus area) and the mobile phone parameter comparison area (the target area), such as mobile phone images and parameter lists, are displayed together, and corresponding operation buttons are set to facilitate user comparison operations.
[0118] Step S138: Combine the spatial layout adjustment rules and interactive element activation strategies into structured instruction data to generate dynamic reconstruction instructions for the virtual building scene.
[0119] The generated spatial layout adjustment rules and interactive element activation strategies are combined into structured instruction data. Spatial layout adjustment rules are used to adjust the spatial structure and layout of the virtual building scene, such as area scaling and adding guide markers; interactive element activation strategies are used to control the display and operation status of interactive elements, such as improving the visibility of core interactive elements and jointly activating related elements.
[0120] These two parts are integrated according to a certain structure to form a dynamic reconstruction instruction for the virtual building scene. This dynamic reconstruction instruction contains all the information required to dynamically reconstruct the virtual building scene.
[0121] Step S140: Drive the virtual building scene to perform real-time spatial reorganization according to the dynamic reconstruction instruction, and generate multimodal feedback signals, which include visual scene update data and tactile interaction response data.
[0122] Based on the generated dynamic reconstruction instructions, the virtual building scene is spatially reorganized in real time. Simultaneously, to allow users to more intuitively perceive the changes in the virtual building scene, multimodal feedback signals are generated, including visual scene update data and tactile interaction response data.
[0123] Step S141: Parse the dynamic reconstruction instruction and extract the target functional area to be reconstructed and the corresponding adjustment parameters.
[0124] The dynamic restructuring instructions are parsed to extract the target functional areas that need to be reorganized and their corresponding adjustment parameters. The dynamic restructuring instructions detail the adjustment operations required for each target functional area, such as the scaling factor for area scaling and the activation state of interactive elements. By parsing the instructions, it is determined which functional areas need to be changed, and the specific adjustment methods and parameters for each area. For example, in a simulated large shopping mall, parsing the instructions determines that the clothing area and food and beverage area need to be reorganized, and obtains the corresponding scaling factor, element display level, and other adjustment parameters for each area.
[0125] Step S142: Call the spatial reorganization engine of the virtual building scene, load the basic layout data of the target functional area, and modify the position parameters in the basic layout data according to the spatial layout adjustment rules, including adjusting the area display size, updating the coordinates of the guide marker points, and modifying the relative positions of interactive elements.
[0126] The spatial reorganization engine of the virtual building scene is invoked. This engine is responsible for performing the actual spatial reorganization operation on the virtual building scene. First, the basic layout data of the target functional area to be reorganized is loaded. This data includes the initial position and state information of fixed structures, interactive elements, etc. within the area.
[0127] Based on the spatial layout adjustment rules, the positional parameters in the basic layout data are modified. If the spatial layout adjustment rules include area scaling and reconstruction, the display size of the area is adjusted according to the scaling factor, changing the size of the area in the virtual building scene; if they include path guidance reconstruction, the coordinates of the guide markers are updated to ensure they are accurately displayed on the optimal movement path; simultaneously, the relative positions of interactive elements are modified to meet the adjusted layout requirements. For example, in a simulated large shopping mall, the clothing area is enlarged according to the spatial layout adjustment rules, the coordinates of the guide markers are updated, and the relative positions of the clothing display racks and fitting rooms are adjusted.
[0128] Step S143: Update the state of interactive elements in the target functional area according to the interactive element activation strategy, activate the operable attributes of core interactive elements, set the semi-transparent display state of non-core interactive elements, and turn off the response function of irrelevant interactive elements.
[0129] Based on the interaction element activation strategy, the state of interactive elements within the target functional area is updated. For core interactive elements, their operable attributes are activated, enabling them to respond to user actions such as clicks and touches. For example, in a simulated dining area of a large shopping mall, if the predicted intent type is ordering food, then core interactive elements such as menus and order buttons are activated, allowing users to click on the menu to select dishes and click on the order button to place an order.
[0130] For non-core interactive elements, set them to a semi-transparent display state to reduce their visibility in the virtual building scene and narrow their interaction response trigger range, retaining only basic display information and simplified interactive feedback. For irrelevant interactive elements, disable their response function so they do not respond to user operations, and set their display transparency to a preset hidden value to release unnecessary texture resources and animation data they occupy, thereby improving the performance of the virtual building scene.
[0131] For example, step S1431: Extract the identifier list and priority sort of the core interactive elements from the interactive element activation strategy, traverse all interactive elements in the target functional area, and identify whether the identifier of each interactive element is in the identifier list of the core interactive elements.
[0132] The activation strategy for interactive elements extracts a list of identifiers and a priority ranking of core interactive elements. The identifier list contains unique identifiers for all elements identified as core interactive elements, while the priority ranking indicates the importance of each core interactive element. All interactive elements within the target functional area are traversed, and each element's identifier is checked against the identifier list. For example, in a simulated electronics section of a large shopping mall, the activation strategy identifies product detail display buttons, purchase buttons, etc., as core interactive elements, and their identifiers are listed in the identifier list. When traversing all interactive elements in this area, it is checked whether each element's identifier is in the identifier list.
[0133] Step S1432: For interactive elements identified in the core interactive element list, enable their operable attributes, set the interactive response trigger range to the complete boundary area of the element, and load the detailed texture resources and interactive feedback animation data of the interactive elements.
[0134] For interactive elements identified in the core interactive element list, enable their operable attributes to allow them to respond to user actions. Set the interaction response trigger range to the complete boundary area of the element, so that the corresponding interactive action will be triggered as long as the user clicks or touches any part of the element. Simultaneously, load detailed texture resources and interactive feedback animation data for the interactive elements to make them appear more realistic and vivid in the virtual building scene, enhancing the user's interactive experience. For example, in the cosmetics area of a simulated large shopping mall, enable the operable attribute for core interactive elements such as the lipstick try-on button. When the user clicks the button, it will trigger the try-on interactive action, while simultaneously loading the detailed texture of the lipstick and the try-on animation effect.
[0135] Step S1433: For interactive elements that are not in the core interactive element list but belong to the target functional area, determine them as non-core interactive elements, set their display transparency to a preset semi-transparent value, narrow the interactive response trigger range to the central area of the non-core interactive element, retain basic texture resources and simplify interactive feedback animation.
[0136] Interactive elements not listed in the core interactive element list but belonging to the target functional area are identified as non-core interactive elements. Their display transparency is set to a preset semi-transparent value, making them less noticeable in the virtual building scene and reducing distraction for the user. The interaction response trigger range is narrowed to the central area of the non-core interactive element; limited interactive operations are only triggered when the user clicks or touches the central part of the element. Basic texture resources are preserved, and interactive feedback animations are simplified to maintain basic element display and simple interactive feedback. For example, in the jewelry section of a simulated large shopping mall, some decorative elements are identified as non-core interactive elements, set to semi-transparent display, and their interaction response trigger range is narrowed, retaining only basic appearance textures and simple blinking animations.
[0137] Step S1434: For interactive elements that do not belong to the target functional area, determine them as irrelevant interactive elements, turn off their interactive response function, set their display transparency to the preset hidden value, and release the unnecessary texture resources and animation data they occupy.
[0138] Interactive elements that do not belong to the target functional area are identified as irrelevant interactive elements. Their interactive response function is disabled to prevent them from responding to user actions and avoid accidental user input. Their display transparency is set to a preset hidden value, making them completely hidden in the virtual building scene. Unnecessary texture resources and animation data they occupy are released to reduce resource consumption in the virtual building scene and improve system efficiency. For example, in a simulated large shopping mall, when the clothing area is reorganized, the interactive elements in the food and beverage area are identified as irrelevant interactive elements, their responsiveness is disabled, they are hidden, and their occupied resources are released.
[0139] Step S1435: Record the state update results of all interactive elements and generate an element state update log as a reference for the subsequent generation of multimodal feedback signals.
[0140] The system records the state updates of all interactive elements, including the activation state of core interactive elements, the semi-transparent display state of non-core interactive elements, and the closed state of irrelevant interactive elements. This information is compiled into an element state update log, which details the state changes of each interactive element. This log serves as a reference for generating subsequent multimodal feedback signals, providing fundamental information for generating accurate visual scene update data and haptic interaction response data. For example, in a simulated large shopping mall, the element state update log records the state changes of each interactive element in the clothing area. These records can then be used to adjust the visual display and haptic feedback when generating multimodal feedback signals.
[0141] Step S144: The spatial reorganization engine performs real-time rendering processing on the target functional area based on the modified layout data and the updated element state, generating visual scene update data containing spatial reorganization effects. The frame rate of the visual scene update data is consistent with the display refresh rate of the VR interactive terminal.
[0142] The spatial reorganization engine performs real-time rendering of the target functional area based on the modified layout data and updated element states. During the rendering process, information such as the adjusted area display size, guide marker positions, and display states of interactive elements are integrated to generate visual scene update data that includes the spatial reorganization effect.
[0143] To ensure a smooth user experience of virtual building scene changes, the frame rate of the visual scene update data is kept consistent with the display refresh rate of the VR interactive terminal. This allows users wearing the VR interactive terminal to observe the real-time, smooth spatial reorganization of the virtual building scene without any stuttering or flickering. For example, in a simulated large shopping mall, after the entertainment area is reorganized, the spatial reorganization engine renders the reorganized entertainment area in real time, transmitting the visual scene update data to the user at the same frame rate as the VR interactive terminal's display refresh rate, allowing the user to perceive the real-time changes in the entertainment area.
[0144] Step S145: Calculate tactile interaction response parameters based on the state of the interactive elements after spatial reorganization and the user's current behavioral unit, including the vibration intensity and vibration mode of the interaction point. The vibration intensity is positively correlated with the activation priority of the interactive element, and the vibration mode corresponds one-to-one with the interaction intent type.
[0145] Based on the state of the interactive elements after spatial reorganization and the user's current behavioral unit, haptic interaction response parameters are calculated. The activation priority of interactive elements reflects their importance and operability in the virtual building scene, and vibration intensity is positively correlated with the activation priority of interactive elements. The higher the activation priority of an interactive element, the greater the vibration intensity generated by the haptic feedback module of the VR interactive terminal when the user interacts with it.
[0146] Vibration patterns correspond one-to-one with interaction intent types, with different interaction intent types corresponding to different vibration patterns. For example, in a simulated large shopping mall, if the user's interaction intent is to click a product details button, the corresponding vibration pattern might be a brief but strong vibration; if it's browsing a product list, the vibration pattern might be a slight but continuous vibration. A specific algorithm calculates appropriate vibration intensity and vibration pattern based on the state of the interactive element and the user's behavior unit, serving as tactile interaction response parameters.
[0147] Step S146: Convert the tactile interaction response parameters into control signals for the VR interactive terminal's tactile feedback module, align them with the visual scene update data by timestamp, and generate a multimodal feedback signal containing both visual scene update data and tactile interaction response data.
[0148] The calculated haptic interaction response parameters are converted into control signals for the haptic feedback module of the VR interactive terminal. The haptic feedback module generates corresponding vibrations based on these control signals, allowing the user to experience interaction within the virtual building scene through touch.
[0149] The converted control signals are aligned with the visual scene update data using timestamps. Each timestamp records the moment a data point is generated. By aligning the timestamps, the visual scene updates and haptic interaction responses are ensured to be synchronized in time, allowing users to experience corresponding haptic feedback while observing changes in the virtual building scene. The aligned visual scene update data and haptic interaction response data are then integrated to generate a multimodal feedback signal that includes both visual scene update data and haptic interaction response data. For example, in a simulated large shopping mall, when a user clicks on a core interactive element, the element in the visual scene will display corresponding feedback. Simultaneously, the haptic feedback module of the VR interactive terminal will vibrate according to the haptic interaction response parameters. This synchronized visual and haptic feedback is transmitted to the user through the multimodal feedback signal, enhancing the user's interactive experience.
[0150] Step S150: Collect the user's real-time response behavior to the multimodal feedback signal, and optimize the prediction model of the intention evolution trajectory based on the real-time response behavior.
[0151] By collecting real-time user responses to multimodal feedback signals and analyzing these responses, the prediction model for intent evolution trajectories can be optimized, enabling the prediction model to more accurately predict user intent and improve the quality of interactive experience in virtual building scenarios.
[0152] Step S151: The behavior capture module of the VR interactive terminal continuously collects the user's real-time response behavior after receiving multimodal feedback signals. The real-time response behavior includes the user's new limb movement trajectory, new gesture sequence, and gaze position.
[0153] The VR interactive terminal's behavior capture module continuously collects the user's real-time response behavior after receiving multimodal feedback signals. The limb sensors in the behavior capture module continue to record the user's new limb movement trajectories. The user may move within the virtual building scene based on updated visual scene data and tactile interaction response data, such as walking towards a target area indicated by a guide marker. The gesture recognition component collects the user's new gesture sequences; the user may make gestures such as clicking or selecting to interact with interactive elements. Simultaneously, eye-tracking technology is used to collect the user's gaze placement, understanding the key areas and elements the user focuses on within the virtual building scene. For example, in a simulated large shopping mall, when the user receives a multimodal feedback signal regarding the reorganization of the clothing section, the behavior capture module records the user's new limb movement trajectory, such as whether they walk towards the reorganized clothing display area; collects new gesture sequences, such as whether they click the purchase button on clothing; and captures the gaze placement, such as whether it lingers on a particular garment.
[0154] Step S152: Extract features from the real-time response behavior to obtain the degree of consistency between the new limb movement trajectory and the optimal movement path, the number of interactions between the new gesture sequence and the core interactive element, and the duration of the gaze position on the core interactive element.
[0155] Feature extraction is performed on the collected real-time response behaviors. For new limb movement trajectories, the degree of similarity between them and the optimal movement path is calculated. This can be achieved by comparing the coordinate sequences of the new limb movement trajectories with those of the optimal movement path, thus calculating the similarity. For new gesture sequences, the number of interactions with core interactive elements is counted to understand the frequency of user operations on these elements. For gaze positions, the duration of gaze on core interactive elements is recorded to reflect the user's level of attention to these elements. For example, in a simulated large shopping mall, the degree of similarity between the user's new limb movement trajectory and the optimal movement path to the fitting room is calculated; the number of times the user clicks on core interactive elements such as clothing purchase buttons is counted; and the duration of the user's gaze on a popular garment is recorded.
[0156] Step S153: Compare the matching degree, number of interactions and duration with the preset feedback evaluation thresholds. When the matching degree is higher than the set matching degree threshold, the number of interactions reaches the set number threshold, and the duration exceeds the set duration threshold, the real-time response behavior is determined to be valid positive feedback; otherwise, it is determined to be invalid or negative feedback.
[0157] The extracted relevance, number of interactions, and duration are compared with preset feedback evaluation thresholds. These thresholds were determined based on extensive experimental data and user behavior analysis. When the relevance exceeds the set threshold, the number of interactions reaches the set threshold, and the duration exceeds the set threshold, it indicates that the user has responded positively to the multimodal feedback signal, and the real-time response behavior is judged as valid positive feedback. For example, in a simulated large shopping mall, if the user's new limb movement trajectory matches the optimal movement path with a relevance exceeding the set threshold, the number of clicks on the core interactive element reaches the set number, and the duration of gaze lingering on the core interactive element exceeds the set duration, then it is judged as valid positive feedback. Conversely, if these conditions are not met, it is judged as invalid or negative feedback.
[0158] Step S154: Based on the real-time response behavior of effective positive feedback, extract the corresponding intent evolution nodes and dynamic reconstruction instruction parameters, increase the weight value of such intent evolution nodes in the prediction model, and optimize the transfer probability of the intent transfer network.
[0159] Based on real-time response behavior with effective positive feedback, corresponding intent evolution nodes and dynamic reconstruction instruction parameters are extracted from the intent evolution trajectory and dynamic reconstruction instructions. These intent evolution nodes and parameters represent predicted intents and reconstruction operations that can elicit positive user responses.
[0160] Increasing the weight of these intent evolution nodes in the prediction model makes it more inclined to predict these types of intents. Simultaneously, optimizing the transition probabilities of the intent transition network by adjusting the magnitude of each intent transition probability based on effective positive feedback allows the model to more accurately reflect the transition patterns of user intents. For example, in a simulated large shopping mall, if a user provides effective positive feedback regarding the zooming and reconstruction of the clothing area and the activation of core interactive elements, the corresponding intent evolution nodes and dynamic reconstruction instruction parameters are extracted. Increasing the weight of these intent evolution nodes and adjusting the transition probabilities of related intents in the intent transition network improves prediction accuracy.
[0161] Step S155: Based on the real-time response behavior of invalid or negative feedback, analyze the reasons for the prediction deviation of the intent evolution trajectory, adjust the prediction probability calculation method of the intent evolution node, and modify the spatial layout adjustment rules or interactive element activation strategy in the dynamic reconstruction instruction.
[0162] For real-time response behaviors with invalid or negative feedback, analyze the reasons for the prediction deviation in the intent evolution trajectory. This could be due to inaccurate prediction of the intent type, or the spatial layout adjustment rules and interactive element activation strategies in the dynamic reconstruction instructions not meeting the user's needs.
[0163] Based on the analysis results, the prediction probability calculation method for intent evolution nodes is adjusted by introducing new factors or adjusting weight allocation to make the prediction probability more accurate. Simultaneously, the spatial layout adjustment rules or interactive element activation strategies in the dynamic reconstruction instructions are modified, such as adjusting the area scaling factor or changing the selection of core interactive elements, to improve the interactive effect of the virtual building scene. For example, in a simulated large shopping mall, if users are not interested in the setting of guide markers, the interval distance and visual display parameters of the guide markers are adjusted after analyzing the reasons, and the path guidance rules in the dynamic reconstruction instructions are modified.
[0164] Step S156: Update the optimized weight values, transition probabilities and calculation methods to the prediction model of the intent evolution trajectory, and verify the optimization effect through a new round of virtual behavior mapping data collection and intent evolution trajectory construction.
[0165] The optimized weights, transition probabilities, and calculation methods are then updated in the prediction model for the intent evolution trajectory. The updated prediction model can more accurately predict user intent data.
[0166] By collecting virtual behavior mapping data and constructing intent evolution trajectories in a new round, continuous user behavior sequences in the physical space are collected again to generate virtual behavior mapping data and construct intent evolution trajectories. The new intent evolution trajectories are compared with those before optimization to verify the optimization effect. If the optimized prediction model can more accurately predict user intent and users respond more positively to interactions in the virtual building scene, it indicates that the optimization has achieved good results; otherwise, further analysis of the reasons is needed, and continued optimization is required. For example, in a simulated large shopping mall, user behavior data is collected again after optimization to construct intent evolution trajectories, and the improvement in user response behavior is observed to verify the effectiveness of the prediction model optimization.
[0167] Figure 2 The illustration shows exemplary hardware and software components of an intelligent building system interactive experience system 100 incorporating VR, which can implement the ideas of this application, according to some embodiments of this application. For example, a processor 120 can be used in the intelligent building system interactive experience system 100 incorporating VR and to perform the functions in this application.
[0168] The VR-integrated intelligent building system interactive experience system 100 can be a general-purpose server or a special-purpose server; both can be used to implement the VR-integrated intelligent building system interactive experience method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the load.
[0169] For example, the VR-integrated intelligent building system interactive experience system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the VR-integrated intelligent building system interactive experience system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to these program instructions. The VR-integrated intelligent building system interactive experience system 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0170] For ease of explanation, only one processor is described in the VR-integrated intelligent building system interactive experience system 100. However, it should be noted that the VR-integrated intelligent building system interactive experience system 100 of this application may also include multiple processors. Therefore, the steps executed by one processor described in this application may also be executed jointly or individually by multiple processors. For example, if the processor of the VR-integrated intelligent building system interactive experience system 100 executes steps A and B, it should be understood that steps A and B may also be executed jointly by two different processors or individually by one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0171] Furthermore, this embodiment of the invention also provides a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-mentioned method for interactive experience of intelligent building system combined with VR is realized.
[0172] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A method for interactive experience of an intelligent building system combining VR, characterized in that, The method includes: The continuous sequence of user behavior in physical space is collected through VR interactive terminals to generate virtual behavior mapping data, which includes the correlation between user physical behavior and functional areas of virtual building scene; Based on the virtual behavior mapping data, a user's intent evolution trajectory in the virtual building scene is constructed. The intent evolution trajectory includes the intent change trend in the time dimension and the regional attention path in the spatial dimension. Based on the intent evolution trajectory, a dynamic reconstruction instruction for the virtual building scene is generated, which includes spatial layout adjustment rules and interactive element activation strategies. The virtual building scene is driven to perform real-time spatial reorganization according to the dynamic reconstruction instructions, and multimodal feedback signals are generated. The multimodal feedback signals include visual scene update data and tactile interaction response data. The system collects real-time response behavior of users to the multimodal feedback signals and optimizes the prediction model of intent evolution trajectory based on the real-time response behavior.
2. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The step of collecting continuous behavioral sequences of users in physical space through a VR interactive terminal to generate virtual behavior mapping data includes: The VR interactive terminal's behavior capture module is activated, and the built-in limb sensor and gesture recognition component are activated. The limb sensor is used to collect the user's limb movement trajectory data, and the gesture recognition component is used to collect the user's gesture action sequence data. The limb movement trajectory data is processed by spatial coordinate transformation to map the trajectory data in the physical space coordinate system to a three-dimensional movement path in the virtual building scene coordinate system. The three-dimensional movement path contains the user's real-time position coordinate sequence in the virtual building scene. The gesture sequence data is processed by keyframe extraction to identify the start frame, peak frame and end frame of the gesture, extract the gesture contour features and motion vectors of each keyframe, and determine the type and execution force of the gesture. The three-dimensional movement path and gesture action type are aligned according to timestamps to generate behavior units containing spatial location information and action feature information. Each behavior unit corresponds to a complete physical behavior segment. The behavioral units are processed by sequence integration and combined in chronological order to form a continuous sequence of user behavior in physical space. Each behavioral unit in the continuous sequence of behavior is associated with a corresponding time stamp and action feature parameters. The continuous behavior sequence is matched with the functional area division data of the virtual building scene to determine the virtual functional area identifier corresponding to each behavior unit. The interaction intent probability of the behavior unit is calculated in combination with the gesture action type to generate virtual behavior mapping data containing virtual area identifier, interaction intent probability and time stamp.
3. The method for interactive experience of intelligent building systems combined with VR according to claim 2, characterized in that, The keyframe extraction process for the gesture sequence data includes identifying the start frame, peak frame, and end frame of the gesture, extracting the gesture contour features and motion vectors of each keyframe, and determining the type and intensity of the gesture. The inter-frame difference calculation is performed on the gesture action sequence data to obtain the pixel change between adjacent frames, and the frame where the pixel change first exceeds a preset threshold is marked as the starting frame of the gesture action. In the frame sequence after the starting frame, the cumulative value of pixel change is continuously calculated. When the cumulative value reaches the maximum value, the corresponding frame is marked as the peak frame of the gesture action. In the frame sequence following the peak frame, when the pixel change is below a preset threshold for multiple consecutive frames, the last frame with a change below the threshold is marked as the end frame of the gesture action. Extract gesture image data from the start frame, peak frame, and end frame; perform edge detection processing on the gesture image data to obtain the set of boundary point coordinates of the gesture contour; and convert the set of boundary point coordinates into a contour feature vector as the gesture contour feature. Calculate the pixel displacement between the peak frame and the start frame to generate the motion direction vector of the gesture. Combine the pixel displacement between the peak frame and the end frame to calculate the magnitude of the motion direction vector as the motion amplitude of the gesture. The gesture contour features are compared with a preset gesture template library to determine the type of gesture action. The motion amplitude is then normalized and used as the execution force parameter of the gesture action.
4. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The construction of the user's intent evolution trajectory in the virtual building scene based on the virtual behavior mapping data includes: The virtual behavior mapping data is parsed, and the virtual area identifier, interaction intent probability, and time stamp of each behavior unit are extracted and arranged in the order of time stamps to form a sequence of behavior units; The frequency and duration of each virtual region identifier in the behavioral unit sequence are statistically analyzed, and the virtual region identifier with the highest frequency and longest duration is determined as the user's current main focus area; The interaction intent probabilities in the sequence of behavioral units are weighted and calculated, with time stamp as the weighting factor. The intent confidence of each behavioral unit is calculated, and behavioral units with intent confidence higher than a preset threshold are selected as valid intent units. Analyze the changes in virtual region identifiers and interaction intent probabilities between adjacent valid intent units, calculate intent transfer probabilities, and construct an intent transfer network, where nodes in the intent transfer network represent interaction intent types and edges represent intent transfer probabilities. Based on the intent transfer network and the current main focus area, the possible type of user's next interaction intent and the corresponding virtual area are predicted, and an intent evolution node is generated. The intent evolution node includes the predicted intent type, target area identifier and predicted probability. The effective intent unit sequence and intent evolution nodes are integrated in chronological order to form an intent evolution trajectory that includes historical intent trajectories and future prediction directions. Each node in the intent evolution trajectory is associated with a corresponding timestamp, virtual region identifier, and intent confidence level.
5. The interactive experience method for intelligent building systems combined with VR according to claim 4, characterized in that, The analysis of changes in virtual region identifiers and interaction intent probabilities between adjacent valid intent units, calculation of intent transfer probabilities, and construction of an intent transfer network includes: Extract two adjacent valid intent units from the valid intent unit sequence, denoted as the preceding intent unit and the following intent unit, and obtain the interaction intent type of the preceding intent unit and the interaction intent type of the following intent unit; The frequency of all combinations of preceding and subsequent intent types in the effective intent unit sequence is counted, and the proportion of the frequency of each combination to the total frequency of combinations is calculated as the intent transfer probability. An initial intent transfer network is constructed using interaction intent types as network nodes and intent transfer probabilities as the weights of directed edges between nodes. Analyze the changes in virtual region identifiers in the sequence of valid intent units. When the virtual region identifiers of adjacent valid intent units are different, adjust the weight value of the corresponding intent transfer probability and increase the probability weight of cross-region intent transfer. When the change in the interaction intent probability of adjacent valid intent units exceeds a preset threshold, the intent transfer probability is adjusted a second time, increasing the transfer weight of intents with significant probability changes. The adjusted intent transfer probabilities are updated to the initial intent transfer network to generate an intent transfer network that includes regional association features.
6. The interactive experience method for intelligent building systems combined with VR according to claim 1, characterized in that, The dynamic reconstruction instruction for generating a virtual building scene based on the intent evolution trajectory includes: The intent evolution trajectory is analyzed to extract the virtual region identifier of the current main focus area, the predicted intent type of the intent evolution node, and the target region identifier; Retrieve the basic layout data corresponding to the current main focus area and target area identifiers from the area configuration library of the virtual building scene. The basic layout data includes the position parameters of fixed structures within the area, the initial state of interactive elements, and spatial connection relationships. Based on the predicted intent type matching preset reconstruction rule base, the target type of dynamic reconstruction is determined. The target type includes area scaling reconstruction, element priority reconstruction, path guidance reconstruction and function combination reconstruction. If the target type is regional scaling and reconstruction, the scaling factor is calculated based on the intent confidence of the current main focus area, and a spatial layout adjustment rule is generated to adjust the display size of the area according to the scaling factor. The higher the intent confidence, the larger the scaling factor. If the target type is element priority reconstruction, then extract the core interactive elements in the area according to the predicted intent type, set the display level of the core interactive elements to be higher than other elements, and generate an interactive element activation strategy to improve the visibility of the core interactive elements. If the target type is path-guided reconstruction, then based on the spatial relationship between the current main focus area and the target area identifier, the optimal movement path is planned, and spatial layout adjustment rules are generated to add visual guidance marks on the optimal movement path. If the target type is functional combination reconstruction, then based on the functional requirements associated with the predicted intent type, the associated interactive elements of the current main focus area and the target area identifier are combined and displayed to generate an interactive element activation strategy that jointly activates associated elements. The spatial layout adjustment rules and interactive element activation strategies are combined into structured instruction data to generate dynamic reconstruction instructions for the virtual building scene.
7. The method for interactive experience of intelligent building systems combined with VR according to claim 6, characterized in that, If the target type is path-guided reconstruction, then based on the spatial relationship between the current primary focus area and the target area identifiers, the optimal movement path is planned, and spatial layout adjustment rules for adding visual guidance markers on the path are generated, including: Extract the center coordinates of the current main area of interest and the center coordinates of the target area identifier from the spatial map data of the virtual building scene, and calculate the straight-line distance and relative direction between the two center coordinates; Based on obstacle distribution data of virtual building scene, obstacle avoidance processing is performed on straight-line distance to generate multiple candidate movement paths. The candidate movement paths are all continuous paths that connect the center coordinates of the current main focus area and the center coordinates of the target area identifier and do not pass through obstacles. Calculate the path length and number of turns for each candidate movement path, and determine the candidate movement path with the highest priority as the optimal movement path. The shorter the path length and the fewer the number of turns, the higher the priority of the candidate movement path. Guide markers are set at equal intervals along the optimal movement path. The interval between the guide markers is dynamically adjusted according to the path length, with the interval increasing as the path length increases. Visual display parameters are configured for each guide marker point, including marker shape, color, and flashing frequency. The marker shape is an arrow pointing in the direction of the target area, the color is a preset color with significant contrast to the surrounding environment, and the flashing frequency is set to a fixed period. The coordinate sequence of the optimal movement path, the position parameters of the guide markers, and the visual display parameters are integrated into a path guidance rule, which is used as a component of the spatial layout adjustment rule.
8. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The process of driving the virtual building scene to perform real-time spatial reorganization according to the dynamic reconstruction instructions and generating multimodal feedback signals includes: The dynamic reconstruction instruction is analyzed to extract the spatial layout adjustment rules and interactive element activation strategies, and to determine the target functional areas that need to be reorganized and their corresponding adjustment parameters. The virtual building scene's spatial reorganization engine is invoked to load the basic layout data of the target functional area. Based on the spatial layout adjustment rules, the position parameters in the basic layout data are modified, including adjusting the area display size, updating the coordinates of the guide marker points, and modifying the relative positions of interactive elements. Based on the interaction element activation strategy, update the state of the interaction elements in the target functional area, activate the operable attributes of the core interaction elements, set the semi-transparent display state of the non-core interaction elements, and turn off the response function of irrelevant interaction elements. The spatial reorganization engine performs real-time rendering processing on the target functional area based on the modified layout data and the updated element status, generating visual scene update data containing spatial reorganization effects. The frame rate of the visual scene update data is consistent with the display refresh rate of the VR interactive terminal. Based on the state of the interactive elements after spatial reorganization and the user's current behavioral unit, tactile interaction response parameters are calculated, including the vibration intensity and vibration mode of the interaction point. The vibration intensity is positively correlated with the activation priority of the interactive element, and the vibration mode corresponds one-to-one with the interaction intention type. The tactile interaction response parameters are converted into control signals for the tactile feedback module of the VR interactive terminal, and aligned with the visual scene update data by timestamp to generate a multimodal feedback signal containing both visual scene update data and tactile interaction response data.
9. The method for interactive experience of intelligent building systems combined with VR according to claim 1, characterized in that, The process of collecting users' real-time response behavior to the multimodal feedback signals and optimizing the prediction model of intent evolution trajectory based on the real-time response behavior includes: The VR interactive terminal continuously collects the user's real-time response behavior after receiving multimodal feedback signals through the behavior capture module. The real-time response behavior includes the user's new limb movement trajectory, new gesture sequence, and gaze position. Feature extraction is performed on the real-time response behavior to obtain the degree of consistency between the new limb movement trajectory and the optimal movement path, the number of interactions between the new gesture sequence and the core interactive element, and the duration of the gaze position on the core interactive element. The matching degree, number of interactions, and duration are compared with preset feedback evaluation thresholds. When the matching degree is higher than the set matching degree threshold, the number of interactions reaches the set number threshold, and the duration exceeds the set duration threshold, the real-time response behavior is determined to be valid positive feedback; otherwise, it is determined to be invalid or negative feedback. Based on the real-time response behavior with effective positive feedback, the corresponding intent evolution nodes and dynamic reconstruction instruction parameters are extracted, the weight values of such intent evolution nodes in the prediction model are increased, and the transfer probability of the intent transfer network is optimized. Based on the real-time response behavior of invalid or negative feedback, analyze the reasons for the prediction deviation of the intent evolution trajectory, adjust the prediction probability calculation method of intent evolution nodes, and modify the spatial layout adjustment rules or interactive element activation strategies in the dynamic reconstruction instructions. The optimized weight values, transition probabilities, and calculation methods are updated to the prediction model of the intent evolution trajectory. The optimization effect is verified through a new round of virtual behavior mapping data collection and intent evolution trajectory construction.
10. An intelligent building system interactive experience system incorporating VR, characterized in that, The system includes a processor and a memory, the memory being connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the VR-integrated intelligent building system interactive experience method as described in any one of claims 1-9.
Citation Information
Patent Citations
Building model modification interaction method, system and terminal based on virtual reality technology
CN115690375A
EEG feedback optimization method and system for building appearance visual features and emotion regulation
CN119646925A
User-interactivity enabled search filter tool optimized for virtualized worlds
US12159364B2
Dynamic user interactions for display control and measuring degree of completeness of user gestures
US20140201683A1
Mixed-reality and CAD architectural design environment
US20180197340A1
Cited By
Immersive experience-oriented digital media interaction control system and method
CN121879589A