Augmented reality content recommendation and expressive presentation method based on user behavior perception
By constructing a state representation and adaptive weight mechanism for virtual characters, the position and orientation of virtual characters in the AR system are optimized, which solves the problem of insufficient user behavior perception in complex open scenes, improves the naturalness of the layout and semantic expression, and enhances the user interaction experience and information transmission effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-17
AI Technical Summary
Existing AR systems lack the ability to perceive user behavior in complex and open scenarios, resulting in simplistic content display strategies, frequent visual conflicts, and negative impacts on user interaction experience and information delivery.
By constructing a state representation of a virtual character, designing a joint objective function, introducing an adaptive weighting mechanism, and iteratively optimizing the position and orientation of the virtual character, visual conflicts are avoided, and the naturalness and semantic expressiveness of the layout are improved.
It enables automatic optimization of character position and orientation in augmented reality scenarios, avoiding visual conflicts, improving the naturalness, harmony and semantic expression of the layout, and enhancing the interactive experience and information transmission effect of the AR system.
Smart Images

Figure CN121685902A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer vision and augmented reality technology, and in particular to an augmented reality content recommendation and expressive presentation method based on user behavior perception. BACKGROUND
[0002] With the rapid development of augmented reality (AR) technology in the fields of computer vision, graphics rendering and human-computer interaction, it has gradually been popularized in application scenarios such as educational display, digital cultural relics, city tour and commercial retail. AR systems improve the spatial expressiveness of information and the immersion of users by superimposing virtual content in the real world, and become an important means of new generation information acquisition and interaction. However, the current mainstream AR systems still face the problem of single display strategy, which limits their adaptability in complex open scenarios.
[0003] Current AR systems mostly use a unified template-driven visualization method in content presentation, emphasizing the information density and visual uniformity of the content, and lack expressive modeling and style diversity design. Especially in the presence of multiple contents or complex scene switching, the contents are prone to occlusion, interference or visual congestion, affecting the efficiency of users receiving important information. At the same time, the existing content layout strategy lacks the ability to perceive the stability of the user's perspective, the moving path and the interest shift, and cannot realize dynamic content organization and visual guidance for user behavior.
[0004] Therefore, how to dynamically optimize the layout of virtual roles, avoid visual conflicts, improve the naturalness, coordination and semantic expression ability of the layout, and thus enhance the interactive experience and information transmission effect of the AR system, is a technical problem to be solved by those skilled in the art. SUMMARY
[0005] The present application provides an augmented reality content recommendation and expressive presentation method based on user behavior perception, which dynamically optimizes the layout of virtual roles, avoids visual conflicts, improves the naturalness, coordination and semantic expression ability of the layout, and thus enhances the interactive experience and information transmission effect of the AR system.
[0006] In one aspect, the present application provides an augmented reality content recommendation and expressive presentation method based on user behavior perception, which comprises: constructing a state representation of a plurality of virtual roles in an augmented reality scene; the state representation includes the three-dimensional spatial position, the facing direction and the occlusion state of each virtual role; construct a joint objective function for evaluating rationality of virtual character layout; the joint objective function comprises a space cost function and a direction cost function, the space cost function is used for quantifying rationality of space relationship between characters, and the direction cost function is used for quantifying coordination and diversity of directions of the characters; According to attribute information of the virtual characters, adaptive weights are generated through a pre-constructed weight prediction model, and the adaptive weights are used to adjust cost items related to importance of the characters in the space cost function; Based on the joint objective function, positions and directions of all virtual characters are iteratively optimized to obtain an optimized layout; The optimized layout is expressively presented in an augmented reality scene.
[0007] The method for recommending and expressively presenting augmented reality content based on user behavior perception provided by the application optimizes positions and directions of characters in a three-dimensional space automatically, avoids visual conflicts, and improves naturalness, coordination and semantic expression capability of the layout, thereby significantly enhancing interactive experience and information transmission effect of an AR system. BRIEF DESCRIPTION OF DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0009] Figure 1 is a flowchart of the method for recommending and expressively presenting augmented reality content based on user behavior perception provided by the embodiments of the application; Figure 2 is a structural diagram of the system for recommending and expressively presenting augmented reality content based on user behavior perception provided by the embodiments of the application. DETAILED DESCRIPTION
[0010] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.
[0011] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing description, and cannot be understood as indicating or implying relative importance.
[0012] Figure 1 is a flowchart of a method for augmented reality content recommendation and expressive presentation based on user behavior perception provided by an embodiment of the present application.
[0013] As shown in Figure 1 , the method for augmented reality content recommendation and expressive presentation based on user behavior perception provided by an embodiment of the present application mainly includes the following steps: 101, constructing a state representation of a plurality of virtual roles in an augmented reality scene; the state representation includes a three-dimensional space position, a face direction and an occlusion state of each virtual role; 102, constructing a joint objective function for evaluating rationality of virtual role layout; the joint objective function includes a space cost function and a direction cost function, the space cost function is used to quantify rationality of spatial relationship between roles, and the direction cost function is used to quantify coordination and diversity of role directions; 103, generating adaptive weights by a pre-constructed weight prediction model according to attribute information of the virtual roles, and using the adaptive weights to adjust cost items related to role importance in the space cost function; 104, based on the joint objective function, iteratively optimizing positions and directions of all virtual roles to obtain an optimized layout; 105, performing expressive presentation of the optimized layout in the augmented reality scene.
[0014] Among them, the virtual role refers to a digitalized image presented in the augmented reality scene, which can have specific attributes, functions and interaction logic, such as a virtual guide, a virtual audience, a virtual exhibit identification, etc.
[0015] The state representation is a set of structured descriptions of core features of the virtual roles in the scene, which is used for subsequent layout optimization and interaction reasoning.
[0016] The three-dimensional spatial position is the specific coordinate of the virtual character in the three-dimensional coordinate system of the augmented reality scene, used to locate the spatial position of the character.
[0017] The facing direction is the orientation of the front of the virtual character, affecting the visibility and interactive perception of the character by the user.
[0018] The occlusion state is an indication of whether the virtual character is occluded by other characters, scene objects, or boundaries, in an invisible or partially visible state.
[0019] The joint objective function is a mathematical model that comprehensively evaluates the pros and cons of the layout of virtual characters, providing a quantitative basis for layout optimization by fusing space and orientation-related cost terms.
[0020] The space cost function is a function used to measure the rationality of the spatial relationship between virtual characters and between characters and the scene, quantifying problems such as collision, boundary crossing, and uneven distribution.
[0021] The orientation cost function is a function used to measure the coordination and diversity of the orientation of virtual characters, ensuring that the orientation of the characters meets the needs of interactive logic and visual expression.
[0022] Attribute information is the inherent characteristic data of virtual characters, including categorical features (such as occupation, character type) and numerical features (such as screen frequency, seniority level).
[0023] Adaptive weight is an importance coefficient dynamically generated based on character attributes, which can be adjusted according to the scene and character state, used to optimize the layout priority.
[0024] Specifically, first, for the target virtual character in the current AR scene, the three-dimensional spatial position of each character is determined one by one, i.e., the specific spatial coordinates of each character are located through the scene coordinate system; the facing direction of each character is determined, taking the preset reference direction in the scene (such as the direction directly in front of the user's perspective, the center direction of the scene) as the reference, to determine the direction of the character's front; at the same time, it is judged whether each character is occluded, and the occlusion state (visible or invisible) of the character is marked through scene object contour detection, character spatial overlap judgment, etc., to finally form a complete state representation of each virtual character containing three-dimensional spatial position, facing direction and occlusion state. The state representation of each virtual character can be denoted as , and the overall representation set can be denoted as formula (1): (1) wherein, represents the absolute position coordinates of the three-dimensional spatial position, usually in meters (m), used to determine the specific geographical location in the scene; represents the facing direction, the included angle with the speaker's front direction, in radians (rad), the value of which affects the visibility and interactivity of the content by the user; A binary flag indicating whether the character is currently in an occluded state, 1 for visible and 0 for occluded or invisible.
[0025] Then a joint objective function is constructed for evaluating the rationality of the layout. Among them, the space cost function mainly quantifies the rationality of the spatial relationship between characters and the scene, such as judging whether the characters collide, whether they are close to the scene boundary, whether they are reasonably distributed around the scene center, whether they evenly cover the scene area, etc. The orientation cost function focuses on the coordination and diversity of the character's orientation, both ensuring that the character's orientation meets the scene's interactive needs (such as facing the user or the interactive focus), and avoiding the visual monotony caused by all characters facing the same direction.
[0026] Subsequently, the attribute information of each virtual character is collected, such as the character's occupation type, functional positioning in the scene, historical appearance frequency, seniority level, etc. After preprocessing these attribute information, input the pre-trained weight prediction model, the model automatically outputs the importance weight of each character according to the attribute characteristics, that is, the adaptive weight. Apply this weight to the space cost function to adjust the cost items related to the importance of the character, such as making important characters more inclined to be close to the scene center and reducing their probability of being occluded.
[0027] Then based on the joint objective function constructed above, start the iterative optimization of the positions and orientations of all virtual characters. In each iteration, according to the calculation results of the joint objective function, adjust the three-dimensional spatial positions (such as moving overlapping characters, adjusting important characters deviating from the center, dispersing too dense characters) and face directions (such as correcting the orientation of characters away from the interactive focus, increasing the diversity of orientation) of part or all virtual characters; then recalculate the cost value of the joint objective function to determine whether the optimization goal is reached. If not, continue to iterate and adjust until the cost value stabilizes in the preset range or meets the iteration times requirement, and get the final optimized layout.
[0028] Finally, the optimized virtual character layout is presented in the augmented reality scene. During the expressive presentation, the size, color, and other presentation attributes of the characters can be fine-tuned to meet the optimization goal and improve the user's visual experience and information reception efficiency.
[0029] The user behavior perception based augmented reality content recommendation and expressive presentation method of the embodiment constructs a state representation of a virtual role, designs a joint optimization target combining spatial rationality and orientation expressiveness, introduces an adaptive weight mechanism based on role attributes, and iteratively optimizes the positions and orientations of all virtual roles based on the joint target function to obtain an optimized layout for expressive presentation in an augmented reality scene, automatically optimizing the positions and orientations of the roles in three-dimensional space, avoiding visual conflicts, and improving the naturalness, coordination and semantic expression ability of the layout, thereby significantly enhancing the interactive experience and information transmission effect of the AR system.
[0030] In some embodiments, the process of constructing a joint target function for evaluating the rationality of the virtual role layout includes: The spatial cost function is constructed, which is the sum of a collision cost function, a boundary cost function, a center cost function, a distribution cost function and a region cost function; The orientation cost function is constructed, which is the sum of a deviation cost function and a diversity cost function. The deviation cost function calculates the deviation of the current orientation of each virtual role from its ideal orientation, which is determined by the position of the virtual role relative to the preset focus point. The diversity cost function encourages the distribution of all virtual role orientations to have diversity.
[0031] The collision cost function is based on the intersection-over-union calculation of the role bounding box, which is a sub-function for quantifying the degree of spatial overlap between virtual roles. The boundary cost function penalizes virtual roles close to the scene boundary, avoiding roles exceeding the user's visual range or being at the edge of the scene affecting the experience. The center cost function encourages virtual roles to approach the scene center according to the corresponding importance weight, which is a sub-function for guiding virtual roles to approach the scene center according to the importance weight, improving the visual priority of important roles. The distribution cost function encourages virtual roles to be evenly distributed in space, which is a sub-function for encouraging virtual roles to be evenly distributed in the scene space, avoiding excessive clustering or dispersion of roles. The region cost function encourages virtual roles to cover multiple partitioned regions of the scene, which is a sub-function for guiding virtual roles to cover multiple partitioned regions of the scene, ensuring the full use of the scene space.
[0032] The role bounding box is a virtual boundary box constructed for a virtual role, which is used to simplify the collision detection calculation of the role's spatial position. The intersection-over-union refers to the ratio of the volume of the overlapping part of two role bounding boxes to the total volume of the two bounding boxes, which is used to determine whether the roles collide and the degree of collision.
[0033] The deviation cost function is used to quantify the difference between the current orientation of the virtual character and the ideal orientation, which can calculate the deviation of the current orientation of each virtual character from its ideal orientation. The diversity cost function is used to encourage virtual characters to face a diverse distribution and avoid visual monotony caused by facing a single direction. The ideal orientation can be determined according to the optimal orientation of the virtual character relative to the preset focus point, ensuring that the character orientation conforms to the interaction logic.
[0034] The preset focus point is a pre-set core interaction position in the scene, such as the current location of the user, the center point of the scene function, the location of important virtual objects, etc.
[0035] Specifically, the spatial cost function quantifies the rationality of the spatial layout through the superposition of multi-dimensional sub-functions. The formula is shown in equation (2): (2) wherein, represents the collision cost function; is the boundary cost function; is the center cost function; represents the distribution cost function; is the area cost function.
[0036] The collision cost function penalizes the spatial overlap between any two virtual characters. The intersection over union (IoU) between the bounding boxes of the characters is used as the basis for judgment. The specific form is shown in equation (3): (3) wherein, represents the bounding box of virtual character ; represents the bounding box of virtual character ; I( ) represents the indicator function, which is 1 if the condition is true, otherwise 0; N represents the number of virtual characters.
[0037] The boundary cost function penalizes characters close to the scene boundary. The formulas are shown in equations (4) and (5): (4) (5) wherein, and represent the half-width of the scene boundary in the x and z directions; represents the x-coordinate of virtual character ; represents the z-coordinate of virtual character ; is the edge buffer distance, usually set to 2m (adjust according to the specific space size).
[0038] The cost of the center cost function encourages important characters to be close to the scene center point, as shown in equation (6): (6) where, is the importance weight of character ; is the scene center coordinate; denotes the Euclidean distance.
[0039] The cost of the distribution cost function encourages uniform distribution of characters and suppresses the phenomenon of character aggregation, as shown in equation (7): (7) that is, the standard deviation of the distance between all pairs of characters is calculated, and the smaller the standard deviation, the more uniform, denotes the Euclidean distance between the i-th virtual character and the j-th virtual character, denotes the traversal of all virtual character pairs that satisfy the i index is less than the j index, which ensures that each unordered virtual character pair is only calculated once (to avoid repeated calculation of the distance of the same virtual character pair).
[0040] The cost of the region cost function divides the scene into M fixed grid regions, and if a region is not occupied by any character, it is penalized, as shown in equation (8): (8) where, denotes the number of characters contained in the k-th region; is the indicator function.
[0041] The orientation cost function achieves coordination and richness of orientation through the combination of deviation penalty and diversity encouragement, as shown in equation (9): (9) where, is the deviation cost function; is the diversity cost function.
[0042] The ideal orientation of each character is determined by its vector relative to the preset focus point (such as the center of the stage), as shown in equation (10): (10) The orientation deviation is calculated by the absolute value of the inverse cosine of the cosine of the angle, and the smaller the value, the closer the character orientation to the stage direction, as shown in equation (11): (11) The diversity cost function encourages all characters to be more diverse in the distribution, and prevents monotony in performance caused by consistency. The diversity cost function can be defined as the negative value of the standard deviation, as shown in equation (12): (12) Here, the standard deviation of the orientations of all characters is calculated by the diversity cost function. The greater the standard deviation, the more diverse the orientation distribution, and the smaller the function value (negative value), the lower the cost. Conversely, the more concentrated the orientation, the greater the function value, and the higher the cost. In some embodiments, entropy or variance can be used instead of standard deviation as a measure of diversity. The core logic remains the same and will not be illustrated here.
[0043] In some embodiments, according to the attribute information of the virtual character, the process of generating the adaptive weight through the pre-constructed weight prediction model can include: embedding and coding the category type features in the attribute information to obtain a category feature vector; normalizing the numerical type features to obtain a numerical feature vector; concatenating the category feature vector and the numerical feature vector to obtain a unified feature vector; inputting the unified feature vector into the weight prediction model to predict the weight, and obtaining the adaptive weight.
[0044] wherein the category type features are classification features in the virtual character attribute information that cannot be directly quantified, such as the role type, function category, identity level, etc. Embedding coding is a processing method for converting category type features into low-dimensional dense vectors that can be processed by computers, preserving the semantic association of the features. The numerical type features are features in the virtual character attribute information that can be directly quantified, such as the number of appearances, the duration of stay, the seniority score, etc. The category feature vector is a low-dimensional vector obtained by embedding coding, used to represent the category type features. The numerical feature vector is obtained by normalization processing, used to represent the numerical type features. The unified feature vector is a single vector formed by combining the category feature vector and the numerical feature vector according to a preset rule, used as the input of the weight prediction model.
[0045] Specifically, first, the attribute information of the virtual role is classified and processed to distinguish between categorical features and numerical features. For categorical features, such as the role type (e.g., lecturer, audience, host), functional category (e.g., interactive, display), identity level (e.g., senior, intermediate, junior), etc., an embedding encoding method is used for processing. Through a pre-trained embedding layer model, each categorical feature is mapped to a low-dimensional dense category feature vector, which can preserve the semantic association between different categories (e.g., the similarity of the category vector of the host and the lecturer is higher than that of the host and the ordinary audience), facilitating subsequent feature learning by the model. For numerical features, such as the historical appearance frequency of the role, the duration of stay in the scene, the number of user interactions, seniority score, etc., normalization processing is performed. According to the value range of the numerical features, the minimum-maximum normalization method is used to map the values of all numerical features to a fixed range of 0-1, eliminating the model training bias caused by the dimensional difference between different features (e.g., appearance frequency is several tens of times, and stay duration is several minutes), and obtaining standardized numerical feature vectors.
[0046] After completing the separate processing of categorical features and numerical features, the obtained category feature vectors and numerical feature vectors are spliced in a predetermined order. For example, first arrange all elements of the category feature vector, and then follow all elements of the numerical feature vector, forming a unified feature vector containing all attribute feature information, ensuring that the vector can fully reflect the attribute features of the virtual role.
[0047] The constructed unified feature vector is input into a pre-trained weight prediction model, which uses a lightweight deep neural network structure and has been trained with a large amount of virtual role attribute data labeled with importance weights. The model extracts and maps the unified feature vector, outputting the corresponding weight value, which is the adaptive weight used to adjust the spatial cost function and accurately reflects the importance level of the virtual role in the current scene layout.
[0048] Further, the process of obtaining the adaptive weight includes: inputting the unified feature vector into the weight prediction model for weight prediction to obtain the predicted weight of the current optimization layout; and dynamically calibrating the predicted weight using a smoothing update mechanism to obtain the adaptive weight.
[0049] Specifically, first, the predicted weight of the current optimization layout is obtained. The unified feature vector of each virtual role is input into the weight prediction model one by one, and the model outputs the initial weight value corresponding to each role based on the characteristics of the current scene and the role attributes, i.e., the predicted weight of the current optimization layout, which reflects the model's preliminary judgment of the importance of the current role.
[0050] Because AR scenes are dynamic, the state of virtual characters may change due to user behavior, scene transitions, and other factors. Relying solely on the predicted weights for the current iteration may lead to fluctuations or deviations. Therefore, a smoothing update mechanism is needed to dynamically calibrate the predicted weights. The core of the smoothing update mechanism is to adjust the predicted weights for the current iteration by combining historical weight information or preset stability rules, avoiding significant fluctuations in weights across different optimization iterations. For example, if the predicted weights for the current iteration differ significantly from those in previous optimization iterations, the smoothing mechanism will appropriately reduce the adjustment magnitude of the predicted weights for the current iteration, ensuring that the calibrated weights can respond to changes in the current scene while maintaining a certain level of stability.
[0051] Through the above smooth update and dynamic calibration process, the initial prediction weights are adjusted to weight values that are more in line with the actual needs of the scenario and have stronger stability. These weight values are the adaptive weights used to adjust the spatial cost function.
[0052] In a specific implementation, the training process of the weight prediction model includes: a. Construct a labeled dataset for weight prediction training, derived from typical scenarios such as conference presentations, interviews, and virtual theaters. Define a set of historical roles. Each historical character It has a set of attribute vectors:
[0053] Each attribute This represents a structured feature of a historical role, such as occupation type, organizational level, historical participation frequency, role seniority, and visual exposure; attributes include categorical (e.g., job title) and numerical (e.g., number of appearances) information. Domain experts label each role in the dataset with "importance," forming the true weights. , used as a training supervision signal.
[0054] b. The attribute information of historical characters is characterized by feature encoding and normalization to obtain the corresponding unified feature vector. For details, please refer to the above record, which will not be repeated here.
[0055] c. Weighted prediction model training: A lightweight deep neural network (such as ResNet-18) is used as the basic prediction structure, and the L1 norm loss is used as the supervision signal, as shown in equation (13): (13) in, This represents the domain expert's label for the importance of the role (value range [0,1]). These are the predicted weights for the model. The model is trained using a large amount of labeled data to minimize the absolute deviation between the predicted and actual weights, thus enabling the model to have accurate weight prediction capabilities.
[0056] In some embodiments, the process of dynamically calibrating the predicted weights using a smooth update mechanism to obtain the adaptive weights may include: Maintain a weight cache pool, which stores the historical weights corresponding to each virtual character in the historical optimized layout; for each virtual character in the current optimized layout, obtain the historical weight of each virtual character from the weight cache pool, and dynamically fuse the predicted weight with the historical weight to obtain the fused weight; use the fused weight as the adaptive weight, and update the adaptive weight to the weight cache pool for smooth updates in subsequent rounds.
[0057] The weight cache pool is a storage unit used to store the adaptive weights of virtual characters in historical optimization rounds, and supports reading, writing and updating operations of weights.
[0058] In this embodiment, when implementing the smooth update mechanism, a dedicated weight cache pool is first maintained. This cache pool is categorized and stored according to the optimization round and the virtual character identifier, saving the historical weights (i.e., the adaptive weights after dynamic calibration in that round) corresponding to each virtual character in each round of historical optimization layout, ensuring that the past weight information of each character can be quickly retrieved.
[0059] For each virtual character in the current optimized layout, the historical weight of that character in all historical optimization rounds is retrieved from the weight cache pool based on the character's identifier. If the character is newly added to the scene and has no historical weight record, a base weight value is assigned by default as a historical weight for calculation.
[0060] Subsequently, the predicted weights obtained in the current iteration are dynamically fused with the extracted historical weights. The fusion process must balance the adaptability of the predicted weights to the current scenario with the stability of the historical weights. For example, based on the time decay characteristics of weights, recent historical weights are given a higher reference weight, while older historical weights are given a lower reference weight, while retaining the core influence of the predicted weights. The fused weights are obtained through weighted calculation. The calculated fused weights are directly used as the adaptive weights for the current round, used to adjust the spatial cost function. Simultaneously, these adaptive weights are updated in the weight cache pool according to the role identifier and the current optimization round, overwriting or supplementing the historical weight records for that role, providing the latest historical data support for smooth updates in subsequent optimization rounds.
[0061] Furthermore, the process of obtaining the fusion weights may include: The absolute difference between the predicted weight and the mean of all historical weights is determined; based on the absolute difference and a preset sensitivity threshold, an influence factor is assigned to the predicted weight and each historical weight; the product of the predicted weight and each historical weight with their respective influence factors is summed to obtain the fused weight; wherein, the influence factor includes a first influence factor for the predicted weight and a second influence factor assigned to each historical weight, the first sum of the first influence factor and all second influence factors being 1; and the second influence factor of each historical weight decreases as the storage time increases; if the absolute difference is greater than or equal to the preset sensitivity threshold, the first influence factor is greater than the second sum of all second influence factors; if the absolute difference is less than the preset sensitivity threshold, the first influence factor is less than the second sum of all second influence factors.
[0062] The preset sensitivity threshold is a pre-defined critical value used to judge the degree of deviation of the predicted weights, determined based on scenario requirements and optimization experience. The influence factor is a coefficient used to adjust the proportion of predicted weights and historical weights in dynamic fusion, including a first influence factor and a second influence factor. The first influence factor is the influence factor assigned to the predicted weights, determining the contribution of the predicted weights in the fusion. The second influence factor is the influence factor assigned to each historical weight, determining the contribution of the corresponding historical weight in the fusion.
[0063] Specifically, the first step is to calculate the average of all historical weights, which is to perform an arithmetic average of all historical weights of the virtual character in the weight cache pool to obtain a value that reflects the overall level of historical weights.
[0064] Next, the absolute difference between the current predicted weight and the average of those weights is calculated. This difference is used to determine the degree of deviation between the current predicted weight and the historical weights. At the same time, a preset sensitivity threshold is retrieved from the system. This threshold is pre-set based on factors such as the dynamic characteristics of the AR scene and the need for optimized stability. It is used to distinguish whether the deviation of the predicted weight is a normal fluctuation or a significant change.
[0065] Based on the comparison between the absolute difference and a preset sensitivity threshold, a first influence factor is assigned to the predicted weights, and a second influence factor is assigned to each historical weight. The sum of the first influence factor and all second influence factors is fixed at 1 to ensure that the calculation of the fused weights conforms to the probability distribution rule. Simultaneously, the second influence factor for each historical weight follows a time decay rule; the longer a historical weight has been stored, the smaller its corresponding second influence factor, meaning that more recent historical weights have a greater impact on the fusion result.
[0066] If the absolute difference is greater than or equal to the preset sensitivity threshold, it indicates that there is a significant difference between the current predicted weight and the historical weight, which may be due to a significant change in the scenario (such as a change in role function or a shift in user focus). In this case, a higher first influence factor is assigned, and the first influence factor is greater than the sum of all second influence factors, so that the predicted weight plays a dominant role in the fusion process to quickly respond to scenario changes. If the absolute difference is less than the preset sensitivity threshold, it indicates that the predicted weight deviates little from the historical weight, and the scenario is in a relatively stable state. In this case, a lower first influence factor is assigned, and the first influence factor is less than the sum of all second influence factors, so that the historical weight plays a major role in the fusion process to ensure the stability of the weight.
[0067] Finally, the predicted weights are multiplied by the first influence factor, each historical weight is multiplied by its respective second influence factor, and all the product results are summed to obtain the fusion weight, which is used as the adaptive weight for the current round.
[0068] In some embodiments, the iterative optimization process may include: By using a neural network model based on an attention mechanism (such as the Transformer model), an initial layout of the global structure is generated based on the semantic features of the current augmented reality scene and the initial state of the virtual character, which serves as the current layout state. Based on the current layout, all virtual characters are divided into multiple spatial function groups according to spatial proximity or functional similarity. In each iteration, a group-level action is sampled from the preset action space and executed for each of the spatial function groups. The group-level action is used to coordinate the position and orientation of all virtual characters in the group in three-dimensional space. Based on the updated layout state and the joint objective function, a group-level reward is calculated for each spatial function group that has performed a group-level action; the group-level reward is the negative of the sum of the spatial cost function and the orientation cost function corresponding to the spatial function group. An optimization algorithm based on policy gradient is adopted to update the action selection strategy of each spatial function group according to the group-level reward, and the updated layout state is used as the current layout state for the next iteration until the preset convergence condition is met, and the final optimized layout is output.
[0069] The initial state is the original state representation of the virtual character before optimization begins, including the initial three-dimensional spatial position, facing direction, and occlusion state.
[0070] The initial layout is an initial character layout with global structural rationality generated by a neural network model, providing a foundation for subsequent iterative optimization.
[0071] Spatial function groups are sets of roles divided according to the spatial proximity or functional similarity of virtual characters, which facilitates coordinated adjustments.
[0072] Spatial proximity refers to the degree of closeness between virtual characters in three-dimensional space; the closer the distance, the higher the spatial proximity.
[0073] Functional similarity refers to the degree of similarity in the functions performed by virtual characters in a scene. The closer the functions are, the higher the functional similarity.
[0074] The preset motion space is a predefined set of actions used to adjust the position and orientation of a virtual character, including actions such as translation and rotation.
[0075] Group-level actions are sampled from a preset action space and are used to coordinate and adjust the roles within the entire functional group.
[0076] Group-level rewards are indicators used to evaluate the layout optimization effect after a spatial functional group performs group-level actions, and are associated with the cost of the joint objective function.
[0077] In detail, the initial layout is generated. An attention-based neural network model is used, which can automatically focus on the key semantic features of the current AR scene (e.g., when the scene type is a virtual meeting, it focuses on features such as the central area of the meeting and the distribution of seats; when the scene type is a city tour, it focuses on features such as the location of attractions and the tour route). At the same time, it combines the initial state of all virtual characters (initial position, orientation, occlusion state), and through the model's feature extraction and layout generation capabilities, outputs an initial layout with a globally reasonable structure. This layout avoids obvious collision and boundary crossing problems in the initial state and serves as the current layout state for subsequent iterative optimization.
[0078] The initial layout aims to provide a well-structured initial solution for subsequent optimization processes, thereby reducing invalid search paths and shortening convergence time. Scene semantic features may include, for example, stage shape, obstacle locations, speaker locations, etc. The initial coarse layout scheme can be spatial regularity initialization or uniform distribution. After structured encoding, it is input into the Transformer model, and the initial estimated state is output, as shown in equation (14): (14) in, This represents the initial optimization point, i.e., the initial state of the virtual character.
[0079] Based on the generated current layout state, all virtual characters are divided into spatial functional groups. The grouping method can be either spatial proximity or functional similarity: if spatial proximity is selected, characters that are close together in 3D space are grouped together; if functional similarity is selected, characters that perform the same or similar functions in the scene (e.g., both are narrators, both are audience members) are grouped together. By grouping these characters, the large-scale character layout optimization problem is decomposed into multiple smaller-scale group optimization problems, reducing the optimization complexity.
[0080] Specifically, to reduce the joint search dimensions, the N virtual characters are divided into G spatial function groups. The criteria include: spatial proximity (such as location-based clustering); and functional similarity (such as identity, role attributes, etc.).
[0081] The formula for defining the local state of the g-th spatial functional group is as shown in equation (15): (15) The formula for the overall global state is as shown in equation (16): (16) in, This represents the "union" operation of sets, which merges all local state sets from group 1 to group G (since each role belongs to only one group, there are no duplicates after merging).
[0082] In the iterative optimization process, in each iteration, a group-level action is sampled from the preset action space for each spatial function group. The preset action space contains various collaborative actions for adjusting character positions and orientations, such as overall translation of characters within the group, overall rotation of characters, and adjustment of spacing between characters within the group. After executing the group-level action, the 3D spatial positions and orientations of all virtual characters within the group will be collaboratively adjusted according to the action rules, generating an updated layout state.
[0083] Specifically, for each group of spatial functions Define the action as a fine-tuning operation in three-dimensional space. The corresponding formula is shown in equation (17): (17) in, This represents the fine-tuning of the spatial function group's position along the x-axis in three-dimensional space (i.e., the change in the x-coordinate of the character within the group). This represents the positional adjustment of a spatial function group in the z-axis direction in three-dimensional space (i.e., the change in the z-coordinate of a character within the group). This represents the amount of fine-tuning of the orientation of all characters (i.e., the change in the character's orientation angle).
[0084] Each group of fine-tuning operations Determined by the policy function, its corresponding formula is shown in equation (18): (18) in, πg ( ) represents the strategy function corresponding to the g-th spatial function group, which calculates and outputs the action to be performed by the group based on the current state of the group. This represents the local state of the g-th spatial functional group.
[0085] Based on the updated layout state, and combined with the aforementioned joint objective function, the spatial cost function value and orientation cost function value corresponding to each spatial function group are calculated. The sum of the two cost function values and the inverse of the sum are taken to obtain the group-level reward for that spatial function group. The higher the group-level reward value, the better the layout optimization effect after the group performs the current group-level action.
[0086] Specifically, to measure the effect of each set of actions, the optimized cost function of that set is calculated and its negative value is taken as the reward. The corresponding calculation formula is shown in equation (19): (19) The global reward is the sum of the rewards for all groups, as shown in formula (20): (20) An optimization algorithm based on policy gradients is employed to adjust the action selection strategy of each spatial function group according to its group-level reward. For example, if a group receives a high group-level reward after performing a certain group-level action, the optimization algorithm will increase the probability of that group selecting that type of action in subsequent iterations; conversely, it will decrease the selection probability. Simultaneously, the updated layout state is used as the current layout state for the next iteration, and the process of group action sampling, execution, reward calculation, and policy updating is repeated.
[0087] Specifically, group-level strategy optimization algorithms (such as GRPO) can be used to update model parameters. The gradient update formula is as shown in equation (21):
[0088] The strategy parameters are updated as follows:
[0089] The iteration process terminates when the preset convergence condition is met. The convergence condition can be set as follows: the cost of the joint objective function remains stable for multiple consecutive iterations (i.e., the decrease in the joint objective function is less than a threshold), or the number of iterations reaches a preset upper limit. The output layout state at this time is the final optimized layout.
[0090] In practical applications, if virtual characters that closely match the user's real-time behavior and areas of interest are not effectively selected, the relevance of the recommended content may be insufficient, affecting the immersive experience of the user and the final effect of layout optimization.
[0091] To address this, the present invention provides the following technical solution: Before constructing state representations for multiple virtual characters in an augmented reality scene, the following steps are also included: Real-time acquisition of user behavior perception data in augmented reality scenarios; the behavior perception data includes the user's gaze direction, head posture, interaction actions, and dwell time; Based on the behavioral perception data, the user's real-time attention area in the current scene is estimated through a pre-trained visual attention mechanism or a deep neural network model. Calculate the correlation score between each candidate virtual character and the real-time attention area; Candidate virtual characters are sorted according to the correlation score, and multiple virtual characters with scores higher than a preset threshold are selected.
[0092] Specifically, the AR device first collects user behavior data in real time through sensors (such as cameras, gyroscopes, gesture recognition modules, etc.), including the user's gaze direction (obtained through eye-tracking technology), head posture (detected by the gyroscope to measure the pitch and yaw angles of the head), interactive actions (such as the user clicking on a virtual character with gestures, asking for relevant information by voice, etc.), and the duration of the user's stay in different areas of the scene.
[0093] The collected behavioral perception data is then input into a pre-trained visual attention mechanism or deep neural network model. This model has been trained using a large number of user behavior samples and can determine the user's attention intent based on the characteristics of the behavioral data. For example, if a user's gaze is continuously directed towards a certain area of the scene, their head posture is facing that area, and they linger there for a relatively long time, the model will identify that area as the user's real-time attention area. If the user clicks on a virtual character through an interactive action, the model will identify the area where the character is located and related areas as the real-time attention area.
[0094] Then, for all candidate virtual characters in the AR scene, the relevance score between each character and the real-time attention area is calculated. The calculation of the relevance score needs to comprehensively consider factors such as the spatial distance between the character and the attention area, the matching degree between the character's semantics and the user's interests, and the consistency between the character's orientation and the user's line of sight. The specific calculation method will be explained in detail in the following content.
[0095] Finally, based on the calculated relevance scores, all candidate virtual characters are sorted in descending order. The higher the score, the stronger the correlation with the user's real-time focus area. A preset threshold is set, determined based on factors such as the total number of virtual characters in the scene and user attention needs. Virtual characters with scores higher than this preset threshold after sorting are selected. These characters are the target virtual characters for subsequent state representation construction, ensuring that subsequent layout optimization focuses on the core characters that users are concerned with, thus improving the targeting of the layout.
[0096] In a specific implementation, the process of calculating the relevance score includes: The system obtains the spatial distance between the candidate virtual character's location and the real-time attention area, the matching degree between the candidate virtual character's semantic tags and the user's historical interests, and the angle between the candidate virtual character's orientation and the user's line of sight. Specifically, the Euclidean distance between the candidate virtual character's location and the geometric center of the real-time attention area is calculated as the spatial distance; the semantic tags of the candidate virtual character are compared with interest tags extracted from the user's historical behavior data, and their overlap is calculated as the matching degree; the minimum angle between the candidate virtual character's facing direction vector and the user's line of sight direction vector is calculated as the angle. A spatial attenuation factor is generated based on the spatial distance, an orientation correction coefficient is generated based on the included angle, and a semantic base score is generated based on the matching degree. Specifically, according to a preset attenuation function, the spatial distance is mapped to a value between zero and one as the spatial attenuation factor; the closer the distance, the higher the mapped spatial attenuation factor value. According to a preset mapping rule, the included angle is mapped to a value between zero and one as the orientation correction coefficient; the smaller the included angle, the more consistent the character's orientation is with the user's line of sight, and the higher the mapped orientation correction coefficient value. Based on a preset scoring rule, an initial semantic base score is assigned to the candidate virtual character according to the degree of matching between the semantic tag and the user's historical interests; the higher the degree of matching, the higher the assigned semantic base score. The spatial attenuation factor and the orientation correction coefficient are combined to obtain the real-time attention gain coefficient; specifically, the spatial attenuation factor and the orientation correction coefficient are multiplied to obtain the real-time attention gain coefficient. The semantic base score is modulated using the real-time attention gain coefficient to obtain the modulated semantic interest score; specifically, the semantic base score is multiplied by the real-time attention gain coefficient to obtain the modulated semantic interest score. The modulated semantic interest score is weighted and combined with the basic visibility score calculated based on spatial distance to output the final relevance score; where the closer the spatial distance, the higher the basic visibility score.
[0097] Specifically, the first step is to obtain three core evaluation metrics: spatial distance, matching degree, and angle. For spatial distance, the geometric center coordinates of the real-time focus area are first determined, and the Euclidean distance (i.e., straight-line distance) between the three-dimensional spatial position of each candidate virtual character and the geometric center is calculated. This distance is used for subsequent calculations.
[0098] For matching degree, semantic tags (such as the type of exhibit and functional attributes corresponding to the character) are extracted for each candidate virtual character. At the same time, interest tags are extracted from the user's historical behavior data (such as exhibits that have been followed and characters that have been interacted with). The semantic tags and interest tags are compared, and the proportion of the number of identical tags to the total number of tags is counted. This proportion is the matching degree.
[0099] For the included angle, the facing direction of the candidate virtual character and the user's line of sight are converted into direction vectors respectively, and the minimum angle between the two direction vectors is calculated. This angle is the included angle used for subsequent calculations.
[0100] Next, corresponding coefficients and scores are generated based on the three indicators mentioned above. For spatial distance, it is mapped to a spatial attenuation factor between 0 and 1 using a preset attenuation function; the closer the distance, the higher the spatial attenuation factor, and vice versa. For angle, it is mapped to an orientation correction coefficient between 0 and 1 using a preset mapping rule; the smaller the angle, the more consistent the character's orientation is with the user's line of sight, and the higher the orientation correction coefficient, and vice versa. For matching degree, a semantic base score is assigned according to a preset scoring rule; the higher the matching degree, the higher the semantic base score. For example, when the matching degree is 100%, the semantic base score is full; when the matching degree is 0, the semantic base score is 0.
[0101] The generated spatial attenuation factor is then fused with the orientation correction coefficient by multiplication to obtain the real-time attention gain coefficient, which comprehensively reflects the spatial correlation and orientation consistency between the character and the user's real-time attention area.
[0102] Then, the semantic base score is modulated using the real-time attention gain coefficient, and the modulated semantic interest score is obtained by multiplication. This score reflects both the relationship between the role and the user's historical interests and the user's current real-time attention status.
[0103] Finally, a basic visibility score is calculated based on spatial distance. The closer the spatial distance, the better the basic visibility of the character, and the higher the basic visibility score. The modulated semantic interest score and the basic visibility score are then weighted and combined according to preset weights. For example, each is assigned a preset weight coefficient, multiplied, and summed to obtain the final relevance score.
[0104] Based on the same general inventive concept, this invention also protects an augmented reality content recommendation and expressive presentation system based on user behavior perception. The augmented reality content recommendation and expressive presentation system based on user behavior perception provided by this invention will be described below. The augmented reality content recommendation and expressive presentation system based on user behavior perception described below can be referred to in correspondence with the augmented reality content recommendation and expressive presentation method based on user behavior perception described above.
[0105] Figure 2 This is a schematic diagram of the structure of the augmented reality content recommendation and expressive presentation system based on user behavior awareness provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the augmented reality content recommendation and expressive presentation system based on user behavior perception in this embodiment includes a first construction module 21, a second construction module 22, a generation module 23, an optimization module 24, and a display module 25.
[0106] In a specific implementation, the first construction module 21 is used to construct the state representation of multiple virtual characters in the augmented reality scene; the state representation includes the three-dimensional spatial position, facing direction and occlusion state of each virtual character; The second construction module 22 is used to construct a joint objective function for evaluating the rationality of the layout of virtual characters; the joint objective function includes a spatial cost function and an orientation cost function, the spatial cost function is used to quantify the rationality of the spatial relationship between characters, and the orientation cost function is used to quantify the coordination and diversity of the character orientations; The generation module 23 is used to generate adaptive weights based on the attribute information of the virtual character through a pre-built weight prediction model, and to use the adaptive weights to adjust the cost terms related to the importance of the character in the spatial cost function; Optimization module 24 is used to iteratively optimize the position and orientation of all virtual characters based on the joint objective function to obtain an optimized layout; Display module 25 is used to expressively present the optimized layout in an augmented reality scene.
[0107] It should be noted that all relevant information that may be involved in the various embodiments of the present invention is processed in strict accordance with the requirements of laws and regulations, following the principles of legality, legitimacy, and necessity, based on the reasonable purpose of the business scenario, and is information that users actively provide or generate during the use of the product / service, as well as information obtained with user authorization.
[0108] The information processed by this invention may vary depending on the specific product / service scenario and should be based on the specific scenario in which the user uses the product / service. This may involve user account information, device information, or other related information. This invention will treat the relevant information and its processing with the utmost diligence.
[0109] This invention places great emphasis on the security of relevant information and has adopted reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent unauthorized access, public disclosure, use, modification, damage or loss of relevant information.
[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for augmented reality content recommendation and expressive presentation based on user behavior perception, characterized in that, The method comprises the following steps: constructing a state representation of a plurality of virtual roles in an augmented reality scene; the state representation comprises a three-dimensional spatial position, a face direction and an occlusion state of each virtual role; constructing a joint objective function for evaluating the rationality of the virtual role layout; the joint objective function comprises a spatial cost function and an orientation cost function, the spatial cost function is used to quantify the rationality of the spatial relationship between the roles, and the orientation cost function is used to quantify the coordination and diversity of the orientations of the roles; according to the attribute information of the virtual roles, an adaptive weight is generated by a pre-constructed weight prediction model, and the adaptive weight is used to adjust a cost term related to the importance of the roles in the spatial cost function; based on the joint objective function, the positions and orientations of all virtual roles are iteratively optimized to obtain an optimized layout; expressively presenting the optimized layout in the augmented reality scene. 2.The user behavior-aware augmented reality content recommendation and expressive presentation method of claim 1, wherein, The method of constructing a joint objective function for evaluating the rationality of the virtual role layout comprises: constructing the spatial cost function, which is the sum of a collision cost function, a boundary cost function, a center cost function, a distribution cost function and a region cost function, wherein the collision cost function is calculated based on the intersection-over-union of the bounding boxes of the roles, the boundary cost function penalizes virtual roles close to the scene boundary, the center cost function encourages virtual roles to approach the scene center according to the corresponding importance weight, the distribution cost function encourages virtual roles to be uniformly distributed in space, and the region cost function encourages virtual roles to cover multiple divided regions of the scene; constructing the orientation cost function, which is the sum of a deviation cost function and a diversity cost function, wherein the deviation cost function calculates the deviation of the current orientation of each virtual role from its ideal orientation, and the ideal orientation is determined by the position of the virtual role relative to a preset focus point, and the diversity cost function encourages the distribution of all virtual role orientations to have diversity. 3.The user behavior-aware augmented reality content recommendation and expressive presentation method of claim 1, wherein, According to the attribute information of the virtual roles, an adaptive weight is generated by a pre-constructed weight prediction model, comprising: embedding and encoding the category type features in the attribute information to obtain a category feature vector; normalizing the numerical type features to obtain a numerical feature vector; concatenating the category feature vector and the numerical feature vector to obtain a unified feature vector; inputting the unified feature vector into the weight prediction model for weight prediction to obtain the adaptive weight. 4.The user behavior-aware augmented reality content recommendation and expressive presentation method of claim 3, wherein, inputting the unified feature vector into the weight prediction model for weight prediction to obtain the adaptive weight, comprising: inputting the unified feature vector into the weight prediction model for weight prediction to obtain the predicted weight of the current optimized layout; using a smoothing update mechanism to dynamically calibrate the predicted weight to obtain the adaptive weight. 5.The user behavior-aware augmented reality content recommendation and expressive presentation method of claim 4, wherein, using a smoothing update mechanism to dynamically calibrate the predicted weight to obtain the adaptive weight, comprising: maintaining a weight cache pool, the weight cache pool saving historical weights corresponding to each virtual role in historical optimized layouts; For each virtual character in the current round of layout optimization, a historical weight of each virtual character is obtained from the weight cache pool, and the predicted weight is dynamically fused with the historical weight to obtain a fused weight; The fused weight is taken as the adaptive weight, and the adaptive weight is updated to the weight cache pool for subsequent round of smooth update. 6.The user behavior-aware augmented reality content recommendation and expressive presentation method of claim 5, wherein, The dynamic fusion of the predicted weight and the historical weight to obtain a fused weight comprises: determining an absolute difference between the predicted weight and a weight mean of all the historical weights; assigning each of the predicted weight and each of the historical weights with a respective influence factor according to the absolute difference and a preset sensitivity threshold; multiplying and summing the predicted weight and each of the historical weights with the respective influence factor to obtain the fused weight; wherein the influence factor includes a first influence factor of the predicted weight and a second influence factor of each of the historical weights, a first sum of the first influence factor and all the second influence factors is 1; and the second influence factor of each of the historical weights decreases as the storage time increases; if the absolute difference is greater than or equal to the preset sensitivity threshold, the first influence factor is greater than a second sum of all the second influence factors; if the absolute difference is less than the preset sensitivity threshold, the first influence factor is less than the second sum of all the second influence factors. 7.The user behavior perception based augmented reality content recommendation and expressive presentation method of claim 1, wherein, Based on the joint objective function, the positions and orientations of all virtual characters are iteratively optimized to obtain an optimized layout, comprising: generating an initial layout of global structure based on the semantic features of the current augmented reality scene and the initial state of the virtual characters through a neural network model based on attention mechanism, as a current layout state; based on the current layout state, all virtual characters are divided into multiple spatial functional groups according to spatial proximity or functional similarity; in each iteration, a group-level action is sampled from a preset action space for each spatial functional group and executed, which is used to cooperatively adjust the positions and orientations of all virtual characters in the group in three-dimensional space, thereby generating an updated layout state; based on the updated layout state and the joint objective function, a group-level reward is calculated for each spatial functional group that has executed a group-level action; the group-level reward is the inverse number of the sum of the spatial cost function and the orientation cost function corresponding to the spatial functional group; an optimization algorithm based on policy gradient is used to update the action selection policy of each spatial functional group according to the group-level reward, and the updated layout state is taken as the current layout state of the next iteration for iteration until a preset convergence condition is met, and the final optimized layout is output. 8.The user behavior perception based augmented reality content recommendation and expressive presentation method of claim 1, wherein, Before constructing the state representation of multiple virtual characters in the augmented reality scene, further comprising: real-time acquisition of behavior perception data of the user in the augmented reality scene; the behavior perception data includes the user's gaze direction, head pose, interaction action and stay time; based on the behavior perception data, the real-time attention area of the user in the current scene is estimated through a pre-trained visual attention mechanism or a deep neural network model; calculate a relevance score between each candidate virtual role and the real-time attention area; sort the candidate virtual roles according to the relevance scores, and select a plurality of virtual roles with scores higher than a preset threshold. 9.The user behavior-aware augmented reality content recommendation and expressive presentation method of claim 8, wherein, The method for calculating the relevance score between each candidate virtual role and the real-time attention area comprises: obtaining a spatial distance between a position of the candidate virtual role and the real-time attention area, a matching degree between a semantic label of the candidate virtual role and a historical interest of the user, and an included angle between an orientation of the candidate virtual role and a line of sight direction of the user; generating a spatial attenuation factor based on the spatial distance, an orientation correction coefficient based on the included angle, and a semantic basic score based on the matching degree; fusing the spatial attenuation factor and the orientation correction coefficient to obtain a real-time attention gain coefficient; modulating the semantic basic score using the real-time attention gain coefficient to obtain a modulated semantic interest score; combining the modulated semantic interest score and a basic visibility score calculated based on the spatial distance by weighting to output a final relevance score. 10.The user behavior perception based augmented reality content recommendation and expressive presentation method of claim 9, wherein, The method for generating a spatial attenuation factor based on the spatial distance comprises: mapping the spatial distance to a value between zero and one as the spatial attenuation factor according to a preset attenuation function; wherein the closer the distance, the higher the value of the spatial attenuation factor obtained by mapping; The method for generating an orientation correction coefficient based on the included angle comprises: mapping the included angle to a value between zero and one as the orientation correction coefficient according to a preset mapping rule; wherein the smaller the included angle, the more consistent the role orientation with the line of sight, and the higher the value of the orientation correction coefficient obtained by mapping; The method for generating a semantic basic score based on the matching degree comprises: assigning an initial semantic basic score value to the candidate virtual role according to the matching degree between the semantic label and the historical interest of the user according to a preset scoring rule; the higher the matching degree, the higher the semantic basic score value assigned.
Citation Information
Patent Citations
Label layout method for quickly positioning virtual scene based on user perception
CN118840514A
Digital exhibition product display method based on deep learning
CN120215704A
Vent travel element universe virtual character cooperation system and method based on large space interaction
CN120653101A
Analog parameter test system based on digital test equipment
CN120779138A
Real-time data distribution and intelligent caching method and system based on edge computing
CN121037449A