Label layout method for fast positioning in virtual scenes based on user perception
By determining the user's interest and perception in the virtual reality scene and updating the guide tag position, the problems of guide tag occlusion and poor clustering are solved, and rapid user positioning and efficient search are achieved.
Patent Information
- Application Number
- CN202410980649.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-07-22
AI Technical Summary
In virtual reality scenarios, existing label layout methods easily lead to occlusion between guiding labels, making them difficult for users to visually perceive. In addition, when the clustering effect is poor, it is difficult to represent the entire cluster and does not consider user interest, making it difficult to filter out labels that users are interested in in a timely manner.
By determining the user interest level of each scene object in the virtual scene, the target perception object that meets the preset conditions is selected, and the user perception force and camera force are generated based on the perception time mapping function and viewport coordinates. Combined with the dynamic adjustment force, the guidance label position is updated to facilitate user perception.
It improves the user's ability to perceive the guide tags of interest, shortens the target search task time, reduces the incidence of motion sickness, and improves positioning efficiency and usability.
Smart Images

Figure CN118840514B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the fields of computer graphics and virtual reality, and more particularly to a label layout method for rapid positioning of a virtual scene based on user perception. Background Art
[0002] In virtual reality scenes, placing guide labels can accelerate the process of users quickly finding their target objects among a large number of scene objects, significantly improving the efficiency of target search tasks. Currently, label placement typically involves clustering similar objects, simplifying and deduplicating guide labels in crowded scenes based on the representative objects in each cluster, and then placing these deduplicated guide labels at the local 3D spatial location of the objects.
[0003] However, in practice, it is found that when the above method is used for label layout, the following technical problems often occur:
[0004] First, when objects in a scene are compactly arranged, placing guide labels in the local 3D spatial location of the objects can easily lead to occlusion between the guide labels, making them difficult for the user to visually perceive.
[0005] Second, when the clustering effect is poor, objects of different categories are easily mistakenly classified into one category, the selected objects are difficult to represent the entire cluster, and the user's interest is not taken into account, which makes it difficult to filter the various guide tags that the user is interested in in a timely manner.
[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0007] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] Some embodiments of the present disclosure propose a label layout method for rapid positioning of a virtual scene based on user perception to solve one or more of the technical problems mentioned in the above background technology section.
[0009] Some embodiments of the present disclosure provide a label layout method for rapid positioning of virtual scenes based on user perception, the method comprising: determining the user interest corresponding to each scene object in the virtual scene within a target time period, wherein each scene object corresponds to a guide label, and the attribute information corresponding to the guide label includes viewport coordinates, label visibility, and environmental contrast; selecting scene objects that meet preset interest conditions from the various scene objects included in the above-mentioned virtual scene as target perception objects, and obtaining a target perception object set; for each target perception object in the above-mentioned target perception object set, performing the following steps: based on a pre-constructed perception time mapping function and the user interest corresponding to the target perception object, determining the user perception corresponding to the target guide label, wherein the target guide label is the guide label corresponding to the above-mentioned target perception object, and the above-mentioned perception time mapping function characterizes the mapping between the guide label and the user perception time. relationship, the user perception force is the force used to move the guide tag to the three-dimensional space position where the user perception time is the shortest; based on the viewport coordinates corresponding to the above-mentioned target guide tag, the camera force corresponding to the above-mentioned target guide tag is determined, wherein the camera force is the force used to keep the guide tag within the user's field of view; the sum of the above-mentioned user perception force and the above-mentioned camera force is determined as the user perceived attraction corresponding to the above-mentioned target guide tag; based on the dynamic adjustment force corresponding to the above-mentioned target guide tag and the above-mentioned user perceived attraction, a tag force is generated, wherein the dynamic adjustment force is a pre-generated force that determines the positional relationship between each guide tag in the virtual scene and between the guide tag and the target perception object according to the dynamic potential field; based on the determined each tag force, the position of each target guide tag corresponding to the above-mentioned target perception object set is updated to obtain a tag position update information set for layout of the above-mentioned each target guide tag.
[0010] Embodiments of the present disclosure have the following advantageous effects: Through the user-perception-based label layout method for rapid virtual scene location, some embodiments of the present disclosure enable users to promptly perceive each guide label. Specifically, the reason why each guide label is difficult for the user to visually perceive is that when objects in a scene are compactly arranged, placing guide labels at local three-dimensional spatial locations of the objects can easily lead to occlusion between the guide labels, making each guide label difficult to visually perceive. Based on this, some embodiments of the present disclosure provide a user-perception-based label layout method for rapid virtual scene location. First, the user's interest level for each scene object in the virtual scene during a target time period is determined. Each scene object corresponds to a guide label, and attribute information associated with the guide label includes viewport coordinates, label visibility, and environmental contrast. This allows the user's interest level in each object in the virtual scene to be determined. Second, scene objects that meet preset interest criteria are selected from the various scene objects included in the virtual scene as target perception objects, thereby obtaining a target perception object set. This facilitates the subsequent layout of guide labels for the objects of greatest interest to the user. Then, for each target perception object in the target perception object set, the following steps are performed: based on a pre-constructed perception time mapping function and the user interest level corresponding to the target perception object, the user perception force corresponding to the target guide tag is determined, wherein the target guide tag is the guide tag corresponding to the target perception object, the perception time mapping function represents the mapping relationship between the guide tag and the user perception time, and the user perception force is the force used to move the guide tag to the three-dimensional spatial position with the minimum user perception time; based on the viewport coordinates corresponding to the target guide tag, the camera force corresponding to the target guide tag is determined, wherein the camera force is the force used to keep the guide tag within the user's field of view; the sum of the user perception force and the camera force is determined as the user perceived attraction corresponding to the target guide tag; based on the dynamic adjustment force corresponding to the target guide tag and the user perceived attraction, a tag action force is generated, wherein the dynamic adjustment force is a pre-generated force that determines the positional relationship between each guide tag in the virtual scene and between the guide tag and the target perception object according to the dynamic potential field. Thus, for each target perception object that the user is most interested in, the user perceived attraction can be determined based on the user interest level, the perception time, and the viewport constraint, and combined with the dynamic adjustment force to ultimately obtain the tag action force used for tag layout. Finally, based on the determined tag interaction forces, the positions of the target guidance tags corresponding to the target perception object set are updated to obtain a tag position update information set for the layout of the target guidance tags. Thus, based on the tag interaction forces, the guidance tags of interest to the user can be placed in the three-dimensional spatial location that minimizes the user's perception time.Therefore, in some embodiments of the present disclosure, the label layout method for rapid virtual scene positioning based on user perception determines the user perceived attractiveness according to the user's interest in the scene object of interest and the user's perception time of the guide label, which can facilitate the layout of the guide label in a three-dimensional space position that is more easily perceived by the user. As a result, the user can perceive each guide label of interest and the corresponding scene object in a timely manner. In addition, because the guide labels of virtual objects that the user is more interested in can be placed in a position with a shorter user perception time, the execution time of the target search task can be shortened, the user's positioning efficiency can be improved in real time, the user's task burden can be reduced, the usability can be improved, and the incidence of motion sickness can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0012] Figure 1 is a flowchart of some embodiments of a label layout method for rapid positioning of a virtual scene based on user perception according to the present disclosure;
[0013] Figure 2 This is an example diagram of the outline and environment image corresponding to an example scene object obtained according to the label layout method for rapid positioning of a virtual scene based on user perception disclosed in the present invention;
[0014] Figure 3 This is a schematic diagram of the forces acting on guide labels corresponding to example scene objects in a label layout method for rapid positioning of a virtual scene based on user perception according to the present disclosure;
[0015] Figure 4 This is an example diagram of a label layout before the disclosed label layout method for rapid virtual scene positioning based on user perception is used;
[0016] Figure 5 1 is an example diagram of a label layout using the user perception-based label layout method for rapid virtual scene positioning disclosed herein. DETAILED DESCRIPTION
[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0023] Figure 1 The process 100 of some embodiments of the label layout method for rapid positioning of a virtual scene based on user perception according to the present disclosure is shown. The label layout method for rapid positioning of a virtual scene based on user perception includes the following steps:
[0024] Step 101: Determine the user interest level corresponding to each scene object in the virtual scene within a target time period.
[0025] In some embodiments, an entity (e.g., a head-mounted display device) executing a label layout method for rapid virtual scene positioning based on user perception can determine the user interest level corresponding to each scene object in the virtual scene within a target time period through various methods. Each scene object can correspond to a guide label. A scene object can be an object in the virtual scene. A scene object can be associated with a scene object identifier. The scene object identifier can be a unique identifier for the scene object. A guide label can be a graphic with the scene object identifier. Attribute information corresponding to the guide label can include viewport coordinates, label visibility, and environmental contrast. The viewport coordinates can be the three-dimensional coordinates corresponding to the position of the guide label in the viewport coordinate system. The label visibility can be the ratio of the number of accessible sampling points corresponding to the guide label to the total number of sampling points. The sampling points can be obtained by sampling the label outline. For example, when the label outline is rectangular, sampling is performed on all four sides of the rectangle. When a ray is projected from the camera angle toward a sampling point, if the ray does not collide with an obstacle during its travel, the sampling point is considered an accessible sampling point. The environmental contrast can represent the degree of color difference between the guide label and the background environment. The target time period may be a time period of a preset duration that is earlier than the current moment and adjacent to the current moment. The preset duration may be a pre-set duration. For example, the preset duration may be 1 second. The user interest level may represent the user's level of interest in the scene object. It should be noted that, during the target time period, the execution entity may obtain continuous frame views of the virtual scene, and each frame view may be an image of the scene as the user gazes at the scene object in the virtual scene.
[0026] In some optional implementations of some embodiments, the execution entity may determine the user interest level corresponding to each scene object in the virtual scene within the target time period through the following steps:
[0027] The first step is to classify each scene object in the virtual scene to obtain a set of target sight intersection objects and a set of non-target sight intersection objects. Each target sight intersection object can be an object that the user is gazing at during a target time period. Each target sight intersection object can correspond to a user's gaze duration. The user gaze duration can be the length of time the user gazes at the target sight intersection object. For each scene object in the virtual scene, perform the following steps:
[0028] In a first sub-step, in response to determining that the scene object satisfies a preset gaze condition, the scene object is determined as a target line of sight intersection object. The preset gaze condition may be that the scene object is the closest object that intersects the center of the user's line of sight. The center of the user's line of sight may be the midpoint of the screen when the user puts on the head-mounted display device. The closest object may be the first object in the virtual scene that intersects the center of the user's line of sight.
[0029] The second sub-step is, in response to determining that the scene object does not meet the preset gaze condition, determining the scene object as a non-target sight-intersection object.
[0030] In practice, during the target time period, when the user's viewpoint changes, the target sight intersection object is also dynamically updated as the user's sight center changes. Therefore, there are multiple target sight intersection objects during the target time period.
[0031] The second step is to determine, for each target intersection object in the target intersection object set, the ratio of the user's gaze duration corresponding to the target intersection object to a preset duration as the current interest level of the target intersection object. The current interest level represents the user's current level of interest in the object in the virtual scene.
[0032] Step 3: For each non-target sight line intersection object in the non-target sight line intersection object set, perform the following steps:
[0033] In a first sub-step, a similarity analysis is performed between the non-target sight intersection object and each target sight intersection object in the target sight intersection object set to obtain an object similarity information set. The object similarity information in the object similarity information set can represent the degree of similarity between the non-target sight intersection object and the scene object of interest to the user. The object similarity information set can be obtained by performing a similarity analysis between the non-target sight intersection object and each target sight intersection object in the target sight intersection object set using various methods.
[0034] In some optional implementations of some embodiments, the execution entity may perform similarity analysis on the non-target sight-line intersection object and each target sight-line intersection object in the target sight-line intersection object set through the following steps to obtain an object similarity information set:
[0035] For each target line of sight intersection object in the above target line of sight intersection object set, perform the following steps:
[0036] Step 1: Determine the contour similarity and environmental similarity between the target sight intersection object and the non-target sight intersection object. The contour similarity can represent the degree of similarity between the contours of the target sight intersection object and the non-target sight intersection object. The environmental similarity can represent the degree of similarity between the background environments of the target sight intersection object and the non-target sight intersection object. The contour similarity and environmental similarity between scene objects can be determined by the following steps:
[0037] Sub-step 1: Determine contour similarity. First, image capture is performed. For each scene object in the target sight intersection object and the non-target sight intersection object, a virtual camera can be placed at a predetermined distance directly above, directly in front of, and directly to the left of the scene object, respectively, to capture three images of the scene object. The predetermined distance can be set to ensure that the scene object is fully displayed in the image. Next, each captured image is preprocessed to obtain a target binary image. The preprocessing includes: converting the image to a grayscale image, performing Gaussian blur on the grayscale image to reduce noise and smooth the image, performing adaptive thresholding on the smoothed grayscale image to convert it to a binary image, and performing morphological operations on the binary image to enhance structural continuity. The target binary image can be a binary image after morphological operations. Then, the scene object contour is extracted from each target binary image using the Suzuki 85 boundary tracking algorithm. Finally, the scene object contours corresponding to each target binary image are grouped according to the corresponding shooting viewpoint to obtain scene object contour matching pairs, and each scene object contour matching pair is subjected to similarity measurement to obtain contour similarity. Wherein, each scene object contour matching pair may include the contours of each scene object corresponding to the same shooting viewpoint. The above-mentioned similarity measurement process may include Hu moment calculation, Hu moment normalization and averaging of the normalized results. Specifically, for each scene object contour matching pair, the Hu moment similarity is determined, and the Hu moment similarity is normalized to scale the Hu moment similarity to the (0, 1) interval. Wherein, the normalization process may be to determine the inverse of the sum of a preset value and the Hu moment similarity as the Hu moment similarity. The above-mentioned preset value may be a pre-set value not less than 1. For example, the above-mentioned preset value may be 1. Then, the average value of the obtained normalized results may be determined as the contour similarity.
[0038] Sub-step 2: Determine environmental similarity. First, for each scene object in the target sight intersection object and the non-target sight intersection object, six images are acquired in six directions around the scene object. The ORIENTEDBRIEF (ORB) algorithm is used to determine the feature point information set corresponding to each image. Each feature point information may include a feature point identifier and a feature descriptor. The feature point identifier may be a unique identifier for the feature point. The feature descriptor may represent the local features corresponding to the feature point. The scene object may be removed from the virtual scene, and a virtual camera may be placed at the center of the removed scene object to capture images of the scene object in six directions. Then, for each image capture direction, a brute force matching method based on the Hamming distance is used to perform feature point matching on the target image matching pairs corresponding to the image capture direction to obtain a set of matching point pairs. The target image matching pairs may be images of different scene objects with the same capture direction. It should be noted that a cross-check is performed during the feature point pairing process to retain only matching point pairs that pass the ratio test, thereby filtering out matching point pairs with larger distances. Next, for each target image matching pair corresponding to each image shooting direction, the average distance and maximum distance of the target image matching pair are determined based on the matching point pairs corresponding to the image shooting direction, and the image similarity between the target image matching pairs is determined based on a preset similarity formula. Finally, the maximum value of the image similarities corresponding to each image shooting direction is selected as the environmental similarity. The preset similarity formula can be:
[0039] φ mage =1-d avg / d max .
[0040] Among them, d avg Denotes the average distance. d max Indicates the maximum distance. φ image Indicates image similarity.
[0041] Step 2: Generate object similarity information based on the above-mentioned outline similarity and the above-mentioned environment similarity. The object similarity information may include the scene object identifier corresponding to the above-mentioned target sight intersection object, the scene object identifier corresponding to the above-mentioned non-target sight intersection object, and the object similarity. The object similarity can be generated by the following formula:
[0042]
[0043] Where o1 represents a non-target line of sight intersection object. o2 represents a target line of sight intersection object. Φ(o1, o2) represents the object similarity information. φ c (o1, o2) represents the contour similarity. φ e(o1, o2) represents the environment similarity.
[0044] As an example, Figure 2 An example diagram of the outline and environment image corresponding to an example scene object obtained by the label layout method for rapid positioning of a virtual scene based on user perception of the present disclosure is shown. Figure 2 The image includes a left sub-image, a right sub-image, and a middle sub-image. The microscope in the left sub-image is an example scene object. The middle sub-image is a left-facing outline of the example scene object. The right sub-image is the environment image corresponding to the example scene object. This facilitates determining the similarity between scene objects based on their corresponding outlines and environment images.
[0045] The second sub-step is to determine the current interest level corresponding to the non-target sight intersection objects based on the object similarity information set and the current interest levels corresponding to the target sight intersection object set. The current interest level corresponding to the non-target sight intersection objects can be determined using the following formula:
[0046]
[0047] Among them, o Indicates a non-target line of sight crossing object. o c Indicates the target line of sight that the user last observed during the target time period. t represents the target time period. i represents the sequence number. represents the i-th target sight crossing object observed by the user in the target time period. N represents the duration of the target time period. n represents the number of target sight crossing objects observed by the user in the target time period. OIS(·) represents the current interest level. o , t) indicates that in the time period t, the non-target sight line crosses the object o o The corresponding current interest level. c , t) means that in the time period t, the target sight crosses the object o c The corresponding current interest level. Φ(·) represents the object similarity between two scene objects. T(·) represents the observation time of the user when observing the scene object. Indicates that during the time period t, the user observes the target sight crossing object The observation time.
[0048] The fourth step is to update the current interest of each scene object in the above virtual scene based on the object historical interest information set to generate user interest. Among them, each object historical interest information in the above object historical interest information set can correspond one-to-one to the scene object in the virtual scene. Each object historical interest information can be the information of the user interest of the corresponding scene object in the virtual scene in the previous time period. The previous time period can be a time interval adjacent to the above target time period and earlier than the above target time period. The previous time period can be the same length as the above target time period. For each scene object in the above virtual scene, the current interest of the scene object can be updated to generate user interest by the following formula:
[0049] OIS(o,t)=β(t)*OIS(o,t-1)+(1-β(t))*OIS(o,t).
[0050] Among them, OIS(o, t) on the left side of the above formula represents the user interest level corresponding to the scene object o in the time period t, and OIS(o, t) on the right side represents the unupdated current interest level corresponding to the scene object o in the time period t. t-1 represents the previous time period. OIS(o, t-1) represents the user interest level corresponding to the scene object o in the time period t-1. β(t) represents the attenuation factor, which determines the change of interest level over time. The attenuation factor β(t) can be dynamically adjusted according to the user's behavior pattern to reflect the attenuation or enhancement of interest over time. If the user continuously looks at the same object for a period of time, so that it continues to intersect with the center of the user's vision and is closest, β(t) will decrease to maintain interest. On the contrary, if the user frequently changes focus during the time period t, reducing the reliability of interest, β(t) will increase to enhance the interest of the previous period. In addition, in order to prevent the OIS(o, t) from dropping too quickly when the user's vision is blank, the present disclosure sets β(t) to decay slowly over time. The attenuation factor β(t) can be expressed according to the following formula:
[0051]
[0052] Among them, ew λ×N represents the natural exponential function. λ represents the decay constant. β0 represents the basic decay factor. SD(t) represents the dispersion of the user's gaze direction in time period t. Indicates normalization processing. G t Represents the set of scene objects that the user is looking at during time period t. Represents the empty set.
[0053] Step 102 : Select scene objects that meet a preset interest condition from various scene objects included in the virtual scene as target perception objects to obtain a target perception object set.
[0054] In some embodiments, the execution entity may select scene objects that meet a preset interest condition from the various scene objects included in the virtual scene as target perception objects to obtain a target perception object set. The preset interest condition may be: the user interest of the scene object is a higher user interest among the various user interest degrees. The various user interest degrees may be user interest degrees corresponding to the various scene objects included in the virtual scene. First, according to the user interest degrees corresponding to the scene objects, the various scene objects included in the virtual scene are sorted in descending order to obtain a scene object sequence. Then, the first preset number of scene objects in the scene object sequence are determined as the target perception object set. The preset number may be a pre-set number. For example, the preset number may be 12.
[0055] The above-mentioned steps for generating each user's interest level and their related content, as an inventive feature of the embodiments of the present disclosure, address the second technical issue mentioned in the background technology: the difficulty in timely filtering individual guide tags of user interest. This difficulty is often caused by the following reasons: when clustering is ineffective, objects of different categories are easily mistakenly grouped together, the selected objects are not representative of the entire cluster, and the user's interest level is not considered. By resolving these issues, the effectiveness of timely filtering individual guide tags of user interest can be achieved. To achieve this, first, each scene object in the virtual scene is classified according to whether it is focused on by the user, resulting in scene objects that the user has focused on and those that the user has not focused on. Then, for each scene object focused on by the user, the current interest level is primarily determined by the user's attention duration. For each scene object that the user has not focused on, its current interest level is primarily determined by its similarity to the scene objects that the user has focused on. This facilitates the selection of scene objects of particular interest to the user in the virtual scene. Then, using a state transition model, the current interest level of each scene object is updated based on the user's interest level in the previous time period. This allows us to determine the user's interest in scene objects over time. Finally, we can select the target perception objects that the user is most interested in from among the various scene objects. This facilitates the subsequent timely filtering of corresponding guidance tags based on the target perception objects of interest to the user. Furthermore, because the similarity between scene objects that the user has not paid attention to and scene objects that the user has paid attention to is combined with the outline similarity and background environment similarity between scene objects, the accuracy of the similarity between scene objects can be improved, thereby improving the accuracy of the guidance tag filtering results.
[0056] Step 103: For each target perception object in the target perception object set, perform the following steps:
[0057] Step 1031 : Determine the user perception corresponding to the target guidance tag based on the pre-built perception time mapping function and the user interest level corresponding to the target perception object.
[0058] In some embodiments, the execution entity may determine the user perception force corresponding to the target guide tag based on a pre-constructed perception time mapping function and the user interest level corresponding to the target perception object. The target guide tag may be the guide tag corresponding to the target perception object. The perception time mapping function may characterize the mapping relationship between the guide tag and the user perception time. The user perception time may be the time it takes for the user's gaze point to shift from the target center to the guide tag. The user perception force may be the force used to move the guide tag to the three-dimensional spatial position where the user perception time is minimized.
[0059] Optionally, the above perceptual time mapping function can be constructed as:
[0060] PT(x,y,Z)=f(V(x,y,z),C(x,y,z),VP(x,y,z)).
[0061] Where (x, y, z) represents the three-dimensional coordinates in the world coordinate system. PT(x, y, z) represents the user-perceived time of the guide label at coordinate (x, y, z). V(x, y, z) represents the visibility of the guide label at coordinate (x, y, z). C(x, y, z) represents the color contrast between the guide label at coordinate (x, y, z) and the environment. VP(x, y, z) represents the viewport coordinates corresponding to the guide label at coordinate (x, y, z). f(·) represents the random forest regressor.
[0062] As an example, when in a virtual environment, the Unity engine randomly generates colored guide labels at different positions in the three-dimensional space every 1-5 seconds. The user needs to look at as many guide labels as possible from a predetermined position. When the gaze point falls on these guide labels, the guide labels that the user looks at disappear immediately. If the guide label is not looked at by the user within 10 seconds, the system removes it. The system records the properties of the disappearing label and the time interval between its appearance and disappearance. Therefore, the user perception time, the visibility value of the guide label, the color contrast between the guide label and the environment, and the viewport coordinates corresponding to the guide label can be determined through the following steps:
[0063] In the first step, the time interval from when the guide tag is generated to when the user first sees the guide tag is determined as the user perception time, wherein the user perception time value of the unseen guide tag can be 10 seconds.
[0064] In the second step, the label contour is sampled and the ratio of accessible sampling points to the total sampling points is determined as the visibility value of the guided label.
[0065] The third step is to determine the color contrast between the guide label and the surrounding environment based on the contrast between the guide label and the surrounding background within a 50x50 pixel area. First, convert the label color and background color from RGB (Red, Green, Blue) format to HSV (Hue, Saturation, Value) format. Then, calculate the absolute difference in hue, saturation, and brightness between the two colors. Finally, perform a weighted sum of these absolute differences to obtain the color contrast between the guide label and the surrounding environment.
[0066] The fourth step is to convert the world coordinates of the guide label into viewport coordinates through the viewport transformation matrix.
[0067] In some optional implementations of some embodiments, the execution entity may determine the user perception corresponding to the target guidance tag based on a pre-built perception time mapping function and the user interest corresponding to the target perception object through the following steps:
[0068] The first step is to determine the perception direction vector corresponding to the current frame based on the above-mentioned perception time mapping function. The above-mentioned current frame can be the video frame corresponding to the current moment. The above-mentioned perception direction vector can characterize the direction of the user's perception. The direction of the perception can be determined by the gradient of the above-mentioned perception time mapping function. It should be noted that since the user perception time mapping function is modeled by a random forest regressor, and the random forest model is implicit, it is difficult to directly calculate the gradient. Therefore, the solution of the present disclosure approximates the gradient by fixing the increment along the coordinate axis to determine the change in the user label perception time. Specifically, the following steps can be performed:
[0069] The first sub-step is to determine the reference direction vector of the perception force. The reference direction vector can be the direction vector to be corrected. The gradient calculation method for the X-axis, Y-axis, and Z-axis of the reference direction vector is the same. Take the gradient calculation of the X-axis as an example:
[0070]
[0071] Where Δx represents the change in the coordinates (x, y, z) on the X-axis. (x+Δx, y, z) represents the three-dimensional coordinates of the coordinates (x, y, z) after moving Δx along the positive direction of the X-axis. (x-Δx, y, z) represents the three-dimensional coordinates of the coordinates (x, y, z) after moving Δx along the negative direction of the X-axis. PT x+Δx PT represents the user perceived time of the guide tag located at coordinate (x+Δx, y, z). x-Δx PT represents the user perceived time of the guide tag located at the coordinate (x-Δx, y, z). xrepresents the user perception time of the guide tag located at the coordinate (x, y, z). The positive and negative gradients can be determined by the above formula group:
[0072]
[0073] Among them, PT x+ Indicates a positive gradient. PT x- represents a negative gradient.
[0074] The X-axis component of the perception force direction vector can be determined by the following formula:
[0075]
[0076] Among them, D x Indicates the X-axis component of the perception force direction vector.
[0077] Similarly, the Y-axis component D of the perception direction vector can be determined according to the above gradient calculation method. y and the Z-axis component D z From this, we can get the reference direction vector of the perception force
[0078] In the second sub-step, the perception direction vector corresponding to the current frame is determined based on the perception direction vector corresponding to the previous frame and the reference direction vector using a spherical linear interpolation method. The previous frame may be a video frame at the previous moment. The perception direction vector corresponding to the current frame can be determined using the following formula:
[0079]
[0080] Among them, Frame represents the current frame, and Frame-1 represents the previous frame. Indicates the perception direction vector corresponding to the current frame. Represents the perception direction vector corresponding to the previous frame. θ represents and α represents the interpolation coefficient controlled by θ.
[0081] In practice, user perception usually needs to be calculated in each frame. Since the direction between video frames may change suddenly, making the label position unstable, the present disclosure uses the above-mentioned spherical linear interpolation method to perform temporal smoothing on the direction.
[0082] The second step is to generate the user perception corresponding to the target guidance tag based on the preset perception adjustment coefficient, the user interest corresponding to the target perception object, and the perception direction vector corresponding to the current frame. The user perception corresponding to the target guidance tag can be generated by the following formula:
[0083]
[0084] in, k represents the user perception of the target guidance label in the time period t. perception is a predefined adjustable parameter used to adjust user perception. For example, k perception It can be between 2-10.
[0085] Step 1032: Determine the camera force corresponding to the target guidance tag based on the viewport coordinates corresponding to the target guidance tag.
[0086] In some embodiments, the execution entity may determine the camera force corresponding to the target guide tag based on the viewport coordinates corresponding to the target guide tag in various ways. The camera force may be a force used to keep the guide tag within the user's field of view.
[0087] In some optional implementations of some embodiments, the execution entity may determine the camera force corresponding to the target guidance tag based on the viewport coordinates corresponding to the target guidance tag through the following steps:
[0088] The first step is to determine the initial horizontal component, initial vertical component, and initial depth component corresponding to the target guide tag based on the preset horizontal coordinate boundary distance, vertical coordinate boundary distance, first depth coordinate boundary distance, second depth coordinate boundary distance, and the viewport coordinate corresponding to the target guide tag. The horizontal coordinate boundary distance can be the preset boundary distance at which the camera force begins to take effect on the X-axis. For example, the viewport coordinate range on the X-axis is (0, 1). Since the guide tag actually has a size, a transition zone needs to be set for the guide tag. When the boundary distance is adjusted to 0.2, the camera force range on the X-axis will become (0.2, 0.8). The vertical coordinate boundary distance can be the preset boundary distance at which the camera force begins to take effect on the Y-axis. The first depth coordinate boundary distance can be the preset minimum boundary distance at which the camera force begins to take effect on the Z-axis. The second depth coordinate boundary distance can be the preset maximum boundary distance at which the camera force begins to take effect on the Z-axis. The initial horizontal component can be the initial value of the camera force in the X-axis direction. The initial vertical axis component can be the initial value of the camera force in the Y-axis direction. The initial depth axis component can be the initial value of the camera force in the Z-axis direction. The initial horizontal axis component corresponding to the target guide tag can be determined by the following formula:
[0089]
[0090] in, Represents the initial horizontal axis component. x represents the position of the target guide label in the X-axis direction in the viewport coordinate system. mx Indicates the horizontal axis coordinate boundary distance. Represents the unit vector in the X-axis direction in the camera coordinate system. The calculation of the above initial vertical axis component can refer to the steps for generating the above initial horizontal axis component, which will not be repeated here. The initial depth axis component corresponding to the target guide label can be determined by the following formula:
[0091]
[0092] in, Represents the initial depth axis component. z represents the position of the target guide label in the Z-axis direction in the viewport coordinate system. Indicates the first depth coordinate boundary distance. Indicates the second depth coordinate boundary distance. The unit vector representing the Z-axis direction in the camera coordinate system.
[0093] The second step is to generate a camera force corresponding to the target guidance tag based on a preset camera force adjustment coefficient, the initial horizontal axis component, the initial vertical axis component, and the initial depth axis component. The camera force adjustment coefficient may be a parameter used to adjust the camera force. The camera force corresponding to the target guidance tag may be determined using the following formula:
[0094]
[0095] in, k represents the camera force corresponding to the target guidance label. camera Indicates the camera force adjustment coefficient. Represents the initial vertical axis component.
[0096] Step 1033 : Determine the sum of the user perception force and the camera force as the user perceived attractiveness corresponding to the target guidance tag.
[0097] In some embodiments, the execution entity may determine the sum of the user perception force and the camera force as the user perceived attractiveness corresponding to the target guidance tag.
[0098] Step 1034 : Generate a tag force based on the dynamic adjustment force corresponding to the target guidance tag and the user-perceived attraction.
[0099] In some embodiments, the execution entity may generate a tag force based on the dynamic adjustment force corresponding to the target guide tag and the user-perceived attraction. The dynamic adjustment force may be a pre-generated force that determines the positional relationships between guide tags in the virtual scene and between the guide tags and the target perceived object based on a dynamic potential field. The tag force may be determined as the sum of the dynamic adjustment force and the user-perceived attraction.
[0100] Optionally, the dynamic adjustment force corresponding to the target guide tag may be composed of a tag spring force, a region correction force, a tag repulsion force, an obstacle repulsion force, a damping force, and a wire crossing force. The tag spring force may be a force that maintains an appropriate distance between the guide tag and the marked scene object according to Hooke's law. When the guide tag is far away from the marked object, the tag spring force pulls the guide tag back; otherwise, it rebounds the guide tag. The region correction force may be a force that keeps the guide tag within a certain area to prevent system crashes. The tag repulsion force may be a repulsive force that prevents collisions or occlusions between tags. The obstacle repulsion force may be a repulsive force that prevents collisions between the guide tag and scene objects. The damping force may be a force that is inversely proportional to speed to prevent image jitter and eliminate system crashes caused by rapid movement. The wire crossing force may be a force that prevents the wires of each guide tag from overlapping. The sum of the tag spring force, the region correction force, the tag repulsion force, the obstacle repulsion force, the damping force, and the wire crossing force may be determined as the dynamic adjustment force.
[0101] As an example, Figure 3 The figure shows a force diagram of a guide label corresponding to an example scene object in a label layout method for rapid positioning of a virtual scene based on user perception according to the present disclosure. Figure 3 This includes the forces and directions acting on the microscope and guide labels of the example scene objects. User-perceived attraction is composed of user-perceived force and camera force. The guide labels of the example scene objects can be laid out based on both user-perceived attraction and dynamic adjustment force.
[0102] Step 104 : Based on the determined forces of each tag, the positions of each target guidance tag corresponding to the target perception object set are updated to obtain a tag position update information set for arranging each target guidance tag.
[0103] In some embodiments, the execution entity may update the positions of the target guidance tags corresponding to the target perception object set based on the determined forces of the tags, and obtain a tag position update information set for use in the layout of the target guidance tags. The tag position update information in the tag position update information set may be information on the optimized three-dimensional spatial position of the guidance tags. Specifically, the following steps may be performed:
[0104] In the first step, for each tag force in each tag force, perform the following steps:
[0105] In the first sub-step, Newton's second law is used to determine the acceleration and velocity of the target guidance tag according to the tag force.
[0106] In the second sub-step, the displacement value of the target guidance tag is determined according to the acceleration and velocity of the target guidance tag through the kinematic equation.
[0107] The third sub-step is to determine the optimized viewport coordinates according to the above displacement value and the viewport coordinates corresponding to the target guide label.
[0108] The fourth sub-step is to determine the scene object identifier corresponding to the target guidance tag and the optimized viewport coordinates as tag position update information.
[0109] In the second step, the information set is updated according to the label position, and each target guide label is placed at the corresponding optimized viewport coordinates.
[0110] As an example, Figure 4 FIG2 shows an example of a label layout before the label layout method for rapid virtual scene positioning based on user perception disclosed in the present invention is used. Figure 4 This is a virtual scene of a scientific laboratory, which includes but is not limited to microscopes, crucibles, and other scene objects. Figure 5 The following diagram shows an example of label layout using the label layout method for rapid virtual scene positioning based on user perception disclosed in the present invention. Figure 5 and Figure 4 Corresponding to the same virtual scene. Figure 5 The object (microscope) framed in the middle is the target sight intersection object that intersects with the center of the user's sight. Figure 5 There are three dark guide labels in the figure, namely Microscope 0, Microscope 1 and Microscope 3. Microscope 0 is the guide label corresponding to the target sight intersection object (microscope). Microscope 1 and Microscope 3 are the guide labels corresponding to the two scene objects (microscopes) with the highest similarity to the target sight intersection object in the virtual scene. In addition, Figure 5 compared to Figure 4 , the number of guide tags in the virtual scene is reduced, and the positions of the same guide tags are optimized. This makes it easier to shorten the execution time of the target search task and improve the user's positioning efficiency in real time.
[0111] Embodiments of the present disclosure have the following advantageous effects: Through the user-perception-based label layout method for rapid virtual scene location, some embodiments of the present disclosure enable users to promptly perceive each guide label. Specifically, the reason why each guide label is difficult for the user to visually perceive is that when objects in a scene are compactly arranged, placing guide labels at local three-dimensional spatial locations of the objects can easily lead to occlusion between the guide labels, making each guide label difficult to visually perceive. Based on this, some embodiments of the present disclosure provide a user-perception-based label layout method for rapid virtual scene location. First, the user's interest level for each scene object in the virtual scene during a target time period is determined. Each scene object corresponds to a guide label, and attribute information associated with the guide label includes viewport coordinates, label visibility, and environmental contrast. This allows the user's interest level in each object in the virtual scene to be determined. Second, scene objects that meet preset interest criteria are selected from the various scene objects included in the virtual scene as target perception objects, thereby obtaining a target perception object set. This facilitates the subsequent layout of guide labels for the objects of greatest interest to the user. Then, for each target perception object in the target perception object set, the following steps are performed: based on a pre-constructed perception time mapping function and the user interest level corresponding to the target perception object, the user perception force corresponding to the target guide tag is determined, wherein the target guide tag is the guide tag corresponding to the target perception object, the perception time mapping function represents the mapping relationship between the guide tag and the user perception time, and the user perception force is the force used to move the guide tag to the three-dimensional spatial position with the minimum user perception time; based on the viewport coordinates corresponding to the target guide tag, the camera force corresponding to the target guide tag is determined, wherein the camera force is the force used to keep the guide tag within the user's field of view; the sum of the user perception force and the camera force is determined as the user perceived attraction corresponding to the target guide tag; based on the dynamic adjustment force corresponding to the target guide tag and the user perceived attraction, a tag action force is generated, wherein the dynamic adjustment force is a pre-generated force that determines the positional relationship between each guide tag in the virtual scene and between the guide tag and the target perception object according to the dynamic potential field. Thus, for each target perception object that the user is most interested in, the user perceived attraction can be determined based on the user interest level, the perception time, and the viewport constraint, and combined with the dynamic adjustment force to ultimately obtain the tag action force used for tag layout. Finally, based on the determined tag interaction forces, the positions of the target guidance tags corresponding to the target perception object set are updated to obtain a tag position update information set for the layout of the target guidance tags. Thus, based on the tag interaction forces, the guidance tags of interest to the user can be placed in the three-dimensional spatial location that minimizes the user's perception time.Therefore, in some embodiments of the present disclosure, the label layout method for rapid virtual scene positioning based on user perception determines the user perceived attractiveness according to the user's interest in the scene object of interest and the user's perception time of the guide label, which can facilitate the layout of the guide label in a three-dimensional space position that is more easily perceived by the user. As a result, the user can perceive each guide label of interest and the corresponding scene object in a timely manner. In addition, because the guide labels of virtual objects that the user is more interested in can be placed in a position with a shorter user perception time, the execution time of the target search task can be shortened, the user's positioning efficiency can be improved in real time, the user's task burden can be reduced, the usability can be improved, and the incidence of motion sickness can be reduced.
[0112] The technical contents not elaborated in detail in the present invention belong to the common knowledge of those skilled in the art.
[0113] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A label layout method for rapid positioning of a virtual scene based on user perception, comprising: Determining the user interest level corresponding to each scene object in the virtual scene during a target time period, wherein each scene object corresponds to a guide label, and attribute information corresponding to the guide label includes viewport coordinates, label visibility, and environment contrast; Selecting scene objects that meet a preset interest condition from each scene object included in the virtual scene as target perception objects to obtain a target perception object set; For each target perception object in the target perception object set, perform the following steps: Determining the user perception force corresponding to the target guidance tag based on a pre-constructed perception time mapping function and the user interest level corresponding to the target perception object, wherein the target guidance tag is the guidance tag corresponding to the target perception object, the perception time mapping function represents the mapping relationship between the guidance tag and the user perception time, and the user perception force is the force used to move the guidance tag to the three-dimensional spatial position where the user perception time is minimized; Determining a camera force corresponding to the target guide tag based on the viewport coordinates corresponding to the target guide tag, wherein the camera force is a force used to keep the guide tag within the user's field of view; determining the sum of the user perception force and the camera force as the user perceived attractiveness corresponding to the target guidance tag; Generate a tag force based on the dynamic adjustment force corresponding to the target guide tag and the user-perceived attraction, wherein the dynamic adjustment force is a pre-generated force that determines the positional relationship between each guide tag in the virtual scene and between the guide tag and the target perception object according to the dynamic potential field; Based on the determined forces of each tag, the positions of each target guidance tag corresponding to the target perception object set are updated to obtain a tag position update information set for arranging the target guidance tags.
2. The method according to claim 1, wherein The determining of the user interest level corresponding to each scene object in the virtual scene within the target time period includes: Classifying each scene object in the virtual scene to obtain a target sight intersection object set and a non-target sight intersection object set, wherein each target sight intersection object is an object gazed at by the user within a target time period, and each target sight intersection object corresponds to a gaze duration of the user; For each target sight intersection object in the target sight intersection object set, determining a ratio of a user's gaze duration corresponding to the target sight intersection object to a preset duration as a current interest level of the target sight intersection object, wherein the preset duration is the duration of the target time period; For each non-target sight-intersecting object in the non-target sight-intersecting object set, perform the following steps: Performing similarity analysis on the non-target sight-intersection object and each target sight-intersection object in the target sight-intersection object set to obtain an object similarity information set; Based on the object similarity information set and the current interest levels corresponding to the target sight intersection object set, the current interest level corresponding to the non-target sight intersection object is determined; based on the object historical interest level information set, the current interest level of each scene object in the virtual scene is updated to generate user interest level, wherein each object historical interest level information in the object historical interest level information set corresponds to a scene object in the virtual scene, and each object historical interest level information is information on the user interest level of the corresponding scene object in the virtual scene in the previous time period.
3. The method according to claim 2, wherein: The performing similarity analysis on the non-target sight-intersection object and each target sight-intersection object in the target sight-intersection object set to obtain an object similarity information set includes: For each target sight intersection object in the target sight intersection object set, perform the following steps: Determining contour similarity and environment similarity between the target sight-intersecting object and the non-target sight-intersecting object; Based on the outline similarity and the environment similarity, object similarity information is generated, wherein the object similarity information is generated by the following formula: Among them, o1 represents the non-target sight intersection object, o2 represents the target sight intersection object, Φ(o1, o2) represents the object similarity information, φ c (o1, o2) represents the contour similarity, φ e (o1, o2) represents the environmental similarity.
4. The method according to any one of claims 1 to 3, wherein: The perception time mapping function is: PT(x,y,z)=f(V(x,y,z),C(x,y,z),VP(x,y,z)), where (x, y, z) represents the three-dimensional coordinates in the world coordinate system, PT(x, y, z) represents the user-perceived time of the guide label at the coordinate (x, y, z), V(x, y, z) represents the visibility of the guide label at the coordinate (x, y, z), C(x, y, z) represents the color contrast between the guide label at the coordinate (x, y, z) and the environment, VP(x, y, z) represents the viewport coordinates corresponding to the coordinate (x, y, z), and f(·) represents the random forest regressor.
5. The method according to claim 4, wherein The determining of the user perception corresponding to the target guidance tag based on the pre-built perception time mapping function and the user interest corresponding to the target perception object includes: Determining a perception force direction vector corresponding to a current frame based on the perception time mapping function; Based on a preset perception adjustment coefficient, the user interest level corresponding to the target perception object, and the perception direction vector corresponding to the current frame, the user perception corresponding to the target guidance tag is generated.
6. The method according to claim 5, wherein: The determining, based on the viewport coordinates corresponding to the target guide tag, the camera force corresponding to the target guide tag includes: Determine an initial horizontal axis component, an initial vertical axis component, and an initial depth axis component corresponding to the target guide tag based on a preset horizontal axis coordinate boundary distance, a vertical axis coordinate boundary distance, a first depth coordinate boundary distance, a second depth coordinate boundary distance, and the viewport coordinates corresponding to the target guide tag; A camera force corresponding to the target guidance tag is generated based on a preset camera force adjustment coefficient, the initial horizontal axis component, the initial vertical axis component, and the initial depth axis component.
7. The method according to claim 6, wherein: The dynamic adjustment force corresponding to the target guidance tag is composed of tag spring force, area correction force, tag repulsion force, obstacle repulsion force, damping force and lead crossing force.
Citation Information
Patent Citations
Dynamic layout optimization of annotation tags in volume rendering
CN116580398A
Label display method and system in three-dimensional scene and virtual engine
CN118051159A